Intestinal tract video target detection method and system based on memory bank
By introducing memory bank-based feature enhancement and optical flow memory assistance modules into the intestinal video object detection system, the problem of low object detection accuracy in colonoscopic videos is solved, and more efficient and accurate colonoscopic object detection is achieved.
Patent Information
- Application Number
- CN202510105722.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-01-23
AI Technical Summary
In the prior art, when performing intestinal object detection based on colonoscopy video, there is a problem that high time cost, experience-dependent and poor target recognition accuracy, especially in the case of picture jitter caused by camera movement and rapid target movement.
A memory bank-based intestinal video object detection system is adopted, including a memory storage module, a memory update module and an optical flow memory auxiliary module. The current image frame is characterized by the image characteristics of multiple memory frames stored in the memory library, and the optical flow method is used to track the optical flow key points to generate an attention feature map to enhance the image.
The efficiency and accuracy of colonoscopic object detection are improved, and the timing characteristics of colonoscopic video and representative image frame information before the current frame can be better utilized, which reduces the dependence on manual screening, and enhances the ability to model target motion trends.
Smart Images

Figure CN120032107A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to an intestinal video target detection system based on a memory library. Background Art
[0002] The unique temporal information of colonoscopy videos can well reflect the motion trajectory of foreground targets and the changes in colonoscopy screening stages, which is very helpful for better colonoscopy target detection.
[0003] The traditional method of target detection based on colonoscopy video is manual screening, which is time-consuming and the accuracy of detection is heavily dependent on the experience of the screener. Furthermore, since colonoscopy video is different from common fixed camera video, camera movement will cause rapid image jitter and rapid movement of the target, resulting in huge differences in adjacent foreground features, and introduce a large number of irrelevant interference frames, which increases the difficulty of applying traditional image or video detection methods to the field of colonoscopy video. Existing intestinal target methods designed based on deep learning are only designed at the image level, or simple adjacent frame interaction is performed at the video level for feature enhancement, resulting in poor target recognition accuracy.
[0004] Therefore, a new technical solution for intestinal target detection based on colonoscopy video is urgently needed. Summary of the invention
[0005] The present invention provides a method and system for intestinal video target detection based on a memory library, so as to solve at least one defect in the prior art.
[0006] In a first aspect, the present invention provides a memory-based intestinal video target detection system, comprising: a memory storage module, a memory update module, and an optical flow memory auxiliary module; The memory storage module is used to enhance the features of the current image frame using the memory frames in the memory bank during the target detection phase of the current image frame; the image features of the multiple memory frames stored in the memory bank are the image features of the multiple image frames adjacent to the current image frame; The memory updating module is used to perform feature similarity and importance analysis on the image features of the current image frame and the image features of the memory frames in the memory bank, and decide whether to update the memory bank according to the analysis results; The optical flow memory auxiliary module is used to use the optical flow method to track the changes of optical flow key points in the video sequence to generate an attention feature map, and perform image enhancement on the current image frame.
[0007] According to a memory-based intestinal video target detection system provided by the present invention, the memory storage module is specifically used to: store image features of a fixed number of previous memory frames adjacent to a current image frame to be detected into the memory bank, and during the detection process of the current image frame, based on the image features of the memory frames in the memory bank, use the attention mechanism to enhance the features of the current image frame.
[0008] According to a memory-based intestinal video target detection system provided by the present invention, the memory storage module uses an attention mechanism to enhance the features of the current image frame based on the image features of the memory frames in the memory library, including: generating a query based on the image features of the current image frame, and generating corresponding key-value pairs based on the image features of the memory frames in the memory library; performing cross-attention calculations on the query of the current image frame and the key values of each memory frame in the memory library to obtain an attention score corresponding to each memory frame; based on the attention score, using the image features of the memory frames in the memory library to enhance the features of the current image frame to generate enhanced features.
[0009] According to a memory-based intestinal video target detection system provided by the present invention, the memory update module is specifically used to: calculate cosine similarity based on image features of the current image frame and the memory frame in the memory library; obtain an affinity matrix based on all cosine similarities; and analyze the affinity matrix to update the memory library.
[0010] According to a memory-based intestinal video target detection system provided by the present invention, the memory update module analyzes the affinity matrix to update the memory bank, including: determining the maximum value in the affinity matrix; when the maximum value is greater than or equal to a preset threshold, using the image features of the current image frame and the memory frame corresponding to the maximum value to perform feature fusion to generate enhanced image features of the memory frame to update the memory bank; when the maximum value is less than the preset threshold, comprehensively evaluating the importance of each memory frame in the memory bank to the current image frame to obtain a comprehensive evaluation score; replacing the image features of the memory frame with the lowest comprehensive evaluation score with the image features of the current image frame to update the memory bank.
[0011] According to a memory-based intestinal video target detection system provided by the present invention, the memory update module comprehensively evaluates the importance of each memory frame in the memory bank to the current image frame to obtain a comprehensive evaluation score, including: determining the comprehensive evaluation score according to the contribution value of the memory frame to the current frame detection, the existence duration of the memory frame, and whether there is a target mutation in the memory frame; wherein the contribution value of the memory frame to the current image frame detection is determined based on the attention score between the memory frame and the current image frame.
[0012] According to a memory-based intestinal video target detection system provided by the present invention, the optical flow memory auxiliary module is specifically used to: determine whether there is a target at the end of the target detection of the previous frame, extract optical flow key points if there is a target, perform optical flow tracking on the optical flow key points to obtain optical flow vectors; analyze the speed and angle of the optical flow vector to generate an attention feature map; and use the attention feature map to perform image enhancement on the current image frame.
[0013] According to the intestinal video target detection system based on a memory library provided by the present invention, the video sequence is a colonoscopy video sequence.
[0014] In a second aspect, the present invention further provides a method for detecting intestinal video targets based on a memory library, using any of the intestinal video target detection systems described above, comprising: Use the optical flow memory auxiliary module to enhance the current image frame; Using the memory storage module to perform feature enhancement on the current image frame using the memory frames in the memory bank during the target detection phase of the current image frame; the memory bank stores image features of the memory frames adjacent to the current image frame; The memory update module is used to perform similarity and importance analysis on the image features of the current image frame and the image features of the memory frames in the memory bank, and decides whether to update the memory bank based on the analysis results.
[0015] In a third aspect, the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of any of the above-described memory-based intestinal video target detection methods are implemented.
[0016] The intestinal video target detection method and system based on memory library provided by the present invention have the following beneficial effects compared with the prior art: (1) The present invention designs a complete set of colonoscopy video detection processes, and provides a complete system for encoding memory features (image features) from the memory bank to construct a memory bank, update the memory bank, interact with the memory bank and the current frame, and perform optical flow motion statistical analysis, thereby improving the detection efficiency and accuracy of colonoscopy targets.
[0017] (2) The present invention constructs a memory library that can well store the temporal and spatial information of image frames at past time points. When detecting the next frame, the detection results of the previous frames are used to perform cross-attention calculations on the features of the current frame and the features of the image frames already in the memory library, thereby learning the temporal information of the video sequence.
[0018] (3) The present invention fully considers the problem that the memory library needs to be representative. Therefore, when updating the memory library, the traditional first-in-first-out principle is not adopted. Instead, a set of evaluation criteria for memory library updates is specially designed to evaluate the importance of existing image frames in the memory library and the specificity of the current truth, so as to determine whether the current frame needs to enter the memory library and whether the existing image frames in the memory library need to be removed.
[0019] (4) The present invention takes into account that videos are dynamic and that targets have a certain motion trend as video image frames progress. Optical flow memory analysis is introduced to assist in modeling the target's motion as the video progresses. By statistically analyzing the speed and direction of the optical flow vector, the motion of the target is learned and an attention map is generated through the optical flow analysis results to enhance image features. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0021] Figure 1 It is a schematic diagram of a framework for performing target detection in a colonoscopy video by constructing a memory library designed in an embodiment of the present invention; Figure 2 It is a schematic diagram of a process for performing target detection in a colonoscopy video by constructing a memory library designed in an embodiment of the present invention; Figure 3 It is a schematic diagram of the process of updating the memory bank provided by the present invention; Figure 4 It is a schematic diagram of the process of optical flow memory assistance provided by the present invention; Figure 5 It is a schematic diagram of image enhancement using optical flow vectors in the present invention; Figure 6 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0022] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0023] It should be noted that, in the description of the embodiments of the present invention, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "include one..." do not exclude the existence of other identical elements in the process, method, article or device including the elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to the specific circumstances.
[0024] The present invention provides a colonoscopy video target detection method and system based on a memory library. The system can make good use of the temporal characteristics of the colonoscopy video and fully utilize the information of the representative image frames before the current frame to assist the detection of the current frame, thereby fully exploring the motion and timing information, achieving better interference avoidance and improving detection performance.
[0025] Figure 1 is a schematic diagram of a framework for performing target detection in colonoscopy video by constructing a memory library designed in an embodiment of the present invention. Figure 2 FIG. 1 is a flow chart of target detection in colonoscopy video by constructing a memory library designed in an embodiment of the present invention, see Figure 1 and Figure 2 As shown, the present invention provides a complete set of memory library construction, memory library interaction, memory library update, optical flow memory-assisted video detection processing based on memory guidance: (1) Memory storage module: used to store image features of a fixed number of memory frames adjacent to the current image frame to be detected into a memory bank, and during the detection process of the current image frame, based on the image features of the memory frames in the memory bank, use an attention mechanism to enhance the features of the current image frame; the image features of the multiple memory frames stored in the memory bank are the image features of the multiple image frames adjacent to the current image frame.
[0026] To achieve video sequence modeling based on memory and enhance the image features of the current frame through the cross-attention mechanism, we first need to build a memory to store the image features of the past in the video. The design goal is to store the image features of the memory frame in the memory, and to enhance the detection effect of the current image by retrieving and updating the image features in the memory. This not only relies on the information of the current frame, but also can use the features of the past image frames to improve the accuracy and robustness of image understanding.
[0027] In this example, each frame of the video will be input into a pre-trained feature encoder network to extract high-dimensional feature vectors. These high-dimensional spatial features will pass through a decoder network after being enhanced with the existing information features in the memory bank from memory attention and optical flow assisted analysis. The output of the decoder will be used to generate detection results after subsequent processing. On the other hand, it will pass through a memory encoder network to re-encode the features of the current frame containing the current frame detection information and store them in the memory bank, so as to serve as memory attention prompts for subsequent frame detection during subsequent frame detection.
[0028] Regarding the process of using the information already in the memory bank to prompt the current frame, specifically, it is to perform cross-attention analysis on the image features of the current frame after encoding and the spatial features of the past image frames stored in the memory bank. Each time a new image is processed, the features of the current image are extracted by the previous encoder, and the features of the memory image frame stored in the memory bank are encoded and generated by the memory encoder when the memory bank is constructed. The cross-attention analysis operation is performed between the current image frame and each frame stored in the memory bank. The features of the current frame are used as queries, and the features in the memory bank are used as keys and values. Through the cross-attention mechanism, the attention score is calculated, so that the effective spatiotemporal information in the memory frame with great reference significance is introduced into the current frame detection to update the features of the current image using the memory frame information.
[0029] Among them, in the detection process of the current image frame, based on the image features of the memory frame in the memory bank, the attention mechanism is used to enhance the features of the current image frame, including: Generate a query based on the image features of the current image frame, and generate corresponding key-value pairs based on the image features of the memory frames in the memory bank; perform cross-attention calculation on the query of the current image frame and the key values of each memory frame in the memory bank to obtain the attention score corresponding to each memory frame; based on the attention score, use the image features of the memory frames in the memory bank to enhance the features of the current image frame and generate enhanced features.
[0030] The specific implementation is as follows: First, the image features of the current frame Generate a query in For each frame feature map in the memory library, the feature extractor is used to obtain the corresponding key-value pair and ; Then the query of the current frame is matched with the key of each frame in the history library (ie, memory library) to get the attention score. The specific calculation method is as follows:
[0031] Among them, the subscript t represents the image frame that needs to be detected at time t, and n represents the nth memory frame in the memory bank. express Dimension. The attention score for the cross-attention interaction between the image at time t and the n-th memory frame.
[0032] The attention score is normalized by softmax to obtain the attention weight of a frame. The weight is used to weight the features of the current frame, and the features of the nth memory frame are optimized for the features of the current frame. Specifically, the normalized attention weighted score is applied to the value in the history library (memory library). superior:
[0033] The final output is obtained by weighted summation. This output represents the characteristics of the current frame, which has been combined with the information of the memory frame and weighted according to the correlation. When the actual calculation is performed, the current frame needs to interact with each frame in the memory bank, so the key values in the memory bank can be spliced together into a matrix to unify the attention calculation:
[0034]
[0035] Replacing k and v in the above formulas 1 and 2 with K and V can directly perform memory attention calculation on all image frames in the memory bank, and use the results of cross attention to weight the current frame to fuse the features of the past frames to obtain the final enhanced features. , N is the number of memory frames in the memory bank.
[0036] (2) Memory update module: used to analyze the feature similarity and importance between the image features of the current image frame and the image features of the memory frames in the memory bank, and decide whether to update the memory bank based on the analysis results.
[0037] Memory update mainly occurs when one frame completes detection and before entering the next frame detection. Because the detection result of the current frame is closer to the next frame detection, in principle it can provide more references. The memory frames that already exist in the memory bank have existed for a long time and may have lost their representativeness in subsequent frame detections due to the long distance. At the same time, due to the existence of irrelevant redundant frames, the image frames already stored in the memory bank may not be representative and need to be removed immediately. Therefore, it is necessary to verify the importance of memory frames and eliminate unimportant memory frames.
[0038] The main reason for updating the memory bank is that the length of the memory bank is limited by the memory. When using the memory image frame features in the memory bank for feature interaction, the speed and efficiency of the interaction are limited by the computing medium. Therefore, the number of image frames that can be stored in the memory bank must be limited. On the other hand, as a long video sequence, a complete colonoscopy video sequence is split into image frames, often with tens of thousands of images. The colonoscopy video goes through multiple stages such as entering the anus, anal canal, intestines, reaching the ileocecal end of the colonoscopy, and exiting. The intestinal image states observed by the colonoscope at different stages are significantly different. Therefore, there is no need to use images from a longer distance to interact with the current frame. Only the temporal characteristics of the neighborhood video segments need to be considered. Therefore, it is necessary to limit the size of the memory bank and update it.
[0039] There are two aspects to updating the memory library: First, evaluate the importance of the existing memory image frames in the memory library. First, the importance can be directly reflected by the contribution value of the memory frame when it is detected in the current frame, and the contribution value can be directly reflected by the similarity score matrix of the attention interaction between the memory library and the current frame. In addition, generally speaking, the longer the image frame is from the current frame, the more limited its reference significance to the current image frame. At the same time, if a state change such as the disappearance and appearance of a target occurs between two frames, it also means that the two adjacent frames reflect an important stage, and the corresponding adjacent frames need to be deleted more carefully.
[0040] Specifically, the memory update module is used to: calculate cosine similarity based on image features of the current image frame and memory frames in the memory library; obtain an affinity matrix based on all cosine similarities; and analyze the affinity matrix to update the memory library.
[0041] The memory update module analyzes the affinity matrix to update the memory library, including: determining the maximum value in the affinity matrix; when the maximum value is greater than or equal to a preset threshold, using the image features of the current image frame and the memory frame corresponding to the maximum value to perform feature fusion to generate enhanced image features of the memory frame to update the memory library; when the maximum value is less than the preset threshold, comprehensively evaluating the importance of each memory frame in the memory library to the current image frame to obtain a comprehensive evaluation score; replacing the image features of the memory frame with the lowest comprehensive evaluation score with the image features of the current image frame to update the memory library.
[0042] Figure 3 is a schematic diagram of the process of updating the memory provided by the present invention, see Figure 3 , the following combination Figure 3 The memory update module of the present invention is further described as follows: First, the detection result of the current frame is processed by the memory encoder to obtain the memory feature of the current frame. (image features of the current frame), the memory features of the current frame and the memory features of the memory frame (Image features of memory frames) Calculate cosine similarity respectively :
[0043] Concatenate the cosine similarities calculated for the current frame and all frames in the memory bank to get the affinity matrix . It reflects the similarity between the memory features of the current frame and the image features already in the memory library. Analyze the extreme values of the affinity matrix. If the maximum value α exceeds 0.7 (the preset threshold), it means that the current frame and the memory frame already in the memory library are similar. If they are highly similar, the features of the current frame are directly fused with the corresponding frame to obtain enhanced memory features. :
[0044] If it is not satisfied, it means that there is no highly similar image in the memory features of the current frame in the memory features of the existing memory library. Therefore, it is necessary to evaluate the importance of the memory frames in the memory library and determine the comprehensive evaluation score, so as to eliminate the least important memory frame (that is, the one with the lowest comprehensive evaluation score) and update the memory features of the current frame to the memory library.
[0045] Among them, for the memory frames already existing in the memory library, the method for determining the comprehensive evaluation score includes: determining the comprehensive evaluation score according to the contribution value of the memory frame to the current frame detection, the existence length of the memory frame, and whether there is a target mutation in the memory frame.
[0046] Specifically, the contribution value of the memory frame can be directly based on the attention score between the memory frame and the current image frame Determine that the existence duration of the memory frame is determined by the difference between the frame number of the memory frame in the entire video and the frame number of the current image frame. Reflection, whether there is a target mutation in the memory frame is analyzed and constructed through the detection results of each frame in the history , The value of is determined by the detection results of the current frame and the previous frame. If there is a target in the previous frame but not in the next frame, or if there is no target in the previous frame but there is a target in the current frame, then the corresponding position of the identification matrix of the existence of the memory feature stored in the current frame is stored with 1, otherwise it is 0.
[0047] Finally, the three indicators are combined to obtain a comprehensive evaluation score for the existing memory frames in the memory library: , (6) in is the weight value of each indicator, which can be 0.4, 0.4, and 0.2 respectively. The memory library can be updated by removing the memory frame with the lowest score and replacing it with the memory feature of the current frame.
[0048] (3) Optical flow memory auxiliary module: Use the optical flow method to track the changes of optical flow key points in the video sequence to generate an attention feature map and perform image enhancement on the current image frame.
[0049] At the end of the target detection in the previous frame, determine whether there is a target. If there is a target, extract the optical flow key points, then perform optical flow tracking on the optical flow key points to obtain the optical flow vector; analyze the speed and angle of the optical flow vector to generate an attention feature map; use the attention feature map to perform image enhancement on the current image frame.
[0050] There are a large number of irrelevant image frames in the colonoscopy video detection process. At the same time, because the camera is movable, the images between adjacent frames appear to be roughly similar, but there will be a small displacement due to the movement of the camera. If the deviation of the detection foreground caused by this tiny movement can be well captured, it will be of great help for subsequent detection. By learning this continuity trend in movement, the detection network's attention on the entire image can be adjusted, so that the previous detection results can provide guidance for the areas that are more concerned in subsequent detections.
[0051] Specifically, when the output detects a target, starting from the current frame, the LK algorithm is used inside the detection frame to first extract the optical flow key points required for LK optical flow tracking, and the results of the optical flow key points extracted by the optical flow operation of this frame are stored. Starting from the next frame, the output of the key points extracted from the detection results of the previous frame is used to track the movement of the optical flow key points of the previous frame in this frame using the LK optical flow method. The optical flow key points of the previous frame are linked to the optical flow key points of the current frame to obtain a series of optical flow vectors, and the changes in the speed and direction of the optical flow vectors are counted to analyze the movement trend of the target in the previous frame in the current frame.
[0052] There are two main situations that need to be distinguished and paid special attention to when the target moves between image frames: one is that the target moves and its position changes but it still exists; the other is that the target disappears in this frame, and the video enters a pure background picture or an irrelevant interference frame such as jitter, water flow, etc. These two situations show differences in optical flow statistics, and there are also differences in the subsequent use of optical flow information. In order to better use the information obtained from the optical flow for analysis, it is necessary to first analyze these two situations.
[0053] Through experimental observation and analysis of the principle of optical flow method, optical flow describes the movement change of pixels in an image between two consecutive frames, reflecting the displacement of each pixel or key point in the image in time. The direction of optical flow tells us the moving direction of the target object in the image. If the optical flow directions of multiple key points are consistent, it means that these points may belong to the same object or the same moving direction; the size of the optical flow reflects the moving speed of the target between two frames. Fast moving targets will have larger optical flow, while stationary or slow moving targets will have smaller optical flow. If the direction and speed of the optical flow are consistent in the image (for example, objects in the entire scene move in a certain direction), then it can be considered that these movements are caused by a unified target, and if the optical flow changes in a disorderly manner, it may be a change in the background or a scene without a target. If there is a sudden change in the change of optical flow in the image (for example, the direction of the optical flow changes very irregularly and the speed increases sharply), this may be due to rapid movement in the scene (such as camera movement, object rapid change of direction, etc.). This sudden change can help detect abnormalities in the image and process them. Therefore, the speed and direction of the optical flow of all nodes are statistically analyzed.
[0054] Since there will be a big difference between the speed and direction of the optical flow when entering a frame without a target from an image frame with a target, the mean of the angle change of the direction of all optical flows and the mean of the speed change are calculated, and the standard deviation is calculated to reflect the fluctuation, and a threshold is set as the judgment standard; if there is the same target between the two frames, the trend of the speed and direction change of the optical flow should show consistency with the movement of the target, and the mean and standard deviation of the speed and direction of the optical flow are also calculated to make a judgment.
[0055] The Gaussian attention map is a weighted map generated by a Gaussian distribution function. The characteristic of the Gaussian distribution is that its weight is the highest at the center and gradually decreases as the distance from the center increases. The basic idea of using the Gaussian attention map to weight image features is to calculate a weighting coefficient based on the feature importance of each position in the image, and then apply this weighting coefficient to the corresponding image feature. In this way, the model can focus on the areas in the image that are most valuable to the task, while weakening unimportant areas, thereby improving performance. Using the speed and direction of the optical flow trajectory statistics to generate the Gaussian center and radius to obtain a suitable Gaussian distribution to generate an attention map to weight image features so that the motion trend learned by the optical flow can be reflected in the attention of the foreground image features in different areas of the picture.
[0056] In the specific implementation, there are two situations: the current frame has a target and the current frame target disappears: In the case of target disappearance, it means that the features of the current frame usually do not need to pay attention to fine details. Through the Gaussian attention map, a module similar to low-pass filtering is designed to reduce the complexity of local textures and make the feature map smoother. Through uniform weighting or center attenuation of the Gaussian map, the feature map of the background area can show continuous smooth characteristics and avoid excessive detail processing; at the same time, the feature map after Gaussian map weighting will become simpler and more regular, which will help reduce the computational complexity of the subsequent network.
[0057] For the case where the target still exists, first calculate the geometric center according to the tracking results of the key nodes of the current frame to obtain the center of gravity of the key points, that is, the center point of the Gaussian distribution. In fact, if the tracking results of the key points are relatively concentrated, the center point will be close to the actual position of the target; if the key points are scattered, the center point tends to the geometric center of the overall position. The standard deviation of the Gaussian distribution determines the smoothing radius of the Gaussian distribution. Through the changes in the speed and direction of the key points, the annotation difference of the Gaussian distribution can be dynamically adjusted to adapt to the situation where the target is concentrated or dispersed. Specifically, the smaller the change in the speed and direction of the optical flow vector, the more concentrated the key point movement is, and the Gaussian distribution should be more focused; when it is larger, the Gaussian distribution should be smoother. Multiply the generated Gaussian attention map with the original feature map pixel by pixel to enhance the area of interest and weaken the background.
[0058] Figure 4 This is a schematic diagram of the optical flow memory assistance process provided by the present invention. Figure 4 The specific implementation process is described below: First, the extraction of the optical flow information of the previous frame occurs at the end of the detection of the previous frame. The decoding result of the decoder is analyzed. If there is a target, the optical flow key points are extracted. If the detection result indicates that there is no target, the optical flow key points are not extracted.
[0059] Before the current frame enters the encoder to extract features, the storage part of the optical flow memory auxiliary module is accessed to confirm whether there are key points extracted from the previous frame. If so, the LK optical flow method is used to track the key nodes of the previous frame in the current frame.
[0060] Match each set of tracking points of the optical flow key points of the previous frame with the results of optical flow tracking of the current frame and obtain a set of optical flow vectors. Calculate the speed (i.e., the Euclidean distance between two points) and direction (i.e., the angle) of these optical flow vectors. Set the clockwise top as 0° and the direction from the point of the previous frame to the point of the current frame. Count the standard deviation and mean of all speeds and directions.
[0061]
[0062] is each value in the data, is the mean of the data,n is the number of data. The data here can be speed and angle values. For the mean and variance of speed and direction, first set a threshold to distinguish the tracking situation of the current frame, and analyze whether the current frame has a target based on the tracking situation. Through experimental analysis, when the speed mean is greater than half of the smaller of the length and width of the image and the angle mean is greater than 180°, it is considered that there may be no target, and the current frame is likely to be an irrelevant frame. It is necessary to further analyze its standard deviation to estimate the target situation of the current frame. Specifically, if the value of the annotation difference is less than 0.3 of the mean, it is more likely to be considered that the current frame has a target, otherwise it is considered that the current frame has no target.
[0063] According to the above analysis of the existence of the target in the current frame, a Gaussian attention map is generated for two cases: judging whether there is no target in the current frame or whether there is a target in the current frame.
[0064] If there is no target in the current frame, the center position of the Gaussian function is considered to be the geometric center of the image of the current frame. The standard deviation of the Gaussian distribution selects a larger value to ensure a smoother weight distribution. The two-dimensional weight value of the Gaussian of the background image without a target is calculated as follows:
[0065] in is the position of the target in the image, For Gaussian distributions where there is no target in the current frame, the center is replaced by the geometric center. Here we take 200 to give a larger standard deviation to smooth the background image details without the target.
[0066] If the current frame is considered to have a target, the calculated mean and standard deviation are used to generate the comprehensive mean and standard deviation values to generate the parameters of the Gaussian attention map: First, the center point of the Gaussian distribution should be the cluster center of the current position of all optical flow tracking, that is, the geometric center point of all points:
[0067] in, is the weight value of each coordinate position, which is calculated by the speed and direction change and the mean change of the optical flow vector. i The weight value of an optical flow vector is calculated as follows:
[0068] in and are the mean values of velocity and angle changes respectively.
[0069] The standard deviation of the Gaussian distribution is obtained from the standard deviation of speed and direction. In order to comprehensively consider the changes in speed and direction, it is calculated by the following formula:
[0070] in is the reference value of the smoothing radius, which is 50 in the experiment. and is the standard deviation of velocity and angle change, respectively expressed as max( , ) and 180° for normalization, max( , ) where w and h are the image size width and height.
[0071] The obtained Gaussian image feature map of optical flow attention G(x, y) and the original image I(x, y) Multiply element-wise to weight the original image using the optical flow information:
[0072] Figure 5 Schematic diagram of the present invention using optical flow vectors for image enhancement, such as Figure 5 As shown in the figure, the image after Gaussian attention weighting can pay more attention to the image area near the target when there is a target, and can smooth the image features to reduce noise or irrelevant information interference and improve detection efficiency when there is no target.
[0073] The optical flow memory-based method is only used when the target is detected in the previous frame, in order to further provide better features of the current frame for subsequent memory interaction.
[0074] The following table is a comparison of the experimental results before and after adding each module: Experimental results comparison table
[0075] in: Flow-A means using optical flow memory to assist enhancement method.
[0076] Mem-A means building a memory library and enhancing it using the memory attention mechanism.
[0077] Mem-UP indicates that the memory bank is updated using the updating method designed in the present invention, otherwise the memory bank is updated using the FIFO first-in-first-out method.
[0078] P, R, and Error represent the accuracy, precision, and false positive rate of detection, respectively.
[0079] The present invention also provides a method for detecting intestinal video targets based on a memory bank, using any of the intestinal video target detection systems described above, the method comprising: Use the optical flow memory auxiliary module to enhance the current image frame; Using the memory storage module to perform feature enhancement on the current image frame using the memory frames in the memory bank during the target detection phase of the current image frame; the memory bank stores image features of the memory frames adjacent to the current image frame; The memory update module is used to perform similarity and importance analysis on the image features of the current image frame and the image features of the memory frames in the memory bank, and decides whether to update the memory bank based on the analysis results.
[0080] The intestinal video target detection method based on a memory library provided by the present invention can also apply the intestinal video target detection system based on a memory library described in any of the above embodiments, which will not be repeated here.
[0081] The intestinal video target detection method and system based on memory library provided by the present invention have the following beneficial effects compared with the prior art: (1) The present invention designs a complete set of colonoscopy video detection processes, and provides a complete system for constructing a memory library from feature encoding in the memory library, updating the memory library, interacting with the memory library and the current frame, and performing optical flow motion statistical analysis, thereby improving the detection efficiency and accuracy of colonoscopy targets.
[0082] (2) The present invention constructs a memory library that can well store the temporal and spatial information of image frames at past time points. When detecting the next frame, the detection results of the previous frames are used to perform cross-attention calculations on the features of the current frame and the features of the image frames already in the memory library, thereby learning the temporal information of the video sequence.
[0083] (3) The present invention fully considers the problem that the memory library needs to be representative. Therefore, when updating the memory library, the traditional first-in-first-out principle is not adopted. Instead, a set of evaluation criteria for memory library updates is specially designed to evaluate the importance of existing image frames in the memory library and the specificity of the current truth, so as to determine whether the current frame needs to enter the memory library and whether the existing image frames in the memory library need to be removed.
[0084] (4) The present invention takes into account that videos are dynamic and that targets have a certain motion trend as video image frames progress. Optical flow memory analysis is introduced to assist in modeling the target's motion as the video progresses. By statistically analyzing the speed and direction of the optical flow vector, the motion of the target is learned and an attention map is generated through the optical flow analysis results to enhance image features.
[0085] Figure 6 is a schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 6As shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630 and a communication bus 640, wherein the processor 610, the communication interface 620 and the memory 630 communicate with each other through the communication bus 640. The processor 610 may call the logic instructions in the memory 630 to execute the intestinal video target detection method based on the memory library.
[0086] On the other hand, the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the memory-based intestinal video target detection method provided in the above-mentioned embodiments.
[0087] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the memory-based intestinal video target detection method provided in the above-mentioned embodiments. Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0088] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A memory-based intestinal video target detection system, characterized in that: include: Memory storage module, memory update module and optical flow memory auxiliary module; The memory storage module is used to perform feature enhancement on the current image frame using the memory frames in the memory bank during the target detection phase of the current image frame; The image features of the multiple memory frames stored in the memory bank are image features of the multiple image frames adjacent to the current image frame; The memory updating module is used to perform feature similarity and importance analysis on the image features of the current image frame and the image features of the memory frames in the memory bank, and decide whether to update the memory bank according to the analysis results; The optical flow memory auxiliary module is used to use the optical flow method to track the changes of optical flow key points in the video sequence to generate an attention feature map, and perform image enhancement on the current image frame.
2. The intestinal video target detection system based on memory library according to claim 1 is characterized in that: The memory storage module is specifically used for: The image features of a fixed number of memory frames adjacent to the current image frame to be detected are stored in the memory bank, and during the detection process of the current image frame, the attention mechanism is used to enhance the features of the current image frame based on the image features of the memory frames in the memory bank.
3. The intestinal video target detection system based on memory library according to claim 2 is characterized in that: The memory storage module uses an attention mechanism to enhance the features of the current image frame based on the image features of the memory frame in the memory bank, including: Generate a query based on the image features of the current image frame, and generate a corresponding key-value pair based on the image features of the memory frame in the memory bank; Perform cross-attention calculation on the query of the current image frame and the key value of each memory frame in the memory library to obtain the attention score corresponding to each memory frame; Based on the attention score, the image features of the memory frame in the memory bank are used to enhance the features of the current image frame to generate enhanced features.
4. The intestinal video target detection system based on memory library according to claim 1, characterized in that: The memory update module is specifically used for: Calculate the cosine similarity based on the image features of the current image frame and the memory frame in the memory library; Get an affinity matrix based on all cosine similarities; The affinity matrix is analyzed to update the memory bank.
5. The intestinal video target detection system based on memory library according to claim 4 is characterized in that: The memory update module analyzes the affinity matrix to update the memory bank, including: Determine the maximum value in the affinity matrix; When the maximum value is greater than or equal to the preset threshold, the image features of the current image frame and the memory frame corresponding to the maximum value are used to perform feature fusion to generate enhanced image features of the memory frame to update the memory library; When the maximum value is less than a preset threshold, comprehensively evaluate the importance of each memory frame in the memory library to the current image frame to obtain a comprehensive evaluation score; The image features of the memory frame with the lowest comprehensive evaluation score are replaced with the image features of the current image frame to update the memory library.
6. The intestinal video target detection system based on memory library according to claim 5, characterized in that: The memory updating module comprehensively evaluates the importance of each memory frame in the memory library to the current image frame and obtains a comprehensive evaluation score, including: Determine the comprehensive evaluation score based on the contribution of the memory frame to the current frame detection, the existence duration of the memory frame, and whether there is a target mutation in the memory frame; Among them, the contribution value of the memory frame to the detection of the current image frame is determined according to the attention score between the memory frame and the current image frame.
7. The intestinal video target detection system based on memory library according to claim 5, characterized in that: The optical flow memory auxiliary module is specifically used for: At the end of the target detection in the previous frame, determine whether there is a target. If there is a target, extract the optical flow key points, then perform optical flow tracking on the optical flow key points to obtain the optical flow vector; Analyze the speed and angle of the optical flow vector to generate an attention feature map; Use the attention feature map to perform image enhancement on the current image frame.
8. The intestinal video target detection system based on memory library according to claim 5, characterized in that: The video sequence is a colonoscopy video sequence.
9. A method for detecting intestinal video targets based on a memory library, using the intestinal video target detection system as described in any one of claims 1 to 8, characterized in that: The method comprises: Use the optical flow memory auxiliary module to enhance the current image frame; Using the memory storage module to perform feature enhancement on the current image frame using the memory frames in the memory bank during the target detection phase of the current image frame; the memory bank stores image features of the memory frames adjacent to the current image frame; The memory update module is used to perform similarity and importance analysis on the image features of the current image frame and the image features of the memory frames in the memory bank, and decides whether to update the memory bank based on the analysis results.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the intestinal video target detection method based on the memory library as described in claim 9 are implemented.
Citation Information
Patent Citations
Multi-mode two-stage unsupervised video anomaly detection method
CN114332053A
Lightweight video object segmentation method based on big data memory storage
CN114882076A
Real-time motion detection method based on multi-scale feature fusion attention
CN115131710A
Video memorability prediction method and device, equipment and storage medium
CN115205745A
Target tracking method based on space-time interaction attention mechanism
CN116563355A