Fundus blood vessel medication region classification method based on time sequence fusion network

Through a time-series fusion network-based method, combined with the timing module and the target detection network, the dynamic change classification problem of drug use in FFA images is solved, and more efficient and accurate analysis of fundus blood vessel drug use is achieved.

CN119942190AActive Publication Date: 2025-05-06HEFEI KERUIKE PHARMACEUTICAL TECHNOLOGY CO LTD
View PDF 14 Cites 0 Cited by

Patent Information

Application Number
CN202510002879.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-05-06
Estimated Expiration
2045-01-02

AI Technical Summary

Technical Problem

The existing FFA image analysis methods rely on expert experience, are time-consuming and subjective, making it difficult to accurately classify the dynamic changes in the area of vascular medicine for the fundus.

Method used

Using a time-series fusion network-based method, combined with a timing module and an object detection network, the timing information in the FFA image sequence is captured to assist in positioning and classifying drug use areas by training a deep learning model.

Benefits of technology

It improves the accuracy and consistency of regional classification of fundus blood vessel medication, reduces labor costs, and reduces the impact of subjective judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942190A_ABST
    Figure CN119942190A_ABST
Patent Text Reader

Abstract

The invention discloses a fundus blood vessel medication region classification method based on a time sequence fusion network, and relates to the technical field of medical image classification. The method comprises the following steps of: 1, acquiring an image sequence and a corresponding label, converting image and label information into a format suitable for neural network learning by using an image labeling tool, and then drawing a bounding box of a medication area on the last image and labeling a corresponding category; according to the fundus blood vessel medication area classification method based on the time sequence fusion network, the characteristics of FFA image medication area classification are considered, a time sequence fusion target detection network is provided, a group of FFA images are input, the medication area can be positioned, and time sequence information of the medication area can be extracted to provide a basis for classification. The defect that a traditional target detection network cannot process sequence images is overcome, the observation process of dynamic changes of a drug action area in the manual classification process is simulated by introducing the time sequence information extraction module, and the classification accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image classification, and in particular to a method for classifying fundus vascular medication areas based on a temporal fusion network. Background Art

[0002] Fluorescein Fundus Angiography (FFA) is an imaging technique used to examine the retinal vascular system at the back of the eye, and is widely used to diagnose and monitor various fundus diseases. It is a non-invasive examination that mainly injects fluorescein dye and then takes dynamic images of fundus blood vessels to help doctors understand the status of retinal and choroidal blood vessels.

[0003] FFA is done by intravenously injecting sodium fluorescein (a yellow fluorescent dye), which flows into the blood vessels of the eye. Fluorescein emits green fluorescence when excited by blue light of a specific wavelength. Using special filters and camera equipment, doctors can capture fluorescent images of the fundus and track the flow of the dye in the retinal and choroidal blood vessels. This process is usually divided into several stages to record the entire process of the dye from entering the eye to its exit.

[0004] Through FFA technology, experts can observe in detail the action process of drugs in the retinal and choroidal blood vessels. On this basis, they can analyze the status of the medication area such as brightness and leakage, evaluate the efficacy of the drug and potential side effects, and optimize the drug and its usage method based on the effects.

[0005] However, this process requires detailed analysis of FFA images, which usually takes a lot of time and is highly dependent on the experience and judgment of experts. Experts must carefully examine each image to capture subtle changes in retinal and choroidal blood vessels, which is crucial to judging the efficacy of drugs, but also increases the complexity and workload of diagnosis. In addition, there may be differences in interpretation between experts, and this subjectivity also poses challenges to subsequent analysis.

[0006] In recent years, artificial intelligence technology, especially deep learning methods, has made remarkable progress in the field of medical image processing. Deep learning methods have been widely used in the analysis and processing of medical images such as MRI, CT and endoscopy, significantly improving the accuracy and efficiency of medical diagnosis. In MRI image processing, deep learning technology is used in many aspects such as image reconstruction, segmentation and disease prediction. By training deep learning networks, researchers can automatically segment structures such as brain tumors and spinal diseases, helping doctors to obtain information about key lesions more quickly. In CT image analysis, deep learning technology is widely used in tasks such as tumor detection, lung disease identification and fracture diagnosis. By training deep learning models, lesions in CT images, such as lung nodules or cerebral hemorrhage, can be automatically identified and classified, helping radiologists to quickly locate lesions in massive images and improve diagnostic efficiency. In endoscopic image processing, deep learning can be applied to image enhancement tasks to improve endoscopic images that are too dark or of poor quality. Through enhancement algorithms, doctors can observe lesions more clearly, thereby reducing the risk of misdiagnosis due to image quality problems. In addition, the researchers also used a neural network based on segmentation tasks to automatically identify and segment polyps and other lesions in endoscopic images. In short, the application of deep learning technology in medical image processing has significantly reduced the workload of doctors in the tedious image review process and improved the efficiency and accuracy of diagnosis.

[0007] Therefore, the use of deep learning technology in FFA image analysis has far-reaching significance. On the one hand, it can save a lot of manpower and material costs. On the other hand, it can help doctors improve the accuracy of analysis, reduce the impact of subjective judgment, and ensure the consistency of analysis results.

[0008] In the task of classifying the medication area of ​​FFA images, there are often multiple key areas that need to be observed. The morphology, structure, and development process of these areas may be significantly different, so uniform classification of the entire image may lead to the omission of important details. Therefore, in order to improve the accuracy of classification and the effectiveness of clinical applications, the ideal solution is to classify each region of interest separately. This method can not only more accurately identify the type of lesion in each area, but also provide more detailed information about the condition.

[0009] To achieve this goal, deep learning models such as object detection networks have become effective tools. Such networks can not only automatically identify different target areas in an image, but also provide accurate location coordinates for each area, thereby achieving more accurate analysis. In the field of deep learning, there are two main types of object detection models, namely, one-stage models represented by YOLO and two-stage models represented by R-CNN, Fast R-CNN, and Faster RCNN. The one-stage model regards the object detection problem as a regression task, directly predicting the category and bounding box on the entire image without generating candidate regions. The two-stage model adopts the strategy of generating candidate regions first, and then performing accurate classification and bounding box regression. The advantage of the one-stage model is that it is fast and suitable for real-time detection tasks, but its accuracy is slightly lower than that of the two-stage model, especially when dealing with complex scenes and small targets. The two-stage model usually performs better in detection accuracy and is suitable for tasks that require high accuracy. FFA images are very complex, mainly reflected in their rich vascular networks and diverse morphological features of the medication area. The fundus contains a large number of small and dense blood vessels, and the morphology, direction and distribution of the blood vessels are very complex. In addition, FFA images contain not only normal blood vessels, but also lesions with subtle differences in morphology. Therefore, the two-stage object detection model is more suitable for FFA images.

[0010] Object detection models usually process single static images, while the analysis of FFA images not only involves information such as brightness and shape of the target area in the late image, but also requires observing its dynamic changes at different times. Therefore, relying solely on traditional object detection models may ignore important information in the time dimension. To solve this problem, a time series processing module must be introduced into the model. Common time series processing modules include recurrent neural networks and their variants, long short-term memory networks, and gated recurrent units, which can effectively capture the time dependencies in the data. Summary of the invention

[0011] 1. Technical issues to be resolved

[0012] In view of the shortcomings of the prior art, the present invention provides a method for classifying fundus vascular medication areas based on a temporal fusion network, which solves the problems mentioned in the above background technology.

[0013] (II) Technical solution

[0014] To achieve the above objectives, the present invention is implemented through the following technical solutions: a method for classifying fundus vascular medication areas based on a temporal fusion network, specifically comprising the following steps:

[0015] Step 1: After obtaining the image sequence and corresponding labels, use the image annotation tool to convert the image and label information into a format suitable for neural network learning, and then draw the bounding box of the medication area on the last image and mark the corresponding category;

[0016] Step 2: Combined with the time series module, it is used to extract the time series information in a set of FFA images, and the auxiliary target detection network uses the time series information to classify more accurately on the basis of locating the medication area. Among them, combined with the time series module, multiple medication areas in the FFA image can be classified;

[0017] Step 3: After the network is built, start training. Input a set of images into the built network to generate predicted bounding boxes, categories, and confidences. Then calculate the coordinate loss between the predicted bounding boxes and the true bounding boxes, as well as the classification loss between the predicted categories and confidences and the true categories.

[0018] Step 4: The network processes the input image and generates the corresponding predicted bounding box, category label, and confidence level. During the test, the predicted result is compared with the true label, and the average accuracy, precision, and recall rate indicators are calculated to evaluate the model performance.

[0019] Preferably, in order to reduce the input amount in step 1, a portion of images needs to be selected in the early, middle and late stages to form an image sequence, wherein the division of the early, middle and late stages is defined by professionals.

[0020] Preferably, when using time series information to more accurately classify in step 2, dynamic changes are an important basis for classifying medication areas.

[0021] Preferably, when the timing module is combined in step 2, dynamic changes in a set of continuous images can be captured.

[0022] Preferably, after calculating the coordinate loss of the predicted bounding box and the true bounding box, as well as the classification loss of the predicted category and confidence and the true category in step three, the gradient information of the loss function is propagated to each layer of the network through the back propagation algorithm, and the network parameters are optimized and updated in combination with the gradient descent method to gradually reduce the loss value.

[0023] Preferably, the gradient information of the loss function is propagated to each layer of the network through the back-propagation algorithm, and the network parameters are optimized and updated in combination with the gradient descent method to gradually reduce the loss value. This training process needs to be repeated until the loss value converges.

[0024] Preferably, when the network processes the input image in step 4, a set of test image sequences that have not participated in the training is prepared, and these images are input into the trained network.

[0025] Preferably, during the test in step 4, the network parameters will not be updated and the loss value will not be calculated.

[0026] (III) Beneficial effects

[0027] The present invention provides a method for classifying fundus vascular medication areas based on a temporal fusion network. Compared with the prior art, it has the following beneficial effects:

[0028] The method for classifying fundus vascular medication areas based on a time series fusion network takes into account the characteristics of FFA image medication area classification and proposes a time series fusion target detection network. When a set of FFA images is input, it can not only locate the medication area, but also extract the time series information of the medication area to provide a basis for classification. This not only makes up for the deficiency that the traditional target detection network cannot process sequence images, but also simulates the observation process of the dynamic changes of the drug action area in the manual classification process by introducing the time series information extraction module, which can improve the classification accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 The network structure diagram of the medication area classification of fluorescein fundus angiography images based on the temporal fusion target detection network;

[0030] Figure 2 The structure diagram of the regional pooling module. DETAILED DESCRIPTION

[0031] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0032] See also Figure 1-Figure 2 , the embodiment of the present invention provides a technical solution:

[0033] The present invention proposes a temporal fusion target detection network for the classification of medication areas in FFA images. Figure 1As shown. First, a set of FFA images is obtained, a total of 10 images, including 5 in the early and middle stages and 5 in the late stages. These 10 images are input into the backbone network to obtain image features, and then the image features of the last image are input into the region proposal network to obtain candidate boxes. The region pooling module maps the candidate boxes to the image features, and after dimensionality reduction processing, the regional features of each image are obtained. On the one hand, the regional features of the last image are sent to the fully connected layer to generate the coordinates of the predicted bounding box, and on the other hand, the regional features of other images are sent to the time series fusion module to obtain the time series features of the group of images. Next, the time series features are fused with the image features of the last image to obtain the fused time series features. The fused time series features are sent to the fully connected layer to generate the category scores. Post-processing will remove redundant and overlapping predicted bounding boxes to facilitate the calculation of performance indicators and visualization.

[0034] Select test platform

[0035] The temporal fusion target detection network of the present invention is implemented using PyTorch, the most popular open source deep learning framework, and is trained and tested on a server equipped with 8 NVIDIA GeForce RTX 3090 GPUs and an Intel Xeon Silver 4214R CPU.

[0036] The specific technical solution includes the following steps:

[0037] 1. Dataset acquisition

[0038] The data used in the present invention is a set of image sequences taken during the early, middle and late stages of fluorescein fundus angiography. In order to reduce the amount of input, a portion of images need to be selected in the early, middle and late stages to form an image sequence. For each medication area, it is necessary to be scored by a professional. After obtaining the image sequence and the corresponding labels, an image annotation tool (such as LabelImg) is used to convert the image and label information into a format suitable for neural network learning. It is necessary to draw the boundary box of the medication area on the last image and mark the corresponding category.

[0039] The data in the training and testing process of the present invention are derived from 14 samples. Each sample was subjected to fluorescein fundus angiography at two weeks, three weeks and four weeks after medication, and FFA images of the left and right eyes were taken continuously. Subsequently, the expert classified each medication area based on his clinical experience. There were 8 medication areas on each image, with a total of four categories 1, 2, 3 and 4. Then, 5 early and middle images and 5 late images were selected from each group of image sequences. After obtaining the image sequence and category labels, the open source image annotation tool LabelImg was used to mark the medication area of ​​the last image of each group with a bounding box and a corresponding category, and the annotation file was saved. Finally, a group of image sequences, a total of 840 images, 84 annotation files, and annotations were obtained, which were used as training and test data for the network.

[0040] 2. Model design

[0041] The present invention proposes a time-series fusion target detection network that can classify multiple medication areas in FFA images. Traditional target detection methods can locate and classify multiple targets in a single image, but cannot capture dynamic changes in a set of continuous images, and dynamic changes are an important basis for classifying medication areas. Therefore, the present invention proposes an improved target detection network that combines a time-series module to extract time-series information from a set of FFA images, and assists the target detection network in locating medication areas and using time-series information to classify more accurately.

[0042] The present invention proposes a temporal fusion target detection network for the classification task of medication areas in fluorescein fundus angiography images. The design motivation of this network is as follows: 1) In FFA images, the classification of medication areas depends on the observation of changes in target areas in sequence images. However, since these areas are usually small and some morphological changes are not obvious enough, and the vascular structure in FFA images is complex and numerous, long-term observation of these images is prone to visual fatigue. In addition, this process needs to be analyzed by experts with rich experience, and fatigue and experience differences between experts may lead to inconsistent classification results. To this end, building a target detection network dedicated to FFA images can not only utilize the powerful ability of neural networks to accurately capture subtle changes in images and reduce labor costs, but also improve the consistency of classification results. 2) Traditional target detection networks can realize the positioning and classification of multiple targets in static images. For example, the two-stage target detection network Faster R-CNN first generates a batch of candidate bounding boxes and then extracts feature information in these bounding boxes. Based on the extracted features, on the one hand, it is used to output the category and probability of the target in the bounding box, and on the other hand, it is used to adjust the coordinates of the bounding box to improve the positioning accuracy of the target. In the task of classifying the medication area of ​​FFA images, it is necessary to locate the medication area and to classify the medication area. Naturally, the target detection model will be considered to be migrated to this task. However, the difference is that the classification of medication areas in FFA images not only needs to consider the brightness, morphology and other information of the medication area on a single FFA image, but also needs to consider the diffusion process of fluorescein in the medication area, which requires multiple images rather than just one image to be examined. For the network, it is necessary to obtain additional time series features. Therefore, the time series fusion target detection network designed by the present invention fuses the time series information into the target detection network to classify the medication areas of FFA images.

[0043] Figure 1 The overall architecture of the target detection network designed by the present invention that integrates temporal information is described. The network input is I1 to I 10 There are 10 images in total, and the output is a set of bounding box coordinates and their corresponding categories for the last image. The following describes the composition of each part of the network.

[0044] Backbone network: Its main function is to extract image features. The backbone network of this network uses ResNet50 and loads its pre-trained weights. This not only accelerates the convergence of the model, but also reduces the risk of overfitting due to limited training data. ResNet50 is a multi-layer convolutional neural network structure. The higher the layer, the more abstract the features are, and the more conducive it is to identifying large-scale targets in the image, and vice versa. The area of ​​the medication area in the FFA image is relatively small. Therefore, this network does not use the top-level features of ResNet50, but uses its features in the third stage (ResNet50 has a total of 4 stages of features).

[0045] Region Proposal Network: Its main function is to generate corresponding candidate boxes based on image features.

[0046] Region pooling module: This module maps the candidate boxes to the image features to obtain the region features. The sizes of the candidate boxes are different, so the scales of the mapped features are also different. Therefore, the region pooling module uses the region pooling layer to map these features of different sizes to a uniform scale, then expands these features and sends them to the continuous fully connected layer to reduce the dimensionality of these features to reduce the number of parameters.

[0047] Time series fusion module: It is used to extract the time series features in the sequence images and add them to the last image in the image to obtain the fusion features as the classification features.

[0048] Fully connected layer 1: A fully connected layer used to regress the original category scores corresponding to the predicted bounding boxes.

[0049] FC2: A fully connected layer that adjusts the coordinates of the predicted bounding box.

[0050] Specifically, for a set of FFA images Where i∈{1,2,...,9,10}, h and w are the width and height of the image. i Input into the backbone network in sequence to obtain their respective image features Where c is the number of feature channels. When using the third stage features of ResNet50, c is 1024. d1 and d2 are the width and height of the feature, respectively. For the feature F of the 10th image 10 , and send it to the region proposal network to generate the candidate box R n , where n is the number of candidate boxes. Next, the candidate box R n And the image feature F i It is input into the regional pooling module and first performs the operation as in formula (1).

[0051] F i R =f roi(R n ,F i ) (1)

[0052] Among them, f roi (·) indicates that the region pooling layer maps candidate boxes of different sizes to image features and resizes them to the same size. Where d is the uniform regional feature width. Next, the operation in formula (2) is performed.

[0053] F i P =W b (W a (flatten(F i R ))) (2)

[0054] Among them, flatten(·) represents the expansion operation, which expands the feature of shape (n,c,d,d) to (n,h0), where h0 = c×d×d. are the parameters of fully connected layer a and fully connected layer b, respectively. h1 and h2 represent the number of neurons in fully connected layer a and fully connected layer b, respectively. Fully connected layer a and fully connected layer b map h0 dimension to h1 and h2 respectively, and finally get

[0055] Regional features of the 10th image It is input into the fully connected layer 2 and used to adjust the coordinates of the predicted bounding box, as shown in formula (3).

[0056] boxes=W2F i P (3)

[0057] in, is the parameter of the fully connected layer 2, 20 represents the 4 position offsets corresponding to 5 categories (4 categories and 1 background class), so

[0058] At the same time F1 P to The features are first stacked and then passed through the gated recurrent unit to generate the temporal features F g , as shown in formula (4).

[0059]

[0060] Among them, stack(·) represents the stacking operation, f gru (·) captures the temporal features between features through the gated recurrent unit, and the feature dimension of its hidden state is set to h2, then Get the time series feature Fg After that, Add to obtain fusion features F t It is sent to the fully connected layer 1 to generate the original category scores, as shown in formula (5).

[0061] classes=W1F t (5)

[0062] in, is the parameter of the fully connected layer 1, 5 represents 5 categories (4 categories and one background category), so The complete output of the network is The predicted bounding boxes have a one-to-one correspondence with the original category scores.

[0063] The post-processing operation will remove the background prediction box, remove the low-confidence prediction box, perform non-maximum suppression to eliminate redundant boxes of the unified target, and so on, so as to obtain clear and accurate detection results.

[0064] 3. Network training

[0065] The temporal fusion target detection network proposed in the present invention is implemented through the widely used open source deep learning framework Pytorch. After the network is built, training begins. First, a set of images are input into the built network to generate predicted bounding boxes, categories, and confidences. Next, the coordinate loss between the predicted bounding box and the true bounding box, as well as the classification loss between the predicted category and confidence and the true category are calculated. Subsequently, the gradient information of the loss function is propagated to each layer of the network through the back propagation algorithm, and the network parameters are optimized and updated in combination with the gradient descent method to gradually reduce the loss value. Repeat this training process until the loss value converges.

[0066] The target detection network that integrates temporal information proposed in the present invention is built through the open source deep learning framework Pytorch. The data set is divided into a training set and a test set in a ratio of 8:2. First, after obtaining the training set, random horizontal flipping is performed for data enhancement. During this process, the images of each group are flipped at the same time, and the coordinates of the candidate boxes in the corresponding annotation files are also flipped. The enhanced data set is sent to the network in batches to obtain the predicted bounding box coordinates and the corresponding categories and confidence levels. Calculate the loss, perform back-propagation training, and then adjust the parameters by the gradient descent method to reduce the loss value until convergence.

[0067] Specifically, the loss of category prediction is calculated using binary cross entropy, and the loss of bounding box coordinate prediction is calculated using smooth L1. During network training, the stochastic gradient descent optimizer is used to adjust the parameters, with the learning rate set to 0.005, the momentum set to 0.9, and the weight decay parameter set to 1×10-4. The batch size is set to 8. The model is trained in an end-to-end manner, with 100 training cycles.

[0068] 4. Network testing

[0069] After completing the network training and obtaining the optimized model, the network needs to be tested. The purpose of the test is to evaluate the performance of the model and verify its generalization ability on unseen data. First, prepare a set of test image sequences that have not participated in the training and input these images into the trained network. The network will process the input image and generate the corresponding predicted bounding box, category label, and confidence. During the test, the network parameters will not be updated, and the loss value will not be calculated. Instead, the predicted results will be compared with the true labels, and the average accuracy, precision, recall and other indicators will be calculated to evaluate the model performance.

[0070] The present invention uses mAP50 as a performance evaluation indicator, which is a commonly used evaluation indicator in object detection tasks. mAP50 refers to the average precision (AP) of the model when the intersection-over-union ratio threshold of the predicted box and the true box is 50%, and the average value is taken for all categories. The average precision AP mainly evaluates the performance of the precision and recall rate of the model under different thresholds.

[0071] The intersection-over-union ratio refers to the degree of overlap between the predicted bounding box and the true bounding box, as shown in formula (6).

[0072]

[0073] Among them, IOU means intersection over union, B p and B gt Represents the predicted bounding box and the true bounding box respectively. When the intersection-over-union ratio exceeds the set threshold (the threshold is set to 50% in the mAP50 indicator), the prediction is considered to be correct and is recorded as TP (True Positive). When the intersection-over-union ratio is less than the threshold, it is recorded as FP (False Positive). For the true bounding box that is not detected, it is recorded as FN (False Negative).

[0074] The precision and recall are further calculated using TP, FP, and FN. The calculation method is as shown in formulas (7) and (8).

[0075]

[0076] Where P is the precision rate, also known as the check rate, which calculates the ratio of correctly predicted bounding boxes to all predicted bounding boxes. R is the recall rate, also known as the check rate, which calculates the ratio of correctly predicted bounding boxes to the true bounding boxes.

[0077] Then, the paired P and R values ​​under different confidence thresholds are calculated. These P and R value pairs are sorted from high to low according to confidence, and the AP is calculated using formula (9).

[0078]

[0079] Where n is the serial number of the P, R value pair, in particular, R0 = 0. interp (R n+1 ) is calculated by formula (10).

[0080]

[0081] in, for The corresponding P value.

[0082] After completing the calculation of the AP for each category, the average value is calculated to calculate the mAP index of the model, as shown in formula (11).

[0083]

[0084] Wherein, N is the total number of categories. In the present invention, there are 4 categories, so N=4.

[0085] At the end of each training cycle, the network weights are saved and then used to perform performance tests on the test set. For the test set, the average precision AP of each category is first calculated, and then the average of these average precisions is taken to get mAP50 as the performance indicator of the current cycle. The network weight with the largest mAP50 within 100 training cycles is saved as the final network parameter.

[0086] The above content only describes the implementation form of the present invention in detail in combination with the relevant preferred embodiments, and it cannot be determined that the specific embodiments of the present invention are limited to these descriptions. The protection scope of the present invention is the same technical means with the same performance and use based on the basic researchers in the field of the present invention.

[0087] The embodiments of the present application have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A method for classifying fundus vascular medication areas based on a temporal fusion network, characterized in that: The specific steps include: Step 1: After obtaining the image sequence and corresponding labels, use the image annotation tool to convert the image and label information into a format suitable for neural network learning, and then draw the bounding box of the medication area on the last image and mark the corresponding category; Step 2: Combined with the time series module, it is used to extract the time series information in a set of FFA images, and the auxiliary target detection network uses the time series information to classify more accurately on the basis of locating the medication area. Among them, combined with the time series module, multiple medication areas in the FFA image can be classified; Step 3: After the network is built, start training. Input a set of images into the built network to generate predicted bounding boxes, categories, and confidences. Then calculate the coordinate loss between the predicted bounding boxes and the true bounding boxes, as well as the classification loss between the predicted categories and confidences and the true categories. Step 4: The network processes the input image and generates the corresponding predicted bounding box, category label, and confidence level. During the test, the predicted result is compared with the true label, and the average accuracy, precision, and recall rate indicators are calculated to evaluate the model performance.

2. According to claim 1, a method for classifying fundus vascular medication areas based on a temporal fusion network is characterized by: In order to reduce the input amount in the step 1, a portion of images need to be selected in the early, middle and late stages to form an image sequence, wherein the division of the early, middle and late stages is defined by professionals.

3. According to claim 1, a method for classifying fundus vascular medication areas based on a temporal fusion network is characterized by: When using time series information to more accurately classify in step 2, dynamic changes are an important basis for classifying medication areas.

4. According to claim 1, a method for classifying fundus vascular medication areas based on a temporal fusion network is characterized by: When the timing module is combined in step 2, dynamic changes in a set of continuous images can be captured.

5. The method for classifying fundus vascular medication areas based on a temporal fusion network according to claim 1, characterized in that: After calculating the coordinate loss of the predicted bounding box and the true bounding box, as well as the classification loss of the predicted category and confidence and the true category in step three, the gradient information of the loss function is propagated to each layer of the network through the back propagation algorithm, and the network parameters are optimized and updated in combination with the gradient descent method to gradually reduce the loss value.

6. The method for classifying fundus vascular medication areas based on a temporal fusion network according to claim 5, characterized in that: The back-propagation algorithm is used to propagate the gradient information of the loss function to each layer of the network, and the gradient descent method is used to optimize and update the network parameters to gradually reduce the loss value. This training process needs to be repeated until the loss value converges.

7. The method for classifying fundus vascular medication areas based on a temporal fusion network according to claim 1 is characterized by: When the network processes the input image in step 4, a set of test image sequences that have not participated in the training are prepared, and these images are input into the trained network.

8. The method for classifying fundus vascular medication areas based on a temporal fusion network according to claim 1 is characterized by: During the test in step 4, the network parameters will not be updated and the loss value will not be calculated.

Citation Information

Patent Citations

  • A video semantic segmentation method and device based on prediction for feature propagation

    CN109919044A

  • Power transmission line defect detection method based on hierarchical region feature fusion learning

    CN110335270A

  • Video processing method, mobile terminal and readable storage medium

    CN111523402A

  • Time sequence behavior detection method and device, time sequence behavior response method and device, equipment and medium

    CN112418114A

  • Space domain limited pixel target detection system and method fusing neural network space-time characteristics

    CN114022759A