Measurement method and device of fundus image based on tight frame mark and network training

By constructing a backbone network and a regression network based on tight-frame labels in deep learning, the accurate identification and measurement of the optic cup and optic disc in fundus images were achieved, solving the problems of high cost and low accuracy in existing technologies and improving the accuracy of measurement.

CN115331050BActive Publication Date: 2026-07-31SHENZHEN SIBRIGHT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN SIBRIGHT TECH CO LTD
Filing Date
2021-10-19
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing deep learning-based methods for target recognition and measurement in fundus images require precise pixel-level labeled data, resulting in high consumption of human and material resources, and the boundary recognition is not accurate enough to meet the needs of precise measurement.

Method used

A deep learning method based on tight-frame markers is adopted. By constructing a backbone network, a weakly supervised learning image segmentation network, and a bounding box regression network, the tight-frame markers are used to identify and measure the optic cup and/or optic disc in fundus images. This includes obtaining the class probability of pixels and the offset of the tight-frame markers to achieve accurate measurement.

Benefits of technology

It reduces the need for pixel-level annotation data, improves the boundary recognition accuracy of the optic cup and optic disc, and can accurately measure the size and ratio of the optic cup and optic disc, thus reducing manpower and material costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115331050B_ABST
    Figure CN115331050B_ABST
Patent Text Reader

Abstract

This disclosure describes a measurement method, apparatus, and network training method for fundus images based on tight-boundary labels. The network module includes a backbone network, a segmentation network for image segmentation, and a regression network based on bounding box regression. Network training includes constructing training samples; obtaining predicted segmentation data output by the segmentation network and predicted offset output by the regression network from fundus image data based on the training samples; wherein the backbone network is used to extract feature maps from the image to be trained, the segmentation network takes the feature maps as input to output predicted segmentation data, and the regression network takes the feature maps as input to output predicted offset; determining the training loss of the network module based on the label data corresponding to the training samples, the predicted segmentation data, and the predicted offset; and training the network module based on the training loss to optimize the network module.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of patent application filed on October 19, 2021, with application number 202111216625.8 and invention title "Measuring Method and Measuring Apparatus for Fundus Images Based on Deep Learning with Tight-Frame Markers". Technical Field

[0002] This disclosure generally relates to the field of deep learning-based recognition technology, and specifically to a measurement method, apparatus, and network training method for fundus images based on tight-frame targets. Background Technology

[0003] Fundus images often contain information about various targets. Image processing techniques can be used to identify these targets in fundus images, enabling automatic analysis of the targets within the fundus. For example, the optic cup and / or optic disc can be identified in fundus images, allowing for the measurement of their dimensions to monitor changes.

[0004] In recent years, artificial intelligence technologies, represented by deep learning, have made significant progress, and their applications in target recognition and measurement have attracted increasing attention. Researchers use deep learning techniques to identify or further measure targets in images. Specifically, in some deep learning-based studies, labeled data is often used to train deep learning-based neural networks to identify and segment the optic cup and / or optic disc in fundus images, thereby enabling the measurement of the optic cup and / or optic disc.

[0005] However, the aforementioned target recognition or measurement methods often require precise pixel-level labeled data for training neural networks, and collecting such data is typically very resource-intensive. Furthermore, some target recognition methods, while not based on pixel-level labeled data, merely identify the optic cup and / or optic disc in fundus images. Their boundary identification of the optic cup and / or optic disc is not precise enough, or the accuracy is often low near the boundaries in fundus images, making them unsuitable for scenarios requiring precise measurements. In such cases, the accuracy of measuring the optic cup and / or optic disc in fundus images needs further improvement. Summary of the Invention

[0006] This disclosure is made in view of the above-mentioned situation, and its purpose is to provide a method and apparatus for measuring fundus images based on tight-frame deep learning that can identify the optic cup and / or optic disc and accurately measure the optic cup and / or optic disc.

[0007] Therefore, the first aspect of this disclosure provides a method for measuring fundus images based on tight-frame markers in deep learning. This method utilizes a network module trained on target-based tight-frame markers to identify at least one target in the fundus image, thereby achieving measurement. The at least one target is the optic cup and / or optic disc, and the tight-frame marker is the minimum bounding rectangle of the target. The measurement method includes: acquiring a fundus image; inputting the fundus image into the network module to obtain a first output and a second output. The first output includes the probability that each pixel in the fundus image belongs to the category of the optic cup and / or optic disc, and the second output includes the position of each pixel in the fundus image and the probability of each pixel belonging to the category of the optic cup and / or optic disc. The offset of the target's bounding box in the category is used as the target offset in the second output. The network module includes a backbone network, a segmentation network based on weakly supervised learning for image segmentation, and a regression network based on bounding box regression. The backbone network is used to extract feature maps from the fundus image. The segmentation network takes the feature maps as input to obtain the first output, and the regression network takes the feature maps as input to obtain the second output. The feature maps have the same resolution as the fundus image. The target is identified based on the first output and the second output to obtain the bounding box of the optic cup and / or optic disc in the fundus image, thereby achieving measurement.

[0008] In this disclosure, a network module is constructed comprising a backbone network, a segmentation network for image segmentation based on weakly supervised learning, and a regression network for bounding box regression. The network module is trained based on the target's tight bounding box. The backbone network receives a fundus image and extracts a feature map with the same resolution as the fundus image. The feature map is input into the segmentation network and the regression network respectively to obtain a first output and a second output. Then, based on the first and second outputs, the tight bounding box of the optic cup and / or optic disc in the fundus image is obtained, thereby enabling measurement. In this configuration, the network module trained based on the tight bounding box of the optic cup and / or optic disc can accurately predict the tight bounding box of the optic cup and / or optic disc in the fundus image, thus enabling accurate measurement based on the tight bounding box of the optic cup and / or optic disc.

[0009] Additionally, in the measurement method according to the first aspect of this disclosure, optionally, the optic cup and / or optic disc are measured based on the rim marker of the optic cup and / or the rim marker of the optic disc in the fundus image to obtain the size of the optic cup and / or optic disc. This allows for accurate measurement of the size of the optic cup and / or optic disc.

[0010] Furthermore, in the measurement method according to the first aspect of this disclosure, optionally, the ratio of the optic cup to the optic disc is obtained based on the dimensions of the optic cup and optic disc in the fundus image. In this case, the ratio of the optic cup to the optic disc is obtained based on the rim marker, thereby enabling accurate measurement of the cup-disc ratio.

[0011] Additionally, in the measurement method according to the first aspect of this disclosure, optionally, the network module is trained by: constructing training samples, wherein the fundus image data of the training samples includes multiple images to be trained, the multiple images to be trained include images containing at least one target among the optic cup and optic disc, and the label data of the training samples includes the gold standard of the category to which the target belongs and the gold standard of the tight-frame label of the target, wherein the images to be trained are fundus images to be trained; obtaining, based on the fundus image data of the training samples, predicted segmentation data output by the segmentation network and predicted offset output by the regression network corresponding to the training samples; determining the training loss of the network module based on the label data corresponding to the training samples, the predicted segmentation data, and the predicted offset; and training the network module based on the training loss to optimize the network module. Thus, an optimized network module can be obtained.

[0012] Furthermore, in the measurement method according to the first aspect of this disclosure, optionally, determining the training loss of the network module based on the label data corresponding to the training samples, the predicted segmentation data, and the predicted offset includes: obtaining the segmentation loss of the segmentation network based on the predicted segmentation data and label data corresponding to the training samples; obtaining the regression loss of the regression network based on the predicted offset corresponding to the training samples and the true offset based on the label data, wherein the true offset is the offset of the position of a pixel in the image to be trained from the gold standard of the tight bounding box of the target in the label data; and obtaining the training loss of the network module based on the segmentation loss and the regression loss. In this case, the predicted segmentation data of the segmentation network can be made to approximate the label data through the segmentation loss, and the predicted offset of the regression network can be made to approximate the true offset through the regression loss.

[0013] Furthermore, in the measurement method according to the first aspect of this disclosure, optionally, the target offset is an offset normalized based on the average size of targets of each category. This improves the accuracy of identifying or measuring targets with small size variations.

[0014] Furthermore, in the measurement method according to the first aspect of this disclosure, optionally, the width and height of the bounding boxes of targets in the label data are averaged according to category to obtain average width and average height, and then the average width and average height are averaged to obtain the average size of targets in each category. Thus, the average size of targets can be obtained from training samples.

[0015] Furthermore, in the measurement method according to the first aspect of this disclosure, optionally, multiple instance learning is used to obtain multiple training packets according to categories based on the gold standard of the tight bounding boxes of targets in each training image. The segmentation loss is obtained based on the multiple training packets of each category. The multiple training packets include multiple positive packets and multiple negative packets. A positive packet is defined as all pixels on each of the multiple straight lines connecting the opposite sides of the gold standard of the tight bounding boxes of the targets. The multiple straight lines include at least one set of mutually parallel first parallel lines and mutually parallel second parallel lines perpendicular to each set of first parallel lines. The negative packet is a single pixel in the region outside the gold standard of the tight bounding boxes of all targets in a category. Thus, the segmentation loss can be obtained based on the positive and negative packets obtained through multiple instance learning.

[0016] Furthermore, in the measurement method according to the first aspect of this disclosure, optionally, the angle of the first parallel line is the angle between the extension of the first parallel line and the extension of any non-intersecting side of the gold standard of the target's tight frame, wherein the angle of the first parallel line is greater than -90° and less than 90°. In this case, it is possible to optimize the segmentation network by dividing the positive envelope at different angles. This improves the accuracy of the segmentation network's predicted segmentation data.

[0017] Furthermore, in the measurement method according to the first aspect of this disclosure, optionally, the segmentation loss includes a unary term and a pairwise term, wherein the unary term describes the degree to which each training bag belongs to the gold standard of each category, and the pairwise term describes the degree to which a pixel in the training image belongs to the same category as its neighboring pixels. In this case, the tight bounding box can be constrained by both positive and negative bags simultaneously through the unary loss, and the predicted segmentation result can be smoothed through the pairwise loss.

[0018] Furthermore, in the measurement method according to the first aspect of this disclosure, optionally, pixels that fall within the gold standard tight bounding box of at least one target are selected from the image to be trained as positive samples to optimize the regression network. In this case, optimizing the regression network based on pixels that fall within the true tight bounding box of at least one target can improve the efficiency of regression network optimization.

[0019] Furthermore, in the measurement method according to the first aspect of this disclosure, optionally, pixels falling within the gold standard tight bounding box of at least one target are selected from the training image according to category as positive samples for each category, and the matching tight bounding box corresponding to the positive sample is obtained to filter the positive samples for each category based on the matching tight bounding box. Then, the filtered positive samples for each category are used to optimize the regression network, wherein the matching tight bounding box is the gold standard tight bounding box into which the positive sample falls. Thus, the regression network can be optimized using the positive samples for each category filtered based on the matching tight bounding box.

[0020] Furthermore, in the measurement method according to the first aspect of this disclosure, optionally, let the position of a pixel be represented as (x, y), the bounding box label of a target corresponding to the pixel be represented as b = (xl, yt, xr, yb), and the offset of the bounding box label b of the target relative to the position of the pixel be represented as t = (tl, tt, tr, tb). Then tl, tt, tr, tb satisfy the formula: tl = (x - xl) / S c tt=(y-yt) / S c tr=(xr-x) / S c tb=(yb-y) / S c Where xl,yt represents the position of the top left corner of the target's bounding box, xr,yb represents the position of the bottom right corner of the target's bounding box, and S c This represents the average size of the target in the c-th category. From this, we can obtain the normalized offset.

[0021] Furthermore, in the measurement method according to the first aspect of this disclosure, optionally, the regression network is optimized by filtering pixels from the training image according to category and using the expected cross-union ratio (CUI) corresponding to the pixels of the training image, selecting pixels whose expected CUI is greater than a preset expected CUI. This allows for the acquisition of positive samples that meet the preset expected CUI.

[0022] Furthermore, in the measurement method according to the first aspect of this disclosure, optionally, multiple bounding boxes of different sizes are constructed with the pixels of the image to be trained as the center points. The maximum value among the intersection-union ratios (IURs) of the multiple bounding boxes with the matching tight bounding boxes of the pixel is obtained and used as the expected IUR, wherein the matching tight bounding box is the gold standard for the tight bounding box to which the pixel of the image to be trained falls. Thus, the expected IUR can be obtained.

[0023] Additionally, in the measurement method according to the first aspect of this disclosure, optionally, the expected crossover ratio satisfies the formula: Among them, r1 and r2 are the relative positions of the pixel points of the to-be-trained image in the matching tight bounding box label, 0 < r1, r2 < 1, IoU1(r1, r2) = 4r1r2, IoU2(r1, r2) = 2r1 / (2r1(1 - 2r2) + 1), IoU3(r1, r2) = 2r2 / (2r2(1 - 2r1) + 1), IoU4(r1, r2) = 1 / (4(1 - r1)(1 - r2)). Thus, the desired intersection over union can be obtained.

[0024] In addition, in the measurement method according to the first aspect of the present disclosure, optionally, the regression loss satisfies the formula: Among them, C represents the number of categories, M c represents the number of positive samples of the c-th category, t ic represents the true offset corresponding to the i-th positive sample of the c-th category, v ic represents the predicted offset corresponding to the i-th positive sample of the c-th category, and s(x) represents the sum of the smooth L1 losses of all elements in x. Thus, the regression loss can be obtained.

[0025] In addition, in the measurement method according to the first aspect of the present disclosure, optionally, the identification of the target based on the first output and the second output to obtain the tight bounding box label of the optic cup and / or optic disc in the fundus image so as to achieve measurement is as follows: obtaining the position of the pixel point with the highest probability belonging to each category from the first output as the first position, and obtaining the tight bounding box label of the target of each category based on the position corresponding to the first position in the second output and the target offset of the corresponding category. Thus, the optic cup and / or optic disc can be identified.

[0026] In addition, in the measurement method according to the first aspect of the present disclosure, optionally, the backbone network includes an encoding module and a decoding module. The encoding module is configured to extract image features at different scales, and the decoding module is configured to map the image features extracted at different scales back to the resolution of the fundus image to output the feature map. Thus, a feature map consistent with the resolution of the fundus image can be obtained.

[0027] A second aspect of this disclosure provides a measurement device for fundus images based on tight-frame markers in deep learning. This device utilizes a network module trained on target-based tight-frame markers to identify at least one target in the fundus image, thereby achieving measurement. The at least one target is the optic cup and / or optic disc, and the tight-frame marker is the minimum bounding rectangle of the target. The measurement device includes an acquisition module, a network module, and a recognition module. The acquisition module is configured to acquire fundus images. The network module is configured to receive the fundus images and acquire a first output and a second output based on the fundus images. The first output includes the probability that each pixel in the fundus image belongs to the category of the optic cup and / or optic disc. The second output includes the probability that each pixel in the fundus image belongs to the category of the optic cup and / or optic disc. The method describes the offset of the position of each pixel in the fundus image with the bounding box of each category of target, and uses the offset in the second output as the target offset. The network module includes a backbone network, a segmentation network based on weakly supervised learning for image segmentation, and a regression network based on bounding box regression. The backbone network is used to extract feature maps from the fundus image. The segmentation network takes the feature maps as input to obtain the first output, and the regression network takes the feature maps as input to obtain the second output. The feature maps have the same resolution as the fundus image. The recognition module is configured to recognize the target based on the first output and the second output to obtain the bounding box of the optic cup and / or optic disc in the fundus image for measurement.

[0028] In this disclosure, a network module is constructed comprising a backbone network, a segmentation network for image segmentation based on weakly supervised learning, and a regression network for bounding box regression. The network module is trained based on the target's tight bounding box. The backbone network receives a fundus image and extracts a feature map with the same resolution as the fundus image. The feature map is input into the segmentation network and the regression network respectively to obtain a first output and a second output. Then, based on the first and second outputs, the tight bounding box of the optic cup and / or optic disc in the fundus image is obtained, thereby enabling measurement. In this configuration, the network module trained based on the tight bounding box of the optic cup and / or optic disc can accurately predict the tight bounding box of the optic cup and / or optic disc in the fundus image, thus enabling accurate measurement based on the tight bounding box of the optic cup and / or optic disc.

[0029] According to this disclosure, a method and apparatus for measuring fundus images based on tight-frame deep learning are provided, which can identify the optic cup and / or optic disc and accurately measure the optic cup and / or optic disc. Attached Figure Description

[0030] This disclosure will now be explained in further detail by way of example only with reference to the accompanying drawings, in which:

[0031] Figure 1This is a schematic diagram illustrating an application scenario of the fundus image measurement method based on tight-frame markers, as described in this disclosure.

[0032] Figure 2(a) is a schematic diagram illustrating a fundus image involved in an example of this disclosure.

[0033] Figure 2(b) is a schematic diagram showing the recognition results of the fundus image involved in the example of this disclosure.

[0034] Figure 3 This is a schematic diagram illustrating an example of a network module involved in the examples of this disclosure.

[0035] Figure 4 This is a schematic diagram illustrating another example of a network module involved in the examples of this disclosure.

[0036] Figure 5 This is a flowchart illustrating the training method of the network module involved in the example of this disclosure.

[0037] Figure 6 This is a schematic diagram illustrating the positive package involved in the examples of this disclosure.

[0038] Figure 7 This is a schematic diagram illustrating a border constructed centered on a pixel as described in the examples of this disclosure.

[0039] Figure 8(a) is a flowchart illustrating a method for measuring fundus images based on deep learning with tight-frame labels, as described in this disclosure.

[0040] Figure 8(b) is a flowchart illustrating another example of a method for measuring fundus images based on tight-frame deep learning, as described in this disclosure.

[0041] Figure 9(a) is a block diagram illustrating a measurement apparatus for fundus images based on deep learning of tight-frame labels, as described in this disclosure example.

[0042] Figure 9(b) is a block diagram illustrating another example of a measurement apparatus for fundus images based on deep learning of tight-frame labels, as described in this disclosure.

[0043] Figure 9(c) is a block diagram illustrating another example of a measurement apparatus for fundus images based on deep learning of tight-frame labels, as described in the examples of this disclosure. Detailed Implementation

[0044] The preferred embodiments of this disclosure are described in detail below with reference to the accompanying drawings. In the following description, the same reference numerals are used for the same components, and repeated descriptions are omitted. Furthermore, the drawings are merely schematic diagrams, and the proportions of the components or the shapes of the components may differ from actual figures. It should be noted that the terms "comprising" and "having," and any variations thereof, in this disclosure, do not necessarily limit the process, method, system, product, or apparatus to the explicitly listed steps or units, but may include or have other steps or units not explicitly listed or inherent to these processes, methods, products, or apparatuses. All methods described in this disclosure may be performed in any suitable order unless otherwise indicated herein or clearly contradicted by the context.

[0045] This disclosure relates to a method and apparatus for measuring fundus images based on deep learning with tight-frame markers. This method can identify targets and improve the accuracy of target measurement. For example, it can identify the optic disc or its tight-frame markers in a fundus image, and then measure the size of the optic disc or its tight-frame markers based on these markers. The method for measuring fundus images based on deep learning with tight-frame markers disclosed herein can also be called a recognition method, a tight-frame marker measurement method, a tight-frame marker identification method, an automatic measurement method, an assisted measurement method, etc. The measurement method disclosed herein can be applied to any application scenario requiring accurate measurement of the width and / or height of targets in an image.

[0046] The measurement method disclosed herein utilizes a network module trained on a target-based tight bounding box to identify and measure targets. The tight bounding box can be the smallest bounding rectangle of the target. In this case, the target is in contact with all four sides of the tight bounding box and does not overlap with any area outside the tight bounding box (i.e., the target is tangent to all four sides of the tight bounding box). Thus, the tight bounding box can represent the width and height of the target. Furthermore, training the network module based on the target's tight bounding box reduces the time and labor costs of collecting pixel-level labeled data (also known as tag data), and the network module can accurately identify the target's tight bounding box.

[0047] Figure 1 Figure 2(a) is a schematic diagram illustrating an application scenario of the fundus image measurement method based on tight-frame label deep learning involved in the present disclosure. Figure 2(b) is a schematic diagram illustrating the fundus image involved in the present disclosure.

[0048] In some examples, the measurement methods disclosed herein can be applied to, for example... Figure 1 In the application scenario shown, fundus images of the fundus 51 can be acquired using acquisition device 52 (e.g., a camera) (see [reference]). Figure 1The fundus image is input into the network module 20 to identify the optic cup and / or optic disc in the fundus image and obtain the tight-frame marker B of the optic cup and / or optic disc (see Figure 1). Measurements of the optic cup and / or optic disc can then be performed based on the tight-frame marker B. For example, inputting the fundus image shown in Figure 2(a) into the network module 20 can obtain the recognition result shown in Figure 2(b). The recognition result may include tight-frame markers for both the optic cup and the optic disc, where tight-frame marker B11 is the tight-frame marker for the optic disc and tight-frame marker B12 is the tight-frame marker for the optic cup. In this case, measurements of the optic cup and optic disc can be performed based on the tight-frame markers.

[0049] The network module 20 disclosed herein can be multi-task based. In some examples, network module 20 can be a deep learning-based neural network. In some examples, network module 20 can include two tasks: one task can be a segmentation network 22 (described later) based on weakly supervised learning for image segmentation, and the other task can be a regression network 23 (described later) based on bounding box regression.

[0050] In some examples, segmentation network 22 can segment fundus images to obtain targets (e.g., optic cup and / or optic disc). In some examples, segmentation network 22 can be based on multiple-instance learning (MIL) and used to supervise tight bounding boxes. In some examples, the problem solved by segmentation network 22 can be a multi-label classification problem. In some examples, the fundus image can contain targets of at least one category of interest (which may be simply referred to as a category). For example, the targets in the fundus image can be the optic cup and / or optic disc, that is, targets of category 1 in the fundus image can be identified. Thus, segmentation network 22 is able to identify fundus images containing targets of at least one category of interest. In some examples, the fundus image may also be devoid of any targets.

[0051] In some examples, regression network 23 can be used to predict tight bounding boxes by category. In some examples, regression network 23 can predict tight bounding boxes by predicting the offset of the tight bounding box relative to the position of each pixel in the fundus image.

[0052] In some examples, network module 20 may also include a backbone network 21. Backbone network 21 can be used to extract feature maps from the fundus image (i.e., the original image input to network module 20). In some examples, backbone network 21 can extract high-level features for object representation. In some examples, the resolution of the feature map can be consistent with the fundus image (i.e., the feature map can be single-scale and consistent with the size of the fundus image). This improves the accuracy of identifying or measuring targets with small size variations. In some examples, feature maps consistent with the scale of the fundus image can be obtained by continuously fusing image features at different scales. In some examples, the feature map can be used as input to segmentation network 22 and regression network 23.

[0053] In some examples, the backbone network 21 may include an encoding module and a decoding module. In some examples, the encoding module may be configured to extract image features at different scales. In some examples, the decoding module may be configured to map the image features extracted at different scales back to the resolution of the fundus image to output a feature map. Thus, a feature map consistent with the resolution of the fundus image can be obtained.

[0054] Figure 3 This is a schematic diagram illustrating an example of a network module 20 involved in the examples of this disclosure.

[0055] In some examples, such as Figure 3 As shown, network module 20 may include a backbone network 21, a segmentation network 22, and a regression network 23. Backbone network 21 can receive fundus images and output feature maps. The feature maps can be used as input to segmentation network 22 and regression network 23 to obtain corresponding outputs. Specifically, segmentation network 22 can use the feature maps as input to obtain a first output, and regression network 23 can use the feature maps as input to obtain a second output. In this case, fundus images can be input into network module 20 to obtain both the first and second outputs.

[0056] In some examples, the first output can be the result of image segmentation prediction. In some examples, the second output can be the result of bounding box regression prediction.

[0057] In some examples, the first output may include the probability that each pixel in the fundus image belongs to each category. In some examples, the probability of each pixel belonging to each category can be obtained through an activation function. In some examples, the first output may be a matrix. In some examples, the size of the matrix corresponding to the first output may be M×N×C, where M×N can represent the resolution of the fundus image, M and N can correspond to the rows and columns of the fundus image, respectively, and C can represent the number of categories. For example, for a fundus image targeting the two categories of optic cup and optic disc, the size of the matrix corresponding to the first output may be M×N×2.

[0058] In some examples, the value corresponding to each pixel location in the fundus image in the first output can be a vector, and the number of elements in the vector can be the same as the number of categories. For example, for the pixel at the k-th location in the fundus image, the corresponding value in the first output can be a vector p. k Vector p k It can include C elements, where C can be the number of categories. In some examples, the vector p k The element value can be a number between 0 and 1.

[0059] In some examples, the second output may include the position of each pixel in the fundus image and the offset of the tight bounding box of each target category. That is, the second output may include the offset of the tight bounding box of a clearly defined target category. In other words, the regression network 23 can predict the offset of the tight bounding box of a clearly defined target category. In this case, when targets of different categories have high overlap, the tight bounding boxes of the corresponding categories can be distinguished, thus enabling the acquisition of the tight bounding boxes of the corresponding categories. Therefore, it is possible to identify or measure targets of different categories with high overlap. In some examples, the offset in the second output can be used as the target offset.

[0060] In some examples, the target offset can be a normalized offset. In other examples, the target offset can be an offset normalized based on the average size of targets across different categories. The target offset and the predicted offset (described later) can correspond to the true offset (described later). That is, if the true offset is normalized when training network module 20 (which can be simply referred to as the training phase), the target offset (corresponding to the measurement phase) and the predicted offset (corresponding to the training phase) predicted using network module 20 can also be automatically normalized accordingly. This improves the accuracy of identifying or measuring targets with small size variations.

[0061] In some examples, the average size of a target can be obtained by averaging its average width and average height. In some examples, the average size of a target can be empirical values ​​(i.e., the average width and average height can be empirical values). In some examples, the average size of a target can be obtained statistically from samples corresponding to acquired fundus images. Specifically, the average width and average height can be obtained by averaging the width and average height of the bounding boxes of targets in the labeled data of the samples by category, and then averaging the average width and average height can be obtained to obtain the average size of targets in that category. In some examples, the samples can be training samples (described later), that is, the average size of a target can be obtained statistically from the training samples. Thus, the average width and average height, or the average size of the target, can be obtained from the training samples.

[0062] In some examples, the second output can be a matrix. In some examples, the size of the matrix corresponding to the second output can be M×N×A, where A represents the total size of the target offsets, M×N represents the resolution of the fundus image, and M and N correspond to the rows and columns of the fundus image, respectively. In some examples, if the size of a target offset is a 4×1 vector (i.e., it can be represented by 4 numbers), then A can be C×4, where C represents the number of categories. For example, for fundus images containing only two categories of targets, the optic cup and optic disc, the size of the matrix corresponding to the second output can be M×N×8.

[0063] In some examples, the value corresponding to a pixel at each location in the fundus image in the second output can be a vector. For example, the value corresponding to the k-th pixel in the fundus image in the second output can be represented as: v k =[v k1 ,v k2 ,…,v kC Where C can be the number of categories, v k Each element in the expression can represent the target displacement for each category of target. This allows for convenient representation of target displacement and its corresponding category. In some examples, v k The elements can be 4-dimensional vectors.

[0064] In some examples, the backbone network 21 may be based on a U-net network. In this embodiment, the encoding module of the backbone network 21 may include unit layers and pooling layers. The decoding module of the backbone network 21 may include unit layers, up-sampling layers, and skip-connection units.

[0065] In some examples, unit layers may include convolutional layers, batch normalized layers, and rectified linear unit layers (ReLU). In some examples, pooling layers may be max pooling layers (max-poooling). In some examples, skip connection units may be used to combine image features from deep layers and image features from shallow layers.

[0066] Additionally, the segmentation network 22 can be a feedforward neural network. In some examples, the segmentation network 22 may include multiple unit layers. In some examples, the segmentation network 22 may include multiple unit layers and convolutional layers (Conv).

[0067] Additionally, the regression network 23 may include dilated convolution layers (DilatedConv) and batch normalization layers (BN). In some examples, the regression network 23 may include dilated convolution layers, batch normalization layers, and convolutional layers.

[0068] Figure 4 This is a schematic diagram illustrating another example of the network module 20 involved in the examples of this disclosure. It should be noted that, in order to more clearly describe the network structure of the network module 20, in... Figure 4 In the diagram, the network layers in the network module 20 are distinguished by the numbers in the arrows. Arrow 1 represents a network layer (i.e., a unit layer) consisting of a convolutional layer, a batch normalization layer, and a modified linear unit layer; arrow 2 represents a network layer consisting of a dilated convolutional layer and a modified linear unit layer; arrow 3 represents a convolutional layer; arrow 4 represents a max pooling layer; arrow 5 represents an upsampling layer; and arrow 6 represents a skip connection unit.

[0069] As an example of network module 20. Figure 4 As shown, a fundus image with a resolution of 256 × 256 can be input into network module 20. Image features are extracted through different levels of unit layers (see arrow 1) and max pooling layers (see arrow 4) of the encoding module. Image features of different scales are continuously fused through different levels of unit layers (see arrow 1), upsampling layers (see arrow 5), and skip connection units (see arrow 6) of the decoding module to obtain a feature map 221 with the same scale as the fundus image. The feature map 221 is then input into segmentation network 22 and regression network 23 to obtain the first output and the second output, respectively.

[0070] In addition, such as Figure 4 As shown, the segmentation network 22 can be composed of unit layers (see arrow 1) and convolutional layers (see arrow 3) in sequence, and the regression network 23 can be composed of multiple network layers consisting of dilated convolutional layers and modified linear unit layers (see arrow 2), as well as convolutional layers (see arrow 3). Among them, the unit layer can be composed of convolutional layers, batch normalization layers, and modified linear unit layers.

[0071] In some examples, the kernel size of the convolutional layers in network module 20 can be set to 3×3. In some examples, the kernel size of the max-pooling layers in network module 20 can be set to 2×2, and the stride can be set to 2. In some examples, the scale-factor of the upsampling layers in network module 20 can be set to 2. In some examples, such as... Figure 4 As shown, the dilation factors of the multiple dilated convolutional layers in network module 20 can be set to 1, 1, 2, 4, 8, and 16 respectively (see the numbers above arrow 2). In some examples, such as Figure 4 As shown, the number of max pooling layers can be 5. This allows the size of the fundus image to be divisible by 32 (32 can be 2 to the power of 5).

[0072] As described above, the measurement method disclosed herein is a measurement method that uses a network module 20 trained based on target tight bounding boxes to identify targets and thus achieve measurement. Hereinafter, the training method (which may be simply referred to as the training method) of the network module 20 disclosed herein will be described in detail with reference to the accompanying drawings. Figure 5 This is a flowchart illustrating the training method of the network module 20 involved in the example of this disclosure.

[0073] In some examples, the segmentation network 22 and the regression network 23 in network module 20 can be trained simultaneously in an end-to-end manner.

[0074] In some examples, the segmentation network 22 and regression network 23 in network module 20 can be jointly trained to simultaneously optimize both. In other examples, through joint training, the segmentation network 22 and regression network 23 can adjust the network parameters of the backbone network 21 via backpropagation so that the feature map output by the backbone network 21 better represents the features of the fundus image and is then input into the segmentation network 22 and regression network 23. In this case, both the segmentation network 22 and regression network 23 process based on the feature map output by the backbone network 21.

[0075] In some examples, multi-instance learning can be used to train the segmentation network 22. In some examples, the expected intersection-over-union ratios (IoUs) of the pixels in the image to be trained can be used to select the pixels to train the regression network 23 (described later).

[0076] In some examples, such as Figure 5As shown, the training method may include constructing training samples (step S120), inputting the training samples into the network module 20 to obtain prediction data (step S140), and determining the training loss of the network module 20 based on the training samples and prediction data, and optimizing the network module 20 based on the training loss (step S160). Thus, an optimized (or trained) network module 20 can be obtained.

[0077] In some examples, training samples may be constructed in step S120. The training samples may include fundus image data and label data. In some examples, the fundus image data may include multiple images to be trained. The images to be trained may be fundus images used for training.

[0078] In some examples, the multiple training images may include images containing the target. For fundus images, the target can be at least one of the optic cup and optic disc. That is, the target can belong to at least one category of the optic cup and optic disc. In some examples, the multiple training images may include images containing the target and images not containing the target. For fundus images, if the optic cup and optic disc are being identified or measured, the target in the fundus image can be one optic disc and one optic cup. That is, there are two types of targets in the fundus image that need to be identified or measured, and the number of each target can be 1; if the target is being identified or measured as either the optic cup or the optic disc, the target in the fundus image can be either one optic cup or one optic disc.

[0079] In some examples, the label data may include the gold standard for the target's category (the gold standard for the category is sometimes also called the true category) and the gold standard for the target's tight bounding box label (the gold standard for the tight bounding box label is sometimes also called the true tight bounding box label). That is, the label data can be the true category of the target in the image to be trained and the true tight bounding box label of the target. It should be noted that, unless otherwise specified, the tight bounding box label or the category of the target in the label data in the training method can be assumed to be the gold standard.

[0080] In some examples, the training images can be labeled to obtain label data. In other examples, annotation tools such as line annotation systems can be used to annotate the training images. Specifically, annotation tools can be used to annotate the bounding boxes (i.e., the smallest bounding rectangles) of targets in the training images, and corresponding categories can be assigned to the bounding boxes to represent the true category to which the targets belong.

[0081] In some examples, to suppress overfitting of network module 20, data augmentation can be performed on the training samples. In some examples, data augmentation may include, but is not limited to, flipping (e.g., vertical or horizontal flipping), magnification, rotation, contrast adjustment, brightness adjustment, or color equalization. In some examples, the same data augmentation can be performed on both the fundus image data and the label data in the training samples. This ensures that the fundus image data and the label data are consistent.

[0082] In some examples, in step S140, training samples can be input into network module 20 to obtain prediction data. As described above, network module 20 may include a segmentation network 22 and a regression network 23. In some examples, network module 20 can obtain prediction data corresponding to the training samples based on the fundus image data of the training samples. The prediction data may include predicted segmentation data output by segmentation network 22 and predicted offset output by regression network 23.

[0083] Furthermore, the predicted segmentation data can correspond to the first output, and the predicted offset can correspond to the second output (i.e., to the target offset). That is, the predicted segmentation data can include the probability that each pixel in the training image belongs to each category, and the predicted offset can include the offset of the position of each pixel in the training image relative to the bounding box of each category of the target. In some examples, corresponding to the target offset, the predicted offset can be an offset normalized based on the average size of the targets for each category. This improves the accuracy of identifying or measuring targets with small size variations.

[0084] To more clearly describe the offset of the pixel position from the target bounding box and the normalized offset, the following description is based on the formula. It should be noted that the predicted offset, target offset, and true offset are all types of offsets and are also applicable to the formula (1) below.

[0085] Specifically, the position of a pixel can be represented as (x, y), the bounding box label of a target corresponding to the pixel can be represented as b = (xl, yt, xr, yb), and the offset of the bounding box label b of the target relative to the position of the pixel (that is, the offset between the position of the pixel and the bounding box label of the target) can be represented as t = (tl, tt, tr, tb). Then tl, tt, tr, tb can satisfy formula (1):

[0086] tl=(x-xl) / S c ,

[0087] tt=(y-yt) / S c ,

[0088] tr=(xr-x) / S c ,

[0089] tb=(yb-y) / S c ,

[0090] Where xl,yt can represent the top-left corner of the target's bounding box, xr,yb can represent the bottom-right corner of the target's bounding box, c can represent the index of the target's category, and S c The average size of the target in the c-th category can be represented. Thus, the normalized offset can be obtained. However, the examples disclosed herein are not limited to this. In other examples, the target's bounding box label can be represented by the position of the lower left corner and the upper right corner, or by the position, length, and width of any corner. Furthermore, in other examples, normalization can be performed in other ways; for example, the offset can be normalized using the length and width of the target's bounding box label.

[0091] In addition, the pixels in formula (1) can be pixels in the training image or the fundus image. That is, formula (1) can be applied to the true offset corresponding to the training image during the training phase and the target offset corresponding to the fundus image during the measurement phase.

[0092] Specifically, during the training phase, the pixel can be a pixel in the image to be trained, the target's bounding box label b can be the gold standard of the target's bounding box label in the image to be trained, and the offset t can be the true offset (also called the gold standard of offset). Thus, the regression loss of the regression network 23 can be obtained based on the predicted offset and the true offset. In addition, if the pixel is a pixel in the image to be trained and the offset t is the predicted offset, the predicted target's bounding box label can be inferred from formula (1).

[0093] Furthermore, during the measurement phase, the pixel can be a pixel in the fundus image, and the offset t can be the target offset. Then, the target's bounding box in the fundus image can be deduced from formula (1) and the target offset (that is, the target offset and the pixel position can be substituted into formula (1) to obtain the target's bounding box). Thus, the target's bounding box in the fundus image can be obtained.

[0094] In some examples, in step S160, the training loss of network module 20 can be determined based on the training samples and prediction data, and network module 20 can be optimized based on the training loss. In some examples, the training loss of network module 20 can be determined based on the label data corresponding to the training samples, the predicted segmentation data, and the predicted offset, and then network module 20 can be trained based on the training loss to optimize network module 20.

[0095] As described above, network module 20 may include a segmentation network 22 and a regression network 23. In some examples, the training loss may include the segmentation loss of segmentation network 22 and the regression loss of regression network 23. That is, the training loss of network module 20 can be obtained based on the segmentation loss and the regression loss. Thus, network module 20 can be optimized based on the training loss. In some examples, the training loss may be the sum of the segmentation loss and the regression loss. In some examples, the segmentation loss may represent the degree to which pixels in the training image in the predicted segmentation data belong to each true category, and the regression loss may represent the degree to which the predicted offset is close to the true offset.

[0096] Figure 6 This is a schematic diagram illustrating the positive package involved in the examples of this disclosure.

[0097] In some examples, the segmentation loss of segmentation network 22 can be obtained based on the predicted segmentation data and label data corresponding to the training samples. Thus, the predicted segmentation data of segmentation network 22 can be made to approximate the label data through the segmentation loss. In some examples, the segmentation loss can be obtained using multi-instance learning. In multi-instance learning, multiple training packets can be obtained based on the ground truth bounding boxes of the targets in each training image (i.e., each category can correspond to multiple training packets). The segmentation loss can be obtained based on the multiple training packets for each category. In some examples, the multiple training packets can include multiple positive packets and multiple negative packets. Thus, the segmentation loss can be obtained based on the positive and negative packets from multi-instance learning. It should be noted that, unless otherwise specified, the positive and negative packets mentioned below refer to each category.

[0098] In some examples, multiple positive packets can be obtained based on the region within the target's true bounding box. For example... Figure 6 As shown, region A2 in the training image P1 is the region within the true bounding box B21 of the target T1.

[0099] In some examples, all pixels on each of the multiple straight lines connecting the two opposite sides of the actual tight bounding box of the target can be grouped into a positive hull (i.e., one straight line can correspond to one positive hull). Specifically, the two ends of each straight line can be the top and bottom, or the left and right ends of the actual tight bounding box. As an example, such as Figure 6 As shown, the pixels on lines D1, D2, D3, D4, D5, D6, D7, and D8 can each be divided into a positive hull. However, the examples disclosed herein are not limited to this; in other examples, positive hulls can be divided in other ways. For example, pixels at specific locations on the true bounding box can be divided into a positive hull.

[0100] In some examples, multiple lines may include at least one set of mutually parallel first parallel lines. For example, multiple lines may include one set of first parallel lines, two sets of first parallel lines, three sets of first parallel lines, or four sets of first parallel lines, etc. In some examples, the number of lines in the first parallel lines may be greater than or equal to two.

[0101] In some examples, multiple straight lines may include at least one set of mutually parallel first parallel lines and mutually parallel second parallel lines, each perpendicular to each set of first parallel lines. Specifically, if multiple straight lines include one set of first parallel lines, then the multiple straight lines may also include a set of second parallel lines perpendicular to that set of first parallel lines; if multiple straight lines include multiple sets of first parallel lines, then the multiple straight lines may also include multiple sets of second parallel lines perpendicular to each set of first parallel lines. Figure 6 As shown, a first set of parallel lines may include parallel lines D1 and D2, and a corresponding second set of parallel lines may include parallel lines D3 and D4, wherein line D1 may be perpendicular to line D3; another set of first parallel lines may include parallel lines D5 and D6, and a corresponding second set of parallel lines may include parallel lines D7 and D8, wherein line D5 may be perpendicular to line D7. In some examples, the number of lines in the first and second parallel lines may be greater than or equal to two.

[0102] As described above, in some examples, multiple straight lines may include multiple sets of first parallel lines (i.e., multiple straight lines may include parallel lines at different angles). In this case, it is possible to optimize the segmentation network 22 by dividing the positive envelope at different angles. As a result, the accuracy of the segmentation network 22 in predicting segmentation data can be improved.

[0103] In some examples, the angle of the first parallel line can be the angle between the extension of the first parallel line and the extension of any non-intersecting edge of the actual frame. The angle of the first parallel line can be greater than -90° and less than 90°. For example, the angle can be -89°, -75°, -50°, -25°, -20°, 0°, 10°, 20°, 25°, 50°, 75°, or 89°, etc. Specifically, if the angle formed by rotating the extension of the non-intersecting edge clockwise by less than 90° to the extension of the first parallel line can be greater than 0° and less than 90°, and if the angle formed by rotating the extension of the non-intersecting edge counterclockwise by less than 90° (that is, clockwise by more than 270°) to the extension of the first parallel line can be greater than -90° and less than 0°, and if the non-intersecting edge is parallel to the first parallel line, then the angle can be 0°. Figure 6As shown, the angles of lines D1, D2, D3, and D4 can be 0°, while the angles (i.e., angle C1) of lines D5, D6, D7, and D8 can be 25°. In some examples, the angle of the first parallel line can be a hyperparameter that can be optimized during training.

[0104] Alternatively, the angle of the first parallel line can be described by rotating the image to be trained. The angle of the first parallel line can be a rotation angle. Specifically, the angle of the first parallel line can be the rotation angle by rotating the image to be trained so that any side of the image to be trained that does not intersect with the first parallel line is parallel to the first parallel line. The angle of the first parallel line can be greater than -90° and less than 90°. The rotation angle for clockwise rotation can be a positive degree, and the rotation angle for counterclockwise rotation can be a negative degree.

[0105] However, the examples disclosed herein are not limited to this. In other examples, the angle of the first parallel line may be in other ranges depending on how it is described. For example, if described based on the edge of the actual tight frame intersecting the first parallel line, the angle of the first parallel line may be greater than 0° and less than 180°.

[0106] In some examples, multiple negative packets can be obtained based on the region outside the target's true bounding box. For example... Figure 6 As shown, region A1 in the training image P1 is the region outside the true bounding box B21 of target T1. In some examples, the negative packet can be a single pixel of the region outside the true bounding box of all targets in a class (i.e., one pixel can correspond to one negative packet).

[0107] As mentioned above, in some examples, segmentation loss can be obtained based on multiple training bags for each category. In some examples, the segmentation loss can include unary terms (also called univariate loss) and pairwise terms (also called pairwise loss). In some examples, the unary term can describe the degree to which each training bag belongs to each ground truth class. In this case, the univariate loss allows the tight bounding boxes to be constrained by both positive and negative bags simultaneously. In some examples, the pairwise term can describe the degree to which a pixel in the training image belongs to the same category as its neighboring pixels. In this case, the pairwise loss smooths the predicted segmentation results.

[0108] In some examples, the segmentation loss can be obtained by class, and the segmentation loss (i.e., the total segmentation loss) can be obtained based on the class segmentation loss. In some examples, the total segmentation loss L... seg The formula can be satisfied:

[0109]

[0110] Among them, Lc C can represent the segmentation loss for category c, where C represents the number of categories. For example, if the optic cup and optic disc are to be identified in a fundus image, C can be 2; if only the optic cup or only the optic disc is to be identified, C can be 1.

[0111] In some examples, the segmentation loss L for class c c The formula can be satisfied:

[0112]

[0113] Where, φ c It can represent a unary term. It can be represented as pairs, where P represents the degree (or probability) that each pixel predicted by the segmentation network belongs to each category, and B... c + can represent a set of multiple positive bags, B c - can represent a set of multiple negative packets, and λ can represent a weighting factor. The weighting factor λ can be a hyperparameter and can be optimized during training. In some examples, the weighting factor λ can be used to switch between two losses (i.e., unary and pairwise losses).

[0114] Generally, in multi-instance learning, if each positive packet of a class contains at least one pixel belonging to that class, then the pixel with the highest probability of belonging to that class in each positive packet can be considered a positive sample of that class. Conversely, if none of the negative packets of a class belong to that class, then even the pixel with the highest probability in a negative packet is considered a negative sample of that class. Based on this, in some examples, the unary term φ corresponding to class c... c The formula can be satisfied:

[0115]

[0116] Among them, P c (b) can represent the probability that a training bag belongs to class c (also called the degree of belonging to class c or the probability of a training bag), where b can represent a training bag. It can represent a set of multiple positive bags. It can represent a set of multiple negative packets. max can represent a function that maximizes the maximum value. β can represent the cardinality (i.e., the number of elements in the set) of multiple positive bags, β can represent the weighting factor, and γ can represent the focusing parameter. In some examples, when the P corresponding to the positive bag... c (b) P corresponding to a negative packet equal to 1 c (b) The value of the univariate term is minimized when it equals 0. That is, the univariate loss is minimized.

[0117] In some examples, the weighting factor β can be between 0 and 1. In some examples, the focusing parameter γ can be greater than or equal to 0.

[0118] In some examples, P c (b) can be the maximum probability that a pixel in a training bag belongs to class c. In some examples, P c (b) can satisfy the formula: P c (b) = max k∈b (p kc ), where p kc It can represent the probability that the pixel at the k-th position of the training packet b belongs to category c.

[0119] In some examples, the maximum probability of a pixel belonging to a class in a training bag can be obtained based on the smooth maximum approximation function (i.e., obtaining P). c (b)). Thus, a relatively stable maximum probability can be obtained.

[0120] In some examples, the maximum smoothing approximation function can be at least one of the α-softmax function and the α-quasimax function.

[0121] In some examples, for the maximum value function f(x) = max 1≤i≤n x i `max` can represent the maximum value function, `n` can represent the number of elements (which can correspond to the number of pixels in the training data set), and `x`. i The value of an element can represent the probability that the pixel at position i in the training bag belongs to a class. In this case, the α-softmax function can satisfy the formula:

[0122]

[0123] Here, α can be a constant. In some examples, the larger α is, the closer it is to the maximum value of the maximum function.

[0124] In addition, the α-quasimax function can satisfy the formula:

[0125]

[0126] Here, α can be a constant. In some examples, the larger α is, the closer it is to the maximum value of the maximum function.

[0127] As mentioned above, in some examples, pairwise terms can describe the degree to which a pixel in the training image belongs to the same category as its neighboring pixels. That is, pairwise terms can assess the similarity in the probabilities that neighboring pixels belong to the same category. In some examples, the pairwise terms corresponding to category c... The formula can be satisfied:

[0128]

[0129] Here, ε can represent the set of all adjacent pixel pairs, (k,k') can represent a pair of adjacent pixels, and k and k' can represent the positions of the two pixels in the adjacent pixel pair, respectively. kc p can represent the probability that the pixel at position k belongs to category c. k'c It can represent the probability that the pixel at position k' belongs to category c.

[0130] In some examples, neighboring pixels can be eight-neighbor or four-neighbor pixels. In some examples, the neighboring pixels of each pixel in the image to be trained can be obtained to obtain a set of neighboring pixel pairs.

[0131] As mentioned above, training loss can include regression loss. In some examples, the regression loss of regression network 23 can be obtained based on the predicted offsets corresponding to the training samples and the true offsets corresponding to the label data. In this case, the predicted offsets of regression network 23 can be approximated to the true offsets through the regression loss.

[0132] In some examples, the true offset can be the offset between the position of a pixel in the training image and the true bounding box of the target in the label data. In other examples, corresponding to the predicted offset, the true offset can be the offset normalized based on the average size of the targets for each class. For details, please refer to the relevant description of the offset in formula (1) above.

[0133] In some examples, corresponding pixels can be selected from the images to be trained as positive samples to train the regression network 23. That is, the regression network 23 can be optimized using positive samples. Specifically, the regression loss can be obtained based on the positive samples, and then the regression network 23 can be optimized using the regression loss.

[0134] In some examples, the regression loss can satisfy the formula:

[0135]

[0136] Where C can represent the number of categories, M c t can represent the number of positive samples in the c-th class. icv can represent the true offset corresponding to the i-th positive sample of the c-th class. ic Let s(x) represent the predicted offset corresponding to the i-th positive sample of the c-th class, and let s(x) represent the sum of the smooth L1 losses of all elements in x. In some examples, for x = t ic -v ic ,s(t ic -v ic The value of can represent the degree to which the predicted offset corresponding to the i-th positive sample of the c-th class, calculated using smooth L1 loss, matches the true offset corresponding to the i-th positive sample. Here, a positive sample can be a pixel in the training image selected for training the regression network 23 (i.e., for calculating the regression loss). Thus, the regression loss can be obtained. In some examples, the true offset corresponding to a positive sample can be the offset corresponding to the true bounding box.

[0137] In some examples, the smooth L1 loss function can satisfy the formula:

[0138]

[0139] Here, σ can represent a hyperparameter used to switch between the smooth L1 loss function and the smooth L2 loss function, and x can represent a variable of the smooth L1 loss function.

[0140] As mentioned above, in some examples, corresponding pixels can be selected from the pixels in the image to be trained as positive samples to train the regression network 23.

[0141] In some examples, positive samples can be pixels in the training image that fall within the true bounding box of at least one target (i.e., pixels in the training image that fall within the true bounding box of at least one target can be selected as positive samples). In this case, optimizing the regression network 23 based on pixels that fall within the true bounding box of at least one target can improve the efficiency of the regression network 23 optimization. In some examples, pixels in the training image that fall within the true bounding box of at least one target can be selected as positive samples for each category. In some examples, the regression loss for each category can be obtained based on the positive samples for each category.

[0142] As described above, pixels falling within the true bounding box of at least one target can be selected from the training image as positive samples for each category. In some examples, the positive samples for each category can be filtered, and the regression network 23 can be optimized based on the filtered positive samples. That is, the positive samples used to calculate the regression loss can be the filtered positive samples.

[0143] In some examples, after obtaining positive samples for each category (i.e., selecting pixels that fall within the true tight bounding box of at least one target from the training image as positive samples), the matching tight bounding box corresponding to the positive sample can be obtained, and then the positive samples for each category can be filtered based on the matching tight bounding box. Thus, the regression network 23 can be optimized using the positive samples for each category filtered based on the matching tight bounding box.

[0144] In some examples, the true tight bounding box of a pixel (e.g., a positive sample) can be filtered to obtain the matching tight bounding box of that pixel. In some examples, for fundus images, the matching tight bounding box can be the true tight bounding box of the pixel. For positive samples, the matching tight bounding box can be the true tight bounding box of the positive sample. That is, the true tight bounding box can be used as the matching tight bounding box of the pixel (e.g., a positive sample).

[0145] In some examples, the expected intersection-union (OCU) of pixels (e.g., positive samples) can be used to filter positive samples from each category. In this case, pixels that are far from the center of the true tight bounding box or the matching tight bounding box can be filtered out. Thus, the adverse effects of pixels far from the center on the optimization of regression network 23 can be reduced, and the efficiency of the optimization of regression network 23 can be improved.

[0146] In some examples, the expected intersection-union ratio (OCR) of positive samples can be obtained based on the matching tight bounding boxes, and positive samples of each category can be filtered based on the OCR. Specifically, after obtaining positive samples of each category, the matching tight bounding boxes corresponding to the positive samples can be obtained, and then the OCR of the positive samples corresponding to the positive samples can be obtained based on the matching tight bounding boxes. Positive samples of each category can be filtered based on the OCR, and finally, the filtered positive samples of each category can be used to optimize the regression network 23. However, the examples disclosed herein are not limited to this. In some examples, the pixels of the training image can be filtered by category and using the OCR of the pixels of the training image (that is, the pixels of the training image can be filtered using the OCR without first selecting pixels that fall within at least one target's true tight bounding box as positive samples). In addition, pixels that do not fall within any true tight bounding boxes (that is, pixels that do not have matching tight bounding boxes) can be marked. This facilitates subsequent filtering of the pixels. For example, the OCR of a pixel can be set to 0 to mark the pixel. Specifically, the regression network 23 can be optimized based on the expected intersection of the pixels corresponding to the pixels in the training image, according to the category, and the pixels in the training image can be filtered.

[0147] In some examples, the regression network 23 can be optimized by selecting pixels from the training image whose expected cross-union ratio (CUI) is greater than a preset expected CUI. In other examples, the regression network 23 can be optimized by selecting positive samples from each class whose expected CUI is greater than a preset expected CUI. This allows for the acquisition of pixels (e.g., positive samples) that meet the preset expected CUI. In some examples, the preset expected CUI can be greater than 0 and less than or equal to 1. For example, the preset expected CUI can be 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or 1. In some examples, the preset expected CUI can be a hyperparameter. The preset expected CUI can be adjusted during the training process of the regression network 23.

[0148] In some examples, the expected intersection-over-union (OCU) ratio for a pixel can be obtained based on the matching tight bounding box of the pixel (e.g., a positive sample). In some examples, if a pixel does not have a matching tight bounding box, the pixel can be ignored or its expected OCU ratio can be set to 0. In this case, pixels without matching tight bounding boxes can be excluded from the training of the regression network 23 or have their contribution to the regression loss reduced. It should be noted that, unless otherwise specified, the following description of the expected OCU ratio for pixels also applies to the expected OCU ratio for positive samples.

[0149] In some examples, the expected intersection-over-union (IOU) ratio can be the maximum of the IOU ratios of the matching tight bounding box of a pixel with the multiple bounding boxes constructed centered on that pixel. Thus, the expected IOU ratio can be obtained. However, the examples in this disclosure are not limited to this; in other examples, the expected IOU ratio can be the maximum of the IOU ratios of the actual tight bounding box of a pixel with the multiple bounding boxes constructed centered on that pixel.

[0150] In some examples, multiple bounding boxes can be constructed centered on pixels in the image to be trained. The maximum value among the intersection-union ratios (IU) of each bounding box with its matching tight bounding box label at that pixel is then used as the expected IU. In some examples, the dimensions of the multiple bounding boxes can be different. Specifically, the width or height of each bounding box can differ from the width of the other bounding boxes.

[0151] Figure 7 This is a schematic diagram illustrating a pixel-centered bounding box as described in the example of this disclosure. To more clearly describe the desired intersection-union ratio, the following is combined with... Figure 7 Describe it. For example... Figure 7 As shown, pixel M1 has a matching tight border label B31, and border B32 is an exemplary border constructed with pixel M1 as the center.

[0152] In some examples, let W be the width of the matching tight bounding box, H be the height of the matching tight bounding box, (r1W, r2H) represent the position of a pixel point, where r1 and r2 are the relative positions of the pixel point within the matching tight bounding box, and satisfy the condition: 0 < r1, r2 < 1. Multiple bounding boxes can be constructed based on the pixel point. As an example, as Figure 7 shown, the position of pixel point M1 can be expressed as (r1W, r2H), and the width and height of the matching tight bounding box B31 can be W and H respectively.

[0153] In some examples, the matching tight bounding box can be divided into four regions by two centerlines of the matching tight bounding box. The four regions can be the upper left region, the upper right region, the lower left region, and the lower right region. For example, as Figure 7 shown, the centerline D9 and centerline D10 of the matching tight bounding box B31 can divide the matching tight bounding box B31 into the upper left region A3, the upper right region A4, the lower left region A5, and the lower right region A6.

[0154] The following takes the case where the pixel point is in the upper left region (that is, r1 and r2 satisfy the condition: 0 < r1, r2 ≤ 0.5) as an example to describe the expected intersection over union. For example, as Figure 7 shown, the pixel point M1 can be a point in the upper left region A3.

[0155] First, construct multiple bounding boxes centered on the pixel point. Specifically, for r1 and r2 satisfying the condition: 0 < r1, r2 ≤ 0.5, the four boundary conditions corresponding to the pixel point M1 can be respectively:

[0156] w1 = 2r1W, h1 = 2r2H;

[0157] w2 = 2r1W, h2 = 2(1 - r2)H;

[0158] w3 = 2(1 - r1)W, h3 = 2r2H;

[0159] w4 = 2(1 - r1)W, h4 = 2(1 - r2)H;

[0160] Among them, w1 and h1 can represent the width and height of the first boundary condition, w2 and h2 can represent the width and height of the second boundary condition, w3 and h3 can represent the width and height of the third boundary condition, and w4 and h4 can represent the width and height of the fourth boundary condition.

[0161] Secondly, calculate the intersection over union of the bounding boxes under each boundary condition with the matching tight bounding box. Specifically, the intersection over union corresponding to the above four boundary conditions can satisfy formula (2):

[0162] IoU1(r1, r2) = 4r1r2,

[0163] IoU2(r1, r2) = 2r1 / (2r1(1 - 2r2) + 1),

[0164] IoU3(r1, r2) = 2r2 / (2r2(1 - 2r1) + 1),

[0165] IoU4(r1, r2) = 1 / (4(1 - r1)(1 - r2)),

[0166] Among them, IoU1(r1, r2) can represent the intersection over union corresponding to the first boundary condition, IoU2(r1, r2) can represent the intersection over union corresponding to the second boundary condition, IoU3(r1, r2) can represent the intersection over union corresponding to the third boundary condition, and IoU4(r1, r2) can represent the intersection over union corresponding to the fourth boundary condition. In this case, the intersection over union corresponding to each boundary condition can be obtained.

[0167] Finally, the maximum intersection over union among the intersection over unions of multiple boundary conditions is the desired intersection over union. In some examples, for r1, r2 satisfying the condition: 0 < r1, r2 ≤ 0.5, the desired intersection over union can satisfy formula (3):

[0168] In addition, the desired intersection over union for pixel points in other regions (i.e., the upper - right region, lower - left region, and lower - right region) can be obtained based on a similar method for the upper - left region. In some examples, for r1 satisfying the condition: 0.5 ≤ r1 < 1, r1 in formula (3) can be replaced by 1 - r1, and for r2 satisfying the condition: 0.5 ≤ r2 < 1, r2 in formula (3) can be replaced by 1 - r2. Thus, the desired intersection over union for pixel points in other regions can be obtained. That is, pixel points in other regions can be mapped to the upper - left region through coordinate transformation, and then the desired intersection over union can be obtained in a manner consistent with that of the upper - left region. Therefore, for r1, r2 satisfying the condition: 0 < r1, r2 < 1, the desired intersection over union can satisfy formula (4):

[0169]

[0170] Among them, IoU1(r1, r2), IoU2(r1, r2), IoU2(r1, r2), and IoU2(r1, r2) can be obtained from formula (2). Thus, the desired intersection over union can be obtained.

[0171] As described above, in some examples, the expected intersection-over-union (IoU) ratio for a pixel can be obtained based on the matching tight bounding box of the pixel (e.g., a positive sample). However, the examples in this disclosure are not limited to this. In other examples, matching tight bounding boxes may not be obtained during the filtering of positive samples of each class or pixels in the image to be trained. Specifically, the expected IoU ratio for a pixel can be obtained based on the true tight bounding box of the pixel (e.g., a positive sample), and pixels of each class can be filtered based on the expected IoU ratio. In this case, the expected IoU ratio can be the maximum value among the expected IoU ratios corresponding to each true tight bounding box. For details on obtaining the expected IoU ratio based on the true tight bounding box, please refer to the relevant description of obtaining the expected IoU ratio based on the matching tight bounding box of the pixel.

[0172] The measurement method involved in this disclosure will now be described in detail with reference to the accompanying drawings. The network module 20 involved in the measurement method can be trained using the training method described above. Figure 8(a) is a flowchart illustrating the measurement method for fundus images based on tight-box label deep learning according to an example of this disclosure.

[0173] The measurement method for fundus images disclosed in this embodiment utilizes a network module 20 trained based on target tight-frame markers to identify at least one target in the fundus image, thereby achieving measurement. The fundus image may include at least one target, which may be the optic cup and / or optic disc. That is, the network module 20 trained based on target tight-frame markers can identify the optic cup and / or optic disc in the fundus image, thereby achieving measurement of the optic cup and / or optic disc. Thus, it is possible to measure the optic cup and / or optic disc in the fundus image based on tight-frame markers. In other examples, microaneurysms in the fundus image can also be identified to achieve measurement of microaneurysms.

[0174] In some examples, as shown in FIG8(a), the measurement method may include acquiring a fundus image (step S420), inputting the fundus image into the network module 20 to obtain a first output and a second output (step S440), and identifying the target based on the first output and the second output to obtain the tight frame of the optic cup and / or optic disc in the fundus image to achieve the measurement (step S460).

[0175] In some examples, a fundus image may be acquired in step S420. In some examples, the fundus image may include at least one target. In some examples, at least one target may be identified to determine the target and its category (i.e., the category of interest). For the fundus image, the category of interest may be the optic cup and / or the optic disc. The target for each category may be the optic cup or the optic disc. Specifically, if one target (optic disc or optic cup) in the fundus image is identified, the category of interest may be the optic cup or the optic disc; if both the optic disc and the optic cup in the fundus image are identified, the category of interest may be both the optic cup and the optic disc. In some examples, the fundus image may also not include the optic disc or the optic cup. In this case, it is possible to determine if a fundus image lacks the optic disc or the optic cup.

[0176] In some examples, in step 440, the fundus image can be input into network module 20 to obtain a first output and a second output. In some examples, the first output may include the probability that each pixel in the fundus image belongs to a category (i.e., optic cup and / or optic disc). In some examples, the second output may include the offset of the position of each pixel in the fundus image from the bounding box of the target for each category. In some examples, the offset in the second output may be used as the target offset. In some examples, network module 20 may include a backbone network 21, a segmentation network 22, and a regression network 23. In some examples, the segmentation network 22 may be image segmentation based on weakly supervised learning. In some examples, the regression network 23 may be based on bounding box regression. In some examples, the backbone network 21 may be used to extract feature maps from the fundus image. In some examples, the segmentation network 22 may use the feature map as input to obtain the first output, and the regression network 23 may use the feature map as input to obtain the second output. In some examples, the resolution of the feature map may be consistent with that of the fundus image. See the relevant description of network module 20 for details.

[0177] In some examples, in step S460, the target can be identified based on the first and second outputs to obtain the tight frame marker of the optic cup and / or optic disc in the fundus image, thereby enabling measurement. This allows for subsequent accurate measurement of the optic cup and / or optic disc based on the tight frame marker. In some examples, the target offset corresponding to the category (i.e., the optic cup and / or optic disc) of a pixel at the corresponding position can be selected from the second output based on the first output, and the tight frame marker of the optic cup and / or optic disc can be obtained based on this target offset.

[0178] In some examples, the position of the pixel with the highest local probability belonging to each category can be obtained from the first output as the first position, and the tight bounding box of the target for each category can be obtained based on the target offset of the corresponding position and category in the second output. In this case, one or more targets of each category can be identified. In some examples, non-maximum suppression (NMS) can be used to obtain the first position. In some examples, the number of first positions corresponding to each category can be greater than or equal to 1. For fundus images, preferably, the position of the pixel with the highest probability belonging to each category can be obtained from the first output as the first position, and the tight bounding box of the optic cup and / or optic disc can be obtained based on the target offset of the corresponding position and category in the second output. Thus, the optic cup and / or optic disc can be identified. In some examples, the first position can be obtained using the maximum method. In some examples, the position of the pixel with the highest probability belonging to each category can be obtained using the maximum method. In some examples, the first position can also be obtained using smoothed maximum suppression.

[0179] In some examples, the bounding box labels of targets for each category can be obtained based on the first position and the target offset. In some examples, the first position and the target offset can be substituted into Equation (1) to inversely deduce the target's bounding box label. Specifically, the first position can be used as the pixel position (x, y) in Equation (1) and the target offset can be used as the offset t to obtain the target's bounding box label b.

[0180] Figure 8(b) is a flowchart illustrating another example of a measurement method for fundus images based on tight-frame markers according to the present disclosure. In some examples, as shown in Figure 8(b), the measurement method may further include obtaining the ratio of the optic cup to the optic disc based on tight-frame markers of the optic cup and optic disc in the fundus image (step S480). Thus, the ratio of the optic cup to the optic disc can be accurately measured based on the tight-frame markers of the optic cup and optic disc.

[0181] In some examples, after obtaining the rim marker of the optic cup and / or optic disc in step S460, the dimensions of the optic cup and / or optic disc can be obtained based on the rim marker of the optic cup and / or optic disc in the fundus image (the dimensions can be, for example, vertical and horizontal diameters). This allows for accurate measurement of the dimensions of the optic cup and / or optic disc. In some examples, the height of the rim marker can be used as the vertical diameter of the optic cup and / or optic disc, and the width of the rim marker can be used as the horizontal diameter of the optic cup and / or optic disc to obtain the dimensions of the optic cup and / or optic disc.

[0182] In some examples, the ratio of the optic cup and optic disc (also known as the cup-to-disc ratio) can be obtained after determining the dimensions of the optic cup and optic disc based on the close-frame index. In this case, obtaining the ratio of the optic cup and optic disc based on the close-frame index allows for a precise measurement of the cup-to-disc ratio.

[0183] In some examples, the cup-to-disc ratio can include a vertical cup-to-disc ratio and a horizontal cup-to-disc ratio. The vertical cup-to-disc ratio can be the ratio of the vertical diameters of the cup and the disk. The horizontal cup-to-disc ratio can be the ratio of the horizontal diameters of the cup and the disk.

[0184] In some examples, the bounding box of the optic cup in the fundus image is labeled b. oc =(xl oc ,yt oc ,xr oc ,yb oc The frame of the display is marked as b. od =(xl od ,yt od ,xr od ,yb od ), where b oc and b od The first two values ​​represent the top-left corner of the heading, and the last two values ​​represent the bottom-right corner.

[0185] The vertical cup-to-disc ratio can satisfy the formula: VCDR=(yb oc -yt oc ) / (yb od -yt od ),

[0186] The horizontal cup-to-plate ratio can satisfy the formula: HCDR=(xr oc -xl oc ) / (xr od -xl od ).

[0187] The measurement apparatus 200 for fundus images based on deep learning and tight-frame markers, as disclosed herein, will be described in detail below with reference to the accompanying drawings. The measurement apparatus 200 may also be referred to as an identification device, tight-frame marker measurement device, tight-frame marker identification device, automatic measurement device, auxiliary measurement device, etc. The measurement apparatus 200 disclosed herein is used to implement the measurement method described above. Figure 9(a) is a block diagram showing the measurement apparatus 200 for fundus images based on deep learning and tight-frame markers, as illustrated in the example of this disclosure.

[0188] As shown in Figure 9(a), in some examples, the measuring device 200 may include an acquisition module 50, a network module 20, and an identification module 60.

[0189] In some examples, the acquisition module 50 may be configured to acquire fundus images. In some examples, the fundus images may include at least one target. In some examples, at least one target may be identified to determine the target and its category (i.e., the category of interest). For the fundus image, the category of interest (or simply category) may be the optic cup and / or optic disc. The target for each category may be the optic cup or the optic disc. See the relevant description in step S420 for details.

[0190] In some examples, network module 20 can be configured to receive fundus images and obtain a first output and a second output based on the fundus images. In some examples, the first output can include the probability that each pixel in the fundus image belongs to each category. In some examples, the second output can include the offset of the position of each pixel in the fundus image from the bounding box of the target for each category. In some examples, the offset in the second output can be used as the target offset. In some examples, network module 20 can include a backbone network 21, a segmentation network 22, and a regression network 23. In some examples, the segmentation network 22 can be image segmentation based on weakly supervised learning. In some examples, the regression network 23 can be based on bounding box regression. In some examples, the backbone network 21 can be used to extract feature maps from the fundus image. In some examples, the segmentation network 22 can take the feature map as input to obtain the first output, and the regression network 23 can take the feature map as input to obtain the second output. In some examples, the resolution of the feature map can be consistent with that of the fundus image. See the relevant description of network module 20 for details.

[0191] In some examples, the recognition module 60 can be configured to recognize the target based on the first and second outputs to obtain the optic cup and / or optic disc outlines in the fundus image for measurement. See the relevant description in step S460 for details.

[0192] Figure 9(b) is a block diagram illustrating another example of the fundus image measurement device 200 based on tight-frame deep learning according to the present disclosure. Figure 9(c) is a block diagram illustrating another example of the fundus image measurement device 200 based on tight-frame deep learning according to the present disclosure.

[0193] As shown in Figures 9(b) and 9(c), in some examples, the measuring device 200 may also include a cup-to-disc ratio module 70. The cup-to-disc ratio module 70 may be configured to obtain the ratio of the optic cup to the optic disc based on the tight borders of the optic cup and optic disc in the fundus image. See the relevant description in step S480 for details.

[0194] The measurement method and apparatus 200 disclosed herein construct a network module 20 comprising a backbone network 21, a segmentation network 22 based on weakly supervised learning for image segmentation, and a regression network 23 based on bounding box regression. The network module 20 is trained based on the target's tight bounding box. The backbone network 21 receives a fundus image and extracts a feature map with the same resolution as the fundus image. The feature map is input into the segmentation network 22 and the regression network 23 respectively to obtain a first output and a second output. Then, based on the first and second outputs, the tight bounding box of the optic cup and / or optic disc in the fundus image is obtained, thereby achieving measurement. In this case, the network module 20 trained based on the tight bounding box of the optic cup and / or optic disc can accurately predict the tight bounding box of the optic cup and / or optic disc in the fundus image, thus enabling accurate measurement based on the tight bounding box of the optic cup and / or optic disc. Furthermore, by predicting the normalized offset through the regression network 23, the accuracy of identifying or measuring the optic cup and / or optic disc, whose size changes little, can be improved. Furthermore, using the expected intersection-union ratio (OCR) to filter the pixels used to optimize the regression network 23 can reduce the adverse effects of pixels far from the center on the optimization of the regression network 23 and improve the efficiency of the optimization. In addition, since the regression network 23 predicts a clear category shift, it can further improve the accuracy of visual cup and / or visual disc recognition or measurement.

[0195] While the present disclosure has been specifically described above in conjunction with the accompanying drawings and examples, it is to be understood that the foregoing description does not limit the present disclosure in any way. Those skilled in the art can make modifications and variations to the present disclosure as needed without departing from its essential spirit and scope, and all such modifications and variations shall fall within the scope of the present disclosure.

Claims

1. A method for training network modules based on tight bounding boxes, characterized in that, The network module includes a backbone network, a segmentation network for image segmentation, and a regression network based on bounding box regression. The tight bounding box is the minimum bounding rectangle of the target. The training method includes: constructing training samples, wherein the fundus image data of the training samples includes multiple images to be trained, wherein the multiple images to be trained include images containing at least one target, namely the optic cup and the optic disc; the label data of the training samples includes the gold standard for the category to which the target belongs and the gold standard for the tight bounding box of the target; and obtaining, through the network module, predicted segmentation data output by the segmentation network and predicted offset output by the regression network corresponding to the training samples based on the fundus image data of the training samples. The backbone network is used to extract feature maps from the images to be trained, and the segmentation network... The feature map is used as input to output the predicted segmentation data. The regression network uses the feature map as input to output the predicted offset. The predicted segmentation data includes the probability that each pixel in the training image belongs to the category of optic cup and / or optic disc. The predicted offset includes the offset of the position of each pixel in the training image from the bounding box label of each category of the target. The value corresponding to each position of the pixel in the training image in the predicted offset is a vector, and each element in the vector represents the target displacement of the target in each category. The training loss of the network module is determined based on the label data corresponding to the training samples, the predicted segmentation data, and the predicted offset. The network module is then trained based on the training loss to optimize the network module. The segmentation network is based on weakly supervised learning. The training loss includes the segmentation loss of the segmentation network obtained based on the predicted segmentation data and label data corresponding to the training samples, and the regression loss of the regression network obtained based on the predicted offset corresponding to the training samples and the true offset corresponding to the label data. The true offset is the offset between the position of the pixel in the image to be trained and the gold standard of the tight bounding box of the target in the label data. Multiple training packets, including multiple positive packets and multiple negative packets, are obtained according to the gold standard of the tight bounding box of the target by category, and the segmentation loss is obtained based on the multiple training packets of each category.

2. The training method according to claim 1, characterized in that: The predicted offset is the offset normalized based on the average size of targets in each category.

3. The training method according to claim 1, characterized in that: The positive bag is defined as the area outside the gold standard of the tight frame of all targets in a category, consisting of all pixels on each of the multiple straight lines connecting the two opposite sides of the gold standard of the tight frame of the target. The multiple straight lines include at least one set of first parallel lines and two parallel lines perpendicular to each set of first parallel lines. The negative bag is defined as the area outside the gold standard of the tight frame of all targets in a category. The angle of the first parallel line is the angle between the extension of the first parallel line and the extension of any non-intersecting side of the gold standard of the tight frame of the target. The angle of the first parallel line is greater than -90° and less than 90°.

4. The training method according to claim 1, characterized in that: The regression network is optimized by filtering pixels from the training image according to category and using the expected intersection-union ratio (OCR) corresponding to the pixels of the training image. Pixels with an expected OCR greater than a preset expected OCR are selected. In this process, multiple bounding boxes of different sizes are constructed with the pixels of the training image as the center point. The maximum value among the OCRs of the multiple bounding boxes and the matching tight frame labels of the pixels is obtained and used as the expected OCR. The matching tight frame label is the gold standard for the tight frame label that the pixels of the training image fall into.

5. The training method according to claim 4, characterized in that: The expected intersection-union ratio satisfies the formula: , in, The desired intersection-union ratio, For intersection, union, and comparison, The relative position of the pixels in the image to be trained within the matching bounding box. , , , , .

6. The training method according to claim 1, characterized in that: The feature map has the same resolution as the image to be trained.

7. An apparatus for measuring an ocular fundus image based on tight frame, characterized by The device includes an acquisition module, a network module trained by the training method according to any one of claims 1 to 6, and a recognition module; the acquisition module is configured to acquire a fundus image; the network module is configured to receive the fundus image and obtain a first output output by a segmentation network and a second output output by a regression network based on the fundus image. The recognition module is configured to identify the optic cup and / or optic disc based on the first output and the second output to obtain the rim markers of the optic cup and / or optic disc in the fundus image, thereby enabling measurement.

8. A measurement method of a fundus image based on tight frame markers, characterized by, The method includes acquiring a fundus image; inputting the fundus image into a network module trained by the training method according to any one of claims 1 to 6 to obtain a first output from a segmentation network and a second output from a regression network; and identifying the optic cup and / or optic disc based on the first output and the second output to obtain the tight frame markers of the optic cup and / or optic disc in the fundus image, thereby achieving measurement.

9. The measuring device according to claim 7 or the measuring method according to claim 8, characterized in that: The position of the pixel with the highest probability belonging to each category of the viewing cup and / or viewing disc is obtained from the first output as the first position, and the tight frame label of the viewing cup and / or viewing disc is obtained based on the position in the second output corresponding to the first position and the offset of the corresponding category.