OCT image detection method and device based on multi-directional image fusion

By employing a multi-directional image fusion OCT image detection method that combines axis detection and projection detection models, OCT images are processed in multiple directions. This solves the problem of low accuracy in traditional OCT fundus image analysis and achieves higher comprehensiveness and accuracy in three-dimensional detection.

CN116934686BActive Publication Date: 2025-12-05WEIZHI MEDICAL TECH (FOSHAN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310677404.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-08
Publication Date
2025-12-05
Estimated Expiration
2043-06-08

AI Technical Summary

Technical Problem

Traditional OCT fundus image analysis relies on manual observation, resulting in low accuracy. In particular, inexperienced professionals are prone to missing crucial details.

Method used

An OCT image detection method based on multi-directional image fusion is adopted. Multiple OCT images are processed through axis detection model and projection detection model, including axis detection processing, image layering and projection processing, and finally detection box fusion is performed to improve the accuracy of analysis.

Benefits of technology

It achieves multi-directional target detection, improves the comprehensiveness and accuracy of 3D detection of OCT images, reduces errors from manual analysis, and enhances the precision of the final detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116934686B_ABST
    Figure CN116934686B_ABST
Patent Text Reader

Abstract

The application discloses an OCT image detection method and device based on multi-direction image fusion, and the method is applied to a target detection model. The method comprises the following steps: when the target detection model inputs multiple OCT images, an axis detection model included in the target detection module performs an axis detection processing operation on all the OCT images to obtain axis detection processing results of all the OCT images; a projection detection model included in the target detection module performs image layering and a first projection processing operation on all the OCT images to obtain projection data of all the OCT images, and performs a second projection processing operation on the axis detection processing results and the projection data to obtain projection processing results corresponding to the projection data; and a detection frame fusion operation is further performed on the projection processing results and the axis detection processing results to obtain target detection information of a target detection factor. It can be seen that the application can reduce the probability that the analysis accuracy of an OCT fundus image is not high due to manual reasons, and is favorable for improving the analysis accuracy of the OCT fundus image.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of OCT image detection and processing technology, and particularly relates to an OCT image detection method and device based on multi-directional image fusion. BACKGROUND

[0002] The optical coherence tomography (OCT) is an imaging technology developed rapidly in recent ten years. It detects the backscattering or multiple scattering signals of incident weak coherent light from different depth layers of biological tissues by using the basic principle of weak coherent light interferometer, and obtains two-dimensional or three-dimensional structure images of biological tissues through scanning. By using this principle, the OCT device for ophthalmology can obtain the fundus image of the scanned person.

[0003] In the traditional process, professionals usually directly observe the fundus image of the scanned person, and then analyze and judge the image information covered by the fundus image according to experience. However, the information covered by the fundus image is relatively complex, and even for professionals with relevant experience, the observation of the fundus image is still difficult, and the professionals may miss some key details due to negligence. If the person in charge with little experience analyzes the fundus image, the probability of missing key details is higher. Therefore, it is particularly important to provide a method for improving the accuracy of OCT fundus image analysis. SUMMARY

[0004] The technical problem to be solved by the present application is to provide an OCT image detection method and device based on multi-directional image fusion, which can reduce the probability of low accuracy of OCT fundus image analysis caused by artificial reasons and improve the accuracy of OCT fundus image analysis.

[0005] To solve the above technical problems, the present application discloses an OCT image detection method based on multi-directional image fusion in the first aspect, which is applied to a target detection model, and the method comprises the following steps:

[0006] When it is determined that the target detection model has inputted multiple frames of OCT images, all the OCT images are transmitted to a sub-detection model included in the target detection model, the sub-detection model at least includes an axis detection model and a projection detection model, and the target detection model is used to determine a target detection factor from all the OCT images;

[0007] performing an axis detection processing operation on all the OCT images by the axis detection model to obtain an axis detection processing result corresponding to all the OCT images, the axis detection processing operation at least including a target detection operation, a neighboring frame detection operation based on a confidence value, a multi-frame fusion operation, and a confidence value determination operation; the axis detection processing result including at least one three-dimensional axis detection box and an axis confidence value corresponding to the target detection factor in each three-dimensional axis detection box;

[0008] performing an image layering and first projection processing operation on all the OCT images by the projection detection model to obtain projection data of all the OCT images, and performing a second projection processing operation on the axis detection processing result and the projection data to obtain a projection processing result corresponding to the projection data, the projection processing result including a three-dimensional projection detection box and a projection confidence value corresponding to the target detection factor in the three-dimensional projection detection box;

[0009] performing a detection box fusion operation on the projection processing result and the axis detection processing result to obtain target detection information corresponding to the target detection factor, the target detection information including a prediction box region where the target detection factor is located and a target confidence value of the target detection factor in the prediction box region.

[0010] As an optional implementation, in the first aspect of the present application, the axis detection model includes a fast-axis detection model and / or a slow-axis detection model, the fast-axis detection model is used to process data in the first dimension and the second dimension of the OCT image, and the slow-axis detection model is used to process data in the first dimension and the third dimension of the OCT image; the projection detection model is used to process data in the second dimension and the third dimension of the OCT image.

[0011] The three-dimensional axis detection box corresponding to the target detection factor is a three-dimensional axis detection box matching the processing dimension corresponding to the fast-axis detection model and / or the slow-axis detection model;

[0012] When the axis detection model performing the axis detection processing operation on all the OCT images includes the fast-axis detection model and the slow-axis detection model, the three-dimensional axis detection box corresponding to the target detection factor is a three-dimensional axis detection box obtained by performing a fusion operation on the three-dimensional axis detection boxes corresponding to the fast-axis detection model and the slow-axis detection model respectively.

[0013] As an optional implementation, in the first aspect of the present application, the performing an axis detection processing operation on all the OCT images by the axis detection model to obtain an axis detection processing result corresponding to all the OCT images includes:

[0014] performing feature detection operations on all the OCT images based on the feature information corresponding to the target detection factor, to obtain at least one frame of first OCT image, the first OCT image comprising a first detection box, and a confidence value of the target detection factor in the first detection box being greater than a first preset confidence threshold;

[0015] performing the feature detection operations on the preceding and subsequent frames of images corresponding to the first OCT image frame by frame based on the first OCT image as a reference image, the coordinate interval of the region where the first detection box is located as a reference, and a preset coordinate interval error value, to obtain a plurality of frames of second OCT image, each frame of the second OCT image comprising a second detection box, and a confidence value of the target detection factor in the second detection box being greater than the first preset confidence threshold;

[0016] determining, by the axis detection model, a maximum overlap region of the first detection box and all the second detection boxes as a two-dimensional axis detection box corresponding to the target detection factor;

[0017] determining a three-dimensional axis detection box corresponding to the target detection factor based on a total number of the first detection box and all the second detection boxes as a third dimension of the two-dimensional axis detection box, and a second preset confidence threshold as a screening condition; the three-dimensional axis detection box comprising an axis confidence value of the target detection factor in the three-dimensional axis detection box, and the axis confidence value being greater than the second preset confidence threshold;

[0018] projecting the three-dimensional axis detection box on a projection plane composed of a second dimension and a third dimension in any of the OCT images to obtain a two-dimensional axis-projection region of the three-dimensional axis detection box on the projection plane, the two-dimensional axis-projection region comprising a confidence value of the target detection factor in the three-dimensional axis detection box projected on the projection plane;

[0019] determining the three-dimensional axis detection box and the two-dimensional axis-projection region as an axis detection processing result corresponding to all the OCT images.

[0020] As an optional implementation, in the first aspect of the present application, the image layering and first projection processing operations performed by the projection detection model on all the OCT images to obtain projection data of all the OCT images comprise:

[0021] determining, by the projection detection model, an interlayer boundary of each frame of the OCT image according to a preset interlayer boundary determination algorithm, and then performing segmentation and layering processing on each frame of the OCT image to obtain a segmentation and layering processing result of each frame of the OCT image;

[0022] performing, by the projection detection model, projection processing operation on the segmentation and layering processing result of each frame of the OCT images according to a preset projection processing algorithm with the first dimension as the projection direction, to obtain the projection processing result corresponding to each frame of the OCT images as the projection data of the frame of OCT images, and the projection processing algorithm includes an average value projection or a maximum / minimum value projection processing algorithm;

[0023] performing, by the projection detection model, image splicing on the projection data of the OCT images of the preset frame number in the channel direction corresponding to the projection detection model, to obtain an input projection image of a projection feature extraction network for inputting the projection detection model as the projection data of all the OCT images.

[0024] As an optional implementation, in the first aspect of the present application, the projection feature extraction network includes a plurality of projection feature extraction layers; and the performing of the second projection processing operation on the axis detection processing result and the projection data to obtain the projection processing result corresponding to the projection data includes:

[0025] performing, by the projection feature extraction network, feature extraction operation on the input projection image to obtain the feature extraction result corresponding to the input projection image inputting any projection feature extraction layer;

[0026] For any projection feature extraction layer, determining, by the projection feature extraction network, a first image size corresponding to the two-dimensional axis-projection region and a second image size of the feature extraction result corresponding to the projection feature extraction layer; and performing, according to a preset size adjustment algorithm with the second image size as the reference, size adjustment operation on the two-dimensional axis-projection region to obtain a size adjustment result corresponding to the two-dimensional axis-projection region, and the size adjustment operation includes at least one of interpolation scaling processing, convolution processing, normalization processing, and weighted value conversion processing;

[0027] performing, by the projection detection model, weighted processing on the feature extraction result and the size adjustment result to obtain a weighted processing result corresponding to the feature extraction result, and the weighted processing result includes a two-dimensional projection detection box corresponding to the target detection factor, and a confidence value corresponding to the target detection factor in the two-dimensional projection detection box is greater than a third preset confidence threshold;

[0028] determining, by the projection detection model, a three-dimensional projection detection box corresponding to the two-dimensional projection detection box as the projection processing result corresponding to the projection data with the first dimension as the image expansion direction and in combination with two image layering boundary lines corresponding to the weighted processing result;

[0029] Two image layer boundary lines are two straight lines with the longest straight line distance among all layer boundary lines included in the weighting processing result.

[0030] As an optional implementation, in the first aspect of the present application, the size adjustment operation performed on the two-dimensional axial-projection region according to the preset size adjustment algorithm based on the second image size to obtain the size adjustment result corresponding to the two-dimensional axial-projection region comprises:

[0031] The interpolation processing is performed on the two-dimensional axial-projection region based on the second image size to obtain an interpolation processing result corresponding to the two-dimensional axial-projection region, and the image size of the two-dimensional axial-projection region in the interpolation processing result is a third image size.

[0032] The interpolation processing result is input into a preset target processing layer to obtain a target processing result with the image size of the second image size, and the target processing result is converted into a weighting value in a preset value interval through a preset activation function, serving as the size adjustment result corresponding to the two-dimensional axial-projection region.

[0033] The target processing layer comprises one or more of a plurality of preset convolution layers, normalization layers and activation layers.

[0034] As an optional implementation, in the first aspect of the present application, the detection frame fusion operation performed on the projection processing result and the axial detection processing result to obtain the target detection information corresponding to the target detection factor comprises:

[0035] The intersection region between the three-dimensional projection detection frame and the three-dimensional axial detection frame is determined as a prediction frame region.

[0036] The average value of the projection confidence value corresponding to the three-dimensional projection detection frame and the axial confidence value corresponding to the three-dimensional axial detection frame is calculated to obtain a target confidence value, serving as the confidence value corresponding to the target detection factor in the prediction frame region.

[0037] The prediction frame region and the target confidence value are determined as the target detection information corresponding to the target detection factor.

[0038] The second aspect of the present application discloses an OCT image detection device based on multi-directional image fusion, which comprises:

[0039] The device is applied to a target detection model, and the device comprises:

[0040] transmitting, when it is determined that the target detection model has inputted the multiple frames of OCT images, all the OCT images to a sub-detection model included in the target detection model, the sub-detection model at least including an axis detection model and a projection detection model, the target detection model being configured to determine a target detection factor from all the OCT images;

[0041] an axis detection module configured to perform an axis detection processing operation on all the OCT images by the axis detection model to obtain an axis detection processing result corresponding to all the OCT images, the axis detection processing operation at least including a target detection operation, a neighboring frame detection operation based on a confidence value, a multi-frame fusion operation, and a confidence value determination operation, the axis detection processing result including at least one three-dimensional axis detection box and an axis confidence value corresponding to the target detection factor in each three-dimensional axis detection box;

[0042] a first projection processing module configured to perform an image layering and first projection processing operation on all the OCT images by the projection detection model to obtain projection data of all the OCT images;

[0043] a second projection processing module configured to perform a second projection processing operation on the axis detection processing result and the projection data by the projection detection model to obtain a projection processing result corresponding to the projection data, the projection processing result including a three-dimensional projection detection box and a projection confidence value corresponding to the target detection factor in the three-dimensional projection detection box;

[0044] a fusion processing module configured to perform a detection box fusion operation on the projection processing result and the axis detection processing result to obtain target detection information corresponding to the target detection factor, the target detection information including a predicted box region in which the target detection factor is located and a target confidence value of the target detection factor in the predicted box region.

[0045] As an optional implementation form, in the second aspect, the axis detection model includes a fast-axis detection model and / or a slow-axis detection model, the fast-axis detection model being configured to process data in a first dimension and a second dimension of the OCT images, the slow-axis detection model being configured to process data in the first dimension and a third dimension of the OCT images; and the projection detection model is configured to process data in the second dimension and the third dimension of the OCT images.

[0046] The three-dimensional axis detection box corresponding to the target detection factor is a three-dimensional axis detection box matching a processing dimension corresponding to the fast-axis detection model and / or the slow-axis detection model;

[0047] When the axis detection model performing the axis detection processing operation on all the OCT images includes the fast axis detection model and the slow axis detection model, the three-dimensional axis detection frame corresponding to the target detection factor is a three-dimensional axis detection frame obtained by performing a fusion operation on the three-dimensional axis detection frames corresponding to the fast axis detection model and the slow axis detection model, respectively.

[0048] As an optional implementation, in the second aspect of the present application, the manner in which the axis detection module performs the axis detection processing operation on all the OCT images by the axis detection model to obtain the axis detection processing result corresponding to all the OCT images specifically includes:

[0049] The axis detection model performs a feature detection operation on all the OCT images based on the feature information corresponding to the target detection factor to obtain at least one first OCT image, wherein the first detection frame is included in the first OCT image, and the confidence value corresponding to the target detection factor in the first detection frame is greater than a first preset confidence threshold value.

[0050] The axis detection model performs the feature detection operation on the preceding and subsequent images corresponding to the first OCT image based on the first OCT image as a reference image and the coordinate interval of the region where the first detection frame is located, and in combination with a preset coordinate interval error value, to obtain a plurality of second OCT images, wherein each second OCT image includes a second detection frame, and the confidence value corresponding to the target detection factor in the second detection frame is greater than the first preset confidence threshold value.

[0051] The axis detection model determines the maximum overlapping region of the first detection frame and all the second detection frames as a two-dimensional axis detection frame corresponding to the target detection factor.

[0052] The total number of the first detection frame and all the second detection frames is taken as a third dimension of the two-dimensional axis detection frame, and a second preset confidence threshold value is taken as a screening condition to determine a three-dimensional axis detection frame corresponding to the target detection factor, wherein the three-dimensional axis detection frame includes an axis confidence value of the target detection factor in the three-dimensional axis detection frame, and the axis confidence value is greater than the second preset confidence threshold value.

[0053] A plane composed of a second dimension and a third dimension in any OCT image is taken as a projection plane to project the three-dimensional axis detection frame to obtain a two-dimensional axis-projection region of the three-dimensional axis detection frame in the projection plane, wherein the two-dimensional axis-projection region includes a confidence value of the target detection factor in the three-dimensional axis detection frame projected on the projection plane.

[0054] The three-dimensional axis detection frame and the two-dimensional axis-projection region are determined as the axis detection processing results corresponding to all the OCT images.

[0055] As an optional implementation, in the second aspect of the present application, the manner in which the first projection processing module performs image layering and first projection processing operations on all the OCT images to obtain projection data of all the OCT images specifically includes:

[0056] The projection detection model determines the inter-layer boundaries of each of the OCT images according to a preset inter-layer boundary determination algorithm, and then performs a segmentation and layering processing on each of the OCT images to obtain a segmentation and layering processing result of each of the OCT images;

[0057] The projection detection model performs a projection processing operation on the segmentation and layering processing result of each of the OCT images in the first dimension as the projection direction according to a preset projection processing algorithm to obtain a projection processing result corresponding to each of the OCT images as the projection data of the frame of OCT image, and the projection processing algorithm includes an average value projection or a maximum / minimum value projection processing algorithm;

[0058] The projection detection model performs image splicing on the projection data of a preset number of frames of the OCT images in the channel direction corresponding to the projection detection model to obtain an input projection image for inputting into a projection feature extraction network of the projection detection model as the projection data of all the OCT images.

[0059] As an optional implementation, in the second aspect of the present application, the projection feature extraction network includes a plurality of projection feature extraction layers; and the manner in which the second projection processing module performs a second projection processing operation on the axis detection processing result and the projection data to obtain a projection processing result corresponding to the projection data specifically includes:

[0060] The projection feature extraction network performs a feature extraction operation on the input projection image to obtain a feature extraction result corresponding to inputting the input projection image into any of the projection feature extraction layers;

[0061] For any of the projection feature extraction layers, the projection feature extraction network determines a first image size corresponding to the two-dimensional axis-projection region and a second image size of the feature extraction result corresponding to the projection feature extraction layer; and performs a size adjustment operation on the two-dimensional axis-projection region according to a preset size adjustment algorithm with the second image size as a reference to obtain a size adjustment result corresponding to the two-dimensional axis-projection region, and the size adjustment operation includes at least one of an interpolation scaling processing, a convolution processing, a normalization processing, and a weighted value conversion processing.

[0062] performing weighted processing on the feature extraction result and the size adjustment result by the projection detection model, to obtain a weighted processing result corresponding to the feature extraction result, the weighted processing result including a two-dimensional projection detection box corresponding to the target detection factor, a confidence corresponding to the target detection factor in the two-dimensional projection detection box being greater than a third preset confidence threshold;

[0063] determining, by the projection detection model, a three-dimensional projection detection box corresponding to the two-dimensional projection detection box as a projection processing result corresponding to the projection data, in combination with two image layered boundary lines corresponding to the weighted processing result, the first dimension being an image expansion direction;

[0064] wherein the two image layered boundary lines are two longest straight line distance layered boundary lines among all the layered boundary lines included in the weighted processing result.

[0065] As an optional implementation, in the second aspect of the present application, the manner in which the second projection processing module performs a size adjustment operation on the two-dimensional axis-projection region according to a preset size adjustment algorithm based on the second image size to obtain a size adjustment result corresponding to the two-dimensional axis-projection region specifically includes:

[0066] performing interpolation processing on the two-dimensional axis-projection region based on the second image size to obtain an interpolation processing result corresponding to the two-dimensional axis-projection region, an image size of the two-dimensional axis-projection region in the interpolation processing result being a third image size;

[0067] inputting the interpolation processing result into a preset target processing layer to obtain a target processing result with an image size of the second image size, and converting the target processing result into a weighted value in a preset value interval through a preset activation function, as a size adjustment result corresponding to the two-dimensional axis-projection region;

[0068] wherein the target processing layer includes one or more of a plurality of preset convolution layers, normalization layers and activation layers.

[0069] As an optional implementation, in the second aspect of the present application, the manner in which the fusion processing module performs a detection box fusion operation on the projection processing result and the axis detection processing result to obtain target detection information corresponding to the target detection factor specifically includes:

[0070] determining an intersection region between the three-dimensional projection detection box and the three-dimensional axis detection box as a prediction box region;

[0071] An average value of a projection confidence value corresponding to the three-dimensional projection detection frame and an axis confidence value corresponding to the three-dimensional axis detection frame is calculated to obtain a target confidence value as a confidence value corresponding to the target detection factor in the prediction frame region;

[0072] The prediction frame region and the target confidence value are determined as target detection information corresponding to the target detection factor.

[0073] The third aspect of the present application discloses another OCT image detection device based on multi-directional image fusion, and the device comprises:

[0074] A memory in which executable program codes are stored;

[0075] A processor coupled with the memory;

[0076] The processor invokes the executable program codes stored in the memory to execute the OCT image detection method based on multi-directional image fusion disclosed in the first aspect of the present application.

[0077] The fourth aspect of the present application discloses a computer storage medium, and the computer storage medium stores computer instructions which are invoked to execute the OCT image detection method based on multi-directional image fusion disclosed in the first aspect of the present application.

[0078] Compared with the prior art, the embodiments of the present application have the following beneficial effects:

[0079] In the embodiment of the present application, a multi-directional image fusion-based OCT image detection method is provided, which is applied to a target detection model. The method comprises the following steps: when it is determined that the target detection model has inputted multiple frames of OCT images, all the OCT images are transmitted to a sub-detection model included in the target detection model, the sub-detection model at least includes an axis detection model and a projection detection model, and the target detection model is used to determine a target detection factor from all the OCT images; the axis detection model is used to perform an axis detection processing operation on all the OCT images to obtain an axis detection processing result corresponding to all the OCT images, the axis detection processing operation at least includes a target detection operation, a neighboring frame detection operation based on a confidence value, a multi-frame fusion operation, and a confidence value determination operation; the axis detection processing result includes at least one three-dimensional axis detection frame and an axis confidence value corresponding to the target detection factor in each three-dimensional axis detection frame; the projection detection model is used to perform an image layering and first projection processing operation on all the OCT images to obtain projection data of all the OCT images, and perform a second projection processing operation on the axis detection processing result and the projection data to obtain a projection processing result corresponding to the projection data, the projection processing result includes a three-dimensional projection detection frame and a projection confidence value corresponding to the target detection factor in the three-dimensional projection detection frame; a detection frame fusion operation is performed on the projection processing result and the axis detection processing result to obtain target detection information corresponding to the target detection factor, the target detection information includes a prediction frame region where the target detection factor is located and a target confidence value of the target detection factor in the prediction frame region. It can be seen that, when the target detection model inputs the OCT images, the axis detection model and the projection detection model are used to perform a detection processing operation of a target detection factor on the OCT images respectively, two detection models represent at least two detection processing directions, and the two detection models detect and process the OCT images respectively to obtain detection processing results (detection frames) of at least two detection directions; multi-directional target detection is realized, which is different from a traditional two-dimensional image detection method of a single direction, and the three-dimensional detection comprehensiveness and accuracy of the OCT images are improved in the present application; further, after the respective detection processing results are obtained, a fusion processing of multiple detection processing results is automatically performed to obtain a unified fusion processing result, and the accuracy of the finally determined three-dimensional detection result is improved. BRIEF DESCRIPTION OF DRAWINGS

[0080] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort.

[0081] Figure 1 is a flowchart of a multi-directional image fusion-based OCT image detection method disclosed in the embodiment of the present application;

[0082] Figure 2 is another flowchart of a method for detecting an OCT image based on multi-directional image fusion according to an embodiment of the present application;

[0083] Figure 3 is a structural diagram of an apparatus for detecting an OCT image based on multi-directional image fusion according to an embodiment of the present application;

[0084] Figure 4 is another structural diagram of an apparatus for detecting an OCT image based on multi-directional image fusion according to an embodiment of the present application;

[0085] Figure 5 is a schematic diagram of three processing directions corresponding to fast axis, slow axis and Enface projection according to an embodiment of the present application. DETAILED DESCRIPTION

[0086] In order to enable persons skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by persons skilled in the art without creative work fall within the scope of protection of the present application.

[0087] The terms "first", "second", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish different objects, and are not used to describe a particular order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product, or the like that includes a series of steps or units is not limited to the listed steps or units, but can optionally include steps or units not listed or can optionally include other steps or units inherent to the process, method, product, or the like.

[0088] In this document, reference to "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the application. The appearance of the phrase in various places in the specification is not necessarily all referring to the same embodiment, nor is it necessarily referring to a particular alternative embodiment. It is explicitly understood that the embodiments described herein can be combined with each other, explicitly or implicitly.

[0089] The application discloses an OCT image detection method and device based on multi-directional image fusion.

[0090] Embodiment one

[0091] Please refer to Figure 1 , Figure 1 is a flowchart of an OCT image detection method based on multi-directional image fusion disclosed by the embodiment of the application. Wherein, Figure 1 The OCT image detection method based on multi-directional image fusion described above can be applied to an OCT image detection device based on multi-directional image fusion, and the method can also be applied to a target detection model, which is not limited in the embodiment of the application. As shown in Figure 1 The OCT image detection method based on multi-directional image fusion can include the following operations:

[0092] 101, when it is determined that the target detection model has input of multiple frames of OCT images, all the OCT images are transmitted to the sub-detection model included in the target detection model, and the sub-detection model at least includes an axis detection model and a projection detection model.

[0093] In the embodiment of the application, the target detection model is used to determine a target detection factor from all the OCT images.

[0094] In the embodiment of the application, please refer to 5, the axis detection model includes a fast-axis detection model and / or a slow-axis detection model (or a detection network); the projection detection model can adopt an Enface (C-Scan) detection model. The OCT image input into the target detection model can be Volume Data data obtained through OCT three-dimensional scanning, and the corresponding image three-dimensional size is L (depth) * W (width) * M (length).

[0095] It should be noted that the structures of the fast-axis detection model and the slow-axis detection model are basically the same, both of which take a 2D image as input (B-Scan) and output a plurality of detection boxes of target detection factors, each of which corresponds to four coordinate points of the upper left, lower left, upper right and lower right of the box and a corresponding classification probability of the target detection factor.

[0096] The fast / slow-axis detection model mainly includes three parts: a feature extraction network, a region proposal network (optional) and a classification box prediction network. The feature extraction network takes a picture as input and outputs a feature map of the image; the region proposal network takes the feature map of the image as input and outputs a region where a target detection factor may exist; and the classification box prediction network takes the feature map of the image and the region where the target detection factor may exist as input and outputs the position coordinates and category of the target detection factor. Typical fast / slow-axis detection models can be Mask-RCNN, Faster-RCNN, Cascade-RCNN, YOLO series, SSD, etc. The present method can be based on different fast / slow-axis detection models, and the structure of the specific detection model is not described in detail in the embodiments of the present application.

[0097] In the embodiments of the present application, the training method for each sub-detection model (fast-axis detection model, slow-axis detection model and projection detection model) is as follows:

[0098] 1. Collect images (including a plurality of OCT images obtained after performing fundus image scanning on a target person), calculate the mean and variance of the images, subtract the mean of the images and divide by the variance to standardize the images, and then perform random flipping, translation, rotation, brightness contrast random change, etc. on the images to augment the data.

[0099] 2. Train the feature extraction network. First, use the ImageNet dataset to train the classification network, use cross-entropy as the loss function, and use the SGD or Adam optimizer to optimize the network parameters. Take the full connection layer of the trained classification network as the feature extraction network of the target detection model.

[0100] 3. After the feature extraction network is trained, the feature extraction network and other parts of the target detection model are combined and trained. The classification of each anchor box in the region proposal network uses binary cross-entropy as the loss function, and the coordinate prediction of the bounding box uses smooth-L1 loss as the loss function; in the classification prediction network, the judgment of the category uses cross-entropy as the loss function, and the coordinate prediction of the bounding box uses smooth-L1 loss as the loss function, and the SGD or Adam optimizer is used to optimize the network parameters. The specific training method is not limited in the embodiments of the present application.

[0101] In this embodiment of the invention, compared with traditional 3D image detection methods, this training method has the advantages of requiring fewer training samples and running faster.

[0102] In this embodiment of the invention, it should be noted that the Enface projection detection model differs from the B-Scan processing model (fast / slow axis detection model) described above. It uses several Enface images stitched together in the channel direction as the overall input.

[0103] In this embodiment of the invention, the fast-axis detection model is used to process the data in the first and second dimensions of the OCT image, and the slow-axis detection model is used to process the data in the first and third dimensions of the OCT image; the projection detection model is used to process the data in the second and third dimensions of the OCT image. In practical applications, the processing dimensions corresponding to the fast-axis B-Scan processed by the fast-axis detection model are depth and width, with an image size of L*W; the processing dimensions corresponding to the slow-axis B-Scan are depth and length, with a corresponding image size of L*M; the processing dimensions corresponding to the fast and slow-axis detection models can be interchanged, and this embodiment of the invention does not impose any limitations. The processing dimensions corresponding to the Enface projection detection model are width and length, with a corresponding image size of W*M.

[0104] In this embodiment of the invention, optionally, when an input OCT image is detected, B-Scan registration can be performed in the fast axis direction to eliminate displacement caused by eye movement during acquisition. This registration method can be rigid or non-rigid; this embodiment of the invention does not limit the method.

[0105] 102. Perform axis detection processing on all OCT images using the axis detection model to obtain the axis detection processing results corresponding to all OCT images. The axis detection processing operation includes at least target detection operation, adjacent frame detection operation based on confidence value, multi-frame fusion operation, and confidence value determination operation.

[0106] In this embodiment of the invention, the axis detection processing result includes at least one three-dimensional axis detection box and the axis confidence value corresponding to the target detection factor in each three-dimensional axis detection box.

[0107] In this embodiment of the invention, the three-dimensional axis detection box corresponding to the target detection factor is a three-dimensional axis detection box that matches the processing dimension corresponding to the fast axis detection model and / or the slow axis detection model;

[0108] When the axis detection model performing axis detection processing on all OCT images includes both fast axis detection model and slow axis detection model, the three-dimensional axis detection box corresponding to the target detection factor is the three-dimensional axis detection box obtained by performing a fusion operation on the three-dimensional axis detection boxes corresponding to the fast axis detection model and the slow axis detection model respectively.

[0109] 103. The projection detection model performs image layering and first projection processing on all OCT images to obtain the projection data of all OCT images.

[0110] 104. The projection detection model performs a second projection processing operation on the axis detection processing results and projection data to obtain the projection processing results corresponding to the projection data.

[0111] In this embodiment of the invention, the projection processing result includes a three-dimensional projection detection box and the projection confidence value corresponding to the target detection factor in the three-dimensional projection detection box.

[0112] 105. Perform a detection frame fusion operation on the projection processing result and the axis detection processing result to obtain the target detection information corresponding to the target detection factor.

[0113] In this embodiment of the invention, the target detection information includes the prediction box region where the target detection factor is located and the target confidence value of the target detection factor in the prediction box region.

[0114] In this embodiment of the invention, step 105, which involves performing a detection box fusion operation on the projection processing result and the axis detection processing result to obtain the target detection information corresponding to the target detection factor, specifically includes:

[0115] Determine the intersection area between the 3D projection detection box and the 3D axis detection box as the prediction box area;

[0116] The average of the projection confidence value corresponding to the 3D projection detection box and the axis confidence value corresponding to the 3D axis detection box is calculated to obtain the target confidence value, which is used as the confidence value corresponding to the target detection factor in the prediction box region;

[0117] The predicted bounding box region and the target confidence value are determined as the target detection information corresponding to the target detection factor.

[0118] It is evident that implementation Figure 1 The described OCT image detection method based on multi-directional image fusion can perform target detection factor detection operations on the OCT image through axis detection model and projection detection model when the target detection model is input into the OCT image. The two detection models represent at least two detection processing directions, and each detection model performs detection processing on the OCT image to obtain detection processing results (detection boxes) in at least two detection directions. This achieves multi-directional target detection, which is different from the traditional single-directional two-dimensional image detection method. This invention improves the comprehensiveness and accuracy of three-dimensional detection of OCT images. Furthermore, after obtaining the corresponding detection processing results, the multi-detection processing results are automatically fused to obtain a precise and unified fusion processing result, which improves the accuracy of the final determined three-dimensional detection result.

[0119] In an optional embodiment, the manner in which the axial detection model performs the axial detection processing operation on all the OCT images in step 102 to obtain the axial detection processing result corresponding to all the OCT images specifically includes:

[0120] The axial detection model performs a feature detection operation on all the OCT images based on the feature information corresponding to the target detection factor to obtain at least one first OCT image, and the first OCT image includes a first detection box, and the confidence of the target detection factor in the first detection box is greater than a first preset confidence threshold.

[0121] The axial detection model performs a feature detection operation on the preceding and subsequent frames of images corresponding to the first OCT image based on the first OCT image as a reference image, the coordinate interval of the region where the first detection box is located as a reference, and a preset coordinate interval error value, to obtain a plurality of second OCT images, and each second OCT image includes a second detection box, and the confidence of the target detection factor in the second detection box is greater than the first preset confidence threshold.

[0122] The axial detection model determines the maximum overlap region of the first detection box and all the second detection boxes as a two-dimensional axial detection box corresponding to the target detection factor.

[0123] The total number of the first detection box and all the second detection boxes is taken as a third dimension of the two-dimensional axial detection box, and a second preset confidence threshold is taken as a screening condition to determine a three-dimensional axial detection box corresponding to the target detection factor. The three-dimensional axial detection box includes an axial confidence value of the target detection factor in the three-dimensional axial detection box, and the axial confidence value is greater than the second preset confidence threshold.

[0124] A plane formed by the second dimension and the third dimension in any OCT image is taken as a projection plane to project the three-dimensional axial detection box to obtain a two-dimensional axial-projection region of the three-dimensional axial detection box in the projection plane, and the two-dimensional axial-projection region includes a confidence value of the target detection factor projected on the projection plane in the three-dimensional axial detection box.

[0125] The three-dimensional axial detection box and the two-dimensional axial-projection region are determined as the axial detection processing result corresponding to all the OCT images.

[0126] In this optional embodiment, the first detection box is a two-dimensional detection box, and corresponds to the processing dimension / direction of the axial detection model. For example, when the axial detection box is a fast-axis detection box and the processing image size is L*W (depth*width), the corresponding first detection box is a detection box in the L*W direction. At this time, the first detection box corresponds to a region where the target detection factor is located.

[0127] In this optional embodiment, the fast-axis direction and the fast-axis detection model are taken as examples for illustration as follows:

[0128] The target detection factor has certain continuity in three-dimensional space, so in adjacent B-Scan images, there should be the same kind of target detection factor in the same or similar position. Using this characteristic, we first perform inter-frame fusion on the detection frame in the fast-axis direction of the B-Scan. Assuming that a target detection factor is detected on the nth frame of the fast-axis B-Scan, a detection frame in which the target detection factor is located is obtained, denoted as S n , the corresponding probability (confidence) p n is greater than a predetermined threshold value, then we start the following processing:

[0129] We then check the next B-Scan (n+1), if there is a detection frame S n of the same type of target detection factor within x, y pixels of the center position x, y, and the probability p n+1 of the detection frame S n+1 is greater than a predetermined threshold value, then we continue to check the next B-Scan (n+2), until there is no detection frame of the same type of target detection factor on the next frame of the B-Scan that meets the condition.

[0130] Similarly, we also check the previous B-Scan (n-1), if there is a detection frame S n of the same type of target detection factor within x, y pixels of the center position x, y, and the probability p n-1 of the detection frame S n-1 is greater than a predetermined threshold value, then we continue to check the previous B-Scan (n-2), until there is no detection frame of the same type of target detection factor on the previous frame of the B-Scan that meets the condition.

[0131] In this optional embodiment, the processing dimension corresponding to the fast-axis detection model is L*W, and the corresponding third dimension is M. Denote the number of all detection frames of the same type of target detection factor detected above as m1. There are several methods to fuse the m1 two-dimensional detection frames into a three-dimensional detection frame. We can calculate the maximum overlapping area of the m1 frames (i.e., find the intersection of multiple detection frames) as the fused frame. We can also calculate the union of the m1 frames as the fused frame. In addition, we can add the probabilities of different frames by pixels to obtain a probability map of the frame, and then take the pixels greater than a certain threshold value on the probability map to obtain the fused frame. Assuming that the pixel size of the finally obtained frame is l1*w1, then the size of the three-dimensional frame is l1*w1*m1, and the probability confidence corresponding to the frame is the average of the confidence corresponding to the m1 frames p1.

[0132] Similarly, we can also perform bounding box fusion on the B-Scan in the slow axis direction to obtain a three-dimensional box with a size of l2*w2*m2. The processing dimension corresponding to the slow axis detection model is L*M, and the corresponding third dimension is W. The number of detection boxes of all target detection factors of the same type detected by the slow axis detection model is w2, and the pixel size of the final box is l2*m2. Therefore, the size of the three-dimensional box is l2*w2*m2, and the probability confidence corresponding to the box is the average confidence p2 of the box.

[0133] In summary, the projection of the three-dimensional detection box of the B-Scan in the fast and slow axis directions in the Enface direction (second and third dimensions) is w1*m1 (with a probability of p1) and w2*m2 (with a probability of p2). After processing all the B-Scans in the fast and slow axis directions, we obtain a target detection factor probability distribution map P lesion with a size of W*M.

[0134] In this case, the projection of the three-dimensional detection box of the slow axis detection model in the Enface direction is the target detection factor probability distribution map P lesion with a size of W*M. lesion .

[0135] In addition, it should be noted that if we want to detect k types of target detection factors, we will have k probability maps. These k probability maps will be used as auxiliary information input in the subsequent Enface image detection.

[0136] As can be seen, in the optional embodiment, the axis detection model can be used to perform target detection factor detection processing on the OCT image in the axis processing direction, including obtaining at least one two-dimensional detection box by bounding box selection of the initial region of the target detection factor, obtaining an axis three-dimensional detection box by frame-by-frame investigation of the same type region of the target detection factor, and projecting the axis three-dimensional detection box, so as to realize the detection of the accurate three-dimensional detection box corresponding to the target detection factor in the axis processing direction and improve the determination accuracy of the axis three-dimensional detection box in the axis direction.

[0137] Embodiment Two

[0138] Please refer to Figure 2 , Figure 2is another flowchart of the OCT image detection method based on multi-directional image fusion disclosed by the embodiment of the present application. Wherein, Figure 2 The OCT image detection method based on multi-directional image fusion described can be applied to an OCT image detection device based on multi-directional image fusion, and the embodiment of the present application is not limited. As shown in Figure 2 The OCT image detection method based on multi-directional image fusion can include the following operations:

[0139] 201、When it is determined that the target detection model exists input multi-frame OCT image, all OCT images are transmitted to the sub-detection model included in the target detection module, and the sub-detection model at least includes an axis detection model and a projection detection model.

[0140] 202、The axis detection model performs axis detection processing operation on all OCT images, and obtains the axis detection processing result corresponding to all OCT images, and the axis detection processing operation at least includes target detection operation, adjacent frame detection operation based on confidence value, multi-frame fusion operation and confidence value determination operation.

[0141] 203、The projection detection model determines the interlayer boundary of each frame of OCT image according to the preset image interlayer boundary determination algorithm, and then performs segmentation and layering processing on each frame of OCT image to obtain the segmentation and layering processing result of each frame of OCT image.

[0142] 204、The projection detection model performs projection processing operation on the segmentation and layering processing result of each frame of OCT image according to the preset projection processing algorithm, and obtains the projection processing result corresponding to each frame of OCT image as the projection data of the frame of OCT image.

[0143] In the embodiment of the present application, the projection processing algorithm includes average value projection or maximum / minimum value projection processing algorithm.

[0144] 205、The projection detection model performs image splicing on the projection data of OCT image with preset frame number in the channel direction corresponding to the projection detection model to obtain the input projection image of the projection feature extraction network of the projection detection model as the projection data of all OCT images.

[0145] 206、The projection detection model performs second projection processing operation on the axis detection processing result and the projection data to obtain the projection processing result corresponding to the projection data.

[0146] 207、The detection frame fusion operation is performed on the projection processing result and the axis detection processing result to obtain the target detection information corresponding to the target detection factor.

[0147] For other descriptions of steps 201-202 and steps 206-207, please refer to other specific descriptions of steps 101-102 and steps 104-105 in Embodiment One, and the present embodiment will not be described again.

[0148] It can be seen that the implementation Figure 2 The described OCT image detection method based on multi-directional image fusion can automatically perform segmentation and layering processing and channel direction splicing for each frame of image after inputting the OCT image into the projection detection model, so as to obtain the required projection data and improve the accuracy of the determined projection data. After the projection direction processing of the OCT image, the preposed axis detection processing result can be comprehensively processed, and the second projection processing and detection frame fusion are sequentially performed, so as to realize multi-directional target detection and improve the determination comprehensiveness and accuracy of the finally determined target detection information.

[0149] In an optional embodiment, the projection feature extraction network comprises a plurality of projection feature extraction layers; and the manner in which the step 206 performs the second projection processing operation on the axis detection processing result and the projection data to obtain the projection processing result corresponding to the projection data specifically comprises:

[0150] The projection feature extraction network performs a feature extraction operation on the input projection image to obtain a feature extraction result corresponding to the input projection image input into any projection feature extraction layer;

[0151] For any projection feature extraction layer, the projection feature extraction network determines a first image size corresponding to the two-dimensional axis-projection region and a second image size of the feature extraction result corresponding to the projection feature extraction layer; and according to a preset size adjustment algorithm, performs a size adjustment operation on the two-dimensional axis-projection region based on the second image size to obtain a size adjustment result corresponding to the two-dimensional axis-projection region, the size adjustment operation comprising at least one of interpolation scaling processing, convolution processing, normalization processing, and weighted value conversion processing;

[0152] The projection detection model performs a weighted processing on the feature extraction result and the size adjustment result to obtain a weighted processing result corresponding to the feature extraction result, the weighted processing result comprising a two-dimensional projection detection frame corresponding to a target detection factor, and a confidence of the target detection factor in the two-dimensional projection detection frame being greater than a third preset confidence threshold;

[0153] The projection detection model determines a three-dimensional projection detection frame corresponding to the two-dimensional projection detection frame as a projection processing result corresponding to the projection data, with the first dimension as an image expansion direction and in combination with two image layering boundary lines corresponding to the weighted processing result;

[0154] Two image layer boundary lines are two straight lines with the longest straight line distance among all layer boundary lines included in the weighted processing result.

[0155] In the optional embodiment, for the projection detection model of Enface, a plurality of Enface images are obtained according to the foregoing steps, and are input into the projection detection model as a whole after splicing in the channel direction. In addition, in the feature extraction aspect, k P lesion lesions are added in a plurality of projection feature extraction layers as auxiliary information for weighted processing of the original feature map.

[0156] In the optional embodiment, further, the format of the size adjustment result corresponding to the two-dimensional axis-projection region obtained by performing a size adjustment operation on the two-dimensional axis-projection region according to a preset size adjustment algorithm with reference to the second image size includes:

[0157] performing interpolation processing on the two-dimensional axis-projection region with reference to the second image size to obtain an interpolation processing result corresponding to the two-dimensional axis-projection region, the image size of the two-dimensional axis-projection region in the interpolation processing result being a third image size;

[0158] inputting the interpolation processing result into a preset target processing layer to obtain a target processing result with the image size being the second image size, and converting the target processing result into a weighted value in a preset value interval through a preset activation function, as the size adjustment result corresponding to the two-dimensional axis-projection region;

[0159] The target processing layer includes one or more of a plurality of preset convolution layers, normalization layers and activation layers.

[0160] In the optional embodiment, specifically, assuming that the size of P lesion is k*W*M (corresponding to the first image size), and the size of the feature map F of a certain projection feature extraction layer is C*X*Y (corresponding to the second image size), we first scale P lesion to the size of k*X*Y through interpolation, and then we convert Plesion into a feature map F att of C*W*Y through a plurality of convolution layers, normalization layers and activation layers, and then we convert the feature map into a weighted value between 0 and 1 through a Sigmoid activation function to obtain a new F att corresponding to the weighted value. We perform weighted calculation on the feature extraction result: F*(1+F att ) to obtain a weighted feature map F’. In this way, the detection result of the B-Scan can be utilized to realize that the Enface projection detection model focuses on the region where the target detection factor is located, and the accuracy of the detected target detection factor is improved.

[0161] In this optional embodiment, it is to be noted that after the two-dimensional projection detection frame (denoted as w3*m3) is determined by the projection detection model, it needs to be expanded in the depth direction. Specifically, we expand the two-dimensional projection frame according to the coordinates of the layering line at the position of the two-dimensional projection detection frame. For example, the projection data determined before includes a plurality of layering lines after performing segmentation layering, and the first and last layering lines are selected to frame the upper and lower boundaries of the retina, so that the two-dimensional projection detection frame is expanded in the depth direction with the coordinates of the first and last layering lines as the boundaries, thereby obtaining a three-dimensional projection detection frame l3*w3*m3, and the probability value of the three-dimensional projection detection frame is the probability value p3 of the original two-dimensional projection frame.

[0162] It can be seen that in this optional embodiment, when the projection detection model is used to perform projection detection processing on the input projection data, the P lesion In summary, the feature weighting processing is performed, the Enface projection detection model is implemented to focus on the region where the target detection factor is located, and the accuracy of the target detection factor, the region where the target detection factor is located, and the probability corresponding to the target detection factor is improved.

[0163] Embodiment three

[0164] Please refer to Figure 3 , Figure 3 is a structure schematic diagram of an OCT image detection device based on multi-directional image fusion disclosed by the embodiment of the present application. The device can be applied to a target detection model. The OCT image detection device based on multi-directional image fusion can be an OCT image detection terminal, equipment, system or server based on multi-directional image fusion. The server can be a local server, a remote server, or a cloud server (also known as a cloud server). When the server is a non-cloud server, the non-cloud server can be in communication connection with the cloud server, and the embodiment of the present application does not make any limitation. As shown in Figure 3 The OCT image detection device based on multi-directional image fusion can include a transmission module 301, an axis detection module 302, a first projection processing module 303, a second projection processing module 304, and a fusion processing module 305, wherein:

[0165] The transmission module 301 is configured to transmit all OCT images to a sub-detection model included in a target detection model when it is determined that the target detection model has inputted a plurality of OCT images. The sub-detection model at least includes an axis detection model and a projection detection model, and the target detection model is configured to determine a target detection factor from all OCT images.

[0166] The shaft detection module 302 is configured to perform shaft detection processing operation on all OCT images by a shaft detection model to obtain shaft detection processing results corresponding to the all OCT images, and the shaft detection processing operation at least includes target detection operation, adjacent frame detection operation based on confidence value, multi-frame fusion operation and confidence value determination operation; the shaft detection processing result includes at least one three-dimensional shaft detection frame and an axis confidence value corresponding to a target detection factor in each three-dimensional shaft detection frame.

[0167] The first projection processing module 303 is configured to perform image layering and first projection processing operation on all OCT images by a projection detection model to obtain projection data of all OCT images.

[0168] The second projection processing module 304 is configured to perform second projection processing operation on the shaft detection processing result and the projection data by the projection detection model to obtain a projection processing result corresponding to the projection data, and the projection processing result includes a three-dimensional projection detection frame and a projection confidence value corresponding to a target detection factor in the three-dimensional projection detection frame.

[0169] The fusion processing module 305 is configured to perform detection frame fusion operation on the projection processing result and the shaft detection processing result to obtain target detection information corresponding to the target detection factor, and the target detection information includes a prediction frame region where the target detection factor is located and a target confidence value of the target detection factor in the prediction frame region.

[0170] In the embodiment of the present application, the shaft detection model includes a fast-axis detection model and / or a slow-axis detection model, the fast-axis detection model is used to process data in the first dimension and the second dimension of the OCT image, and the slow-axis detection model is used to process data in the first dimension and the third dimension of the OCT image; the projection detection model is used to process data in the second dimension and the third dimension of the OCT image.

[0171] The three-dimensional shaft detection frame corresponding to the target detection factor is a three-dimensional shaft detection frame matched with the processing dimension corresponding to the fast-axis detection model and / or the slow-axis detection model.

[0172] When the shaft detection model performing the shaft detection processing operation on all OCT images includes the fast-axis detection model and the slow-axis detection model, the three-dimensional shaft detection frame corresponding to the target detection factor is a three-dimensional shaft detection frame obtained by performing fusion operation on the three-dimensional shaft detection frames corresponding to the fast-axis detection model and the slow-axis detection model respectively.

[0173] In the embodiment of the present application, the fusion processing module 305 performs detection frame fusion operation on the projection processing result and the shaft detection processing result to obtain the target detection information corresponding to the target detection factor, and the way of the fusion processing module 305 includes:

[0174] The intersection region between the three-dimensional projection detection frame and the three-dimensional shaft detection frame is determined as the prediction frame region.

[0175] The average value of the projection confidence value corresponding to the three-dimensional projection detection box and the axis confidence value corresponding to the three-dimensional axis detection box is calculated to obtain a target confidence value as the confidence value corresponding to the target detection factor in the prediction box region;

[0176] The prediction box region and the target confidence value are determined as the target detection information corresponding to the target detection factor.

[0177] It can be seen that the implementation Figure 3 The described OCT image detection device based on multi-directional image fusion can perform detection processing operations of target detection factors on OCT images through an axis detection model and a projection detection model when the target detection model inputs the OCT images. The two detection models represent at least two detection processing directions, and the two detection models respectively perform detection processing on the OCT images to obtain detection processing results (detection boxes) in at least two detection directions. Multi-directional target detection is achieved, which is different from the traditional two-dimensional image detection method in a single direction. The three-dimensional detection comprehensiveness and accuracy of the OCT images are improved in the present application. Further, after obtaining the respective detection processing results, automatic fusion processing of the multiple detection processing results is performed to obtain a precise and unified fusion processing result, thereby improving the accuracy of the final determined three-dimensional detection result.

[0178] In an optional embodiment, the axis detection module 302 performs axis detection processing operations on all OCT images by the axis detection model, and the manner specifically includes:

[0179] The axis detection model performs feature detection operations on all OCT images based on the feature information corresponding to the target detection factor to obtain at least one first OCT image. The first detection box in the first OCT image includes a target detection factor with a confidence greater than a first preset confidence threshold.

[0180] The axis detection model performs feature detection operations on the preceding and subsequent frames of images corresponding to the first OCT image based on the first OCT image as a reference image and the coordinate interval of the region where the first detection box is located as a reference, and combines a preset coordinate interval error value to obtain multiple second OCT images. Each second OCT image includes a second detection box, and the target detection factor in the second detection box has a confidence greater than the first preset confidence threshold.

[0181] The axis detection model determines the maximum overlapping region of the first detection box and all second detection boxes as a two-dimensional axis detection box corresponding to the target detection factor.

[0182] The total number of the first detection frame and all the second detection frames is taken as a third dimension of the two-dimensional axis detection frame, and a second preset signal threshold is taken as a screening condition to determine a three-dimensional axis detection frame corresponding to the target detection factor; the three-dimensional axis detection frame includes an axis confidence value of the target detection factor in the three-dimensional axis detection frame, and the axis confidence value is greater than the second preset signal threshold;

[0183] A plane composed of the second dimension and the third dimension in any OCT image is taken as a projection plane to project the three-dimensional axis detection frame to obtain a two-dimensional axis-projection area of the three-dimensional axis detection frame in the projection plane, and the two-dimensional axis-projection area includes a confidence value of the target detection factor in the three-dimensional axis detection frame projected on the projection plane;

[0184] The three-dimensional axis detection frame and the two-dimensional axis-projection area are determined as an axis detection processing result corresponding to all the OCT images.

[0185] It can be seen that in the optional embodiment, the axis detection model can be used to perform the detection processing of the target detection factor on the OCT image in the axis processing direction: obtaining at least one two-dimensional detection frame through the frame selection of the initial area where the target detection factor is located, obtaining the axis three-dimensional detection frame through the frame-by-frame investigation of the same type area of the target detection factor, and projecting the axis three-dimensional detection frame, so as to realize the accurate three-dimensional detection frame corresponding to the target detection factor in the axis processing direction, and improve the determination accuracy of the axis three-dimensional detection frame in the axis direction.

[0186] In another optional embodiment, the first projection processing module 303 performs image layering and first projection processing operations on all the OCT images by the projection detection model, and the manner of obtaining the projection data of all the OCT images specifically includes:

[0187] The projection detection model determines the interlayer boundary of each frame of OCT image according to a preset image interlayer boundary determination algorithm, and then performs segmentation and layering processing on each frame of OCT image to obtain the segmentation and layering processing result of each frame of OCT image;

[0188] The projection detection model takes the first dimension as the projection direction and performs projection processing operation on the segmentation and layering processing result of each frame of OCT image according to a preset projection processing algorithm to obtain the projection processing result corresponding to each frame of OCT image as the projection data of the frame of OCT image, and the projection processing algorithm includes the average value projection or the maximum / minimum value projection processing algorithm;

[0189] The projection detection model performs image splicing on the projection data of a preset number of frames of OCT images in the channel direction corresponding to the projection detection model to obtain an input projection image of the projection feature extraction network of the projection detection model as the projection data of all the OCT images.

[0190] It can be seen that in the optional embodiment, after the OCT image is input into the projection detection model, for each frame of image, the segmentation and layering processing, the splicing in the channel direction are automatically performed, the required projection data is obtained, and the accuracy of the determined projection data is improved; after the OCT image is processed in the projection direction, the pre-posed axis detection processing result can be integrated, the second projection processing and the detection frame fusion are sequentially performed, the multi-direction target detection is realized, and the determination comprehensiveness and the accuracy of the finally determined target detection information are improved.

[0191] In still another optional embodiment, the projection feature extraction network includes a plurality of projection feature extraction layers; the second projection processing module 304 performs a second projection processing operation on the axis detection processing result and the projection data to obtain a projection processing result corresponding to the projection data, and the manner specifically includes:

[0192] The projection feature extraction network performs a feature extraction operation on the input projection image to obtain a feature extraction result corresponding to the input projection image input into any projection feature extraction layer;

[0193] For any projection feature extraction layer, the projection feature extraction network determines a first image size corresponding to the two-dimensional axis-projection region and a second image size of the feature extraction result corresponding to the projection feature extraction layer; based on the second image size, a size adjustment operation is performed on the two-dimensional axis-projection region according to a preset size adjustment algorithm to obtain a size adjustment result corresponding to the two-dimensional axis-projection region, and the size adjustment operation includes at least one of an interpolation scaling processing, a convolution processing, a normalization processing, and a weighted value conversion processing;

[0194] The projection detection model performs a weighted processing on the feature extraction result and the size adjustment result to obtain a weighted processing result corresponding to the feature extraction result, and the weighted processing result includes a two-dimensional projection detection frame corresponding to a target detection factor, and a confidence of the target detection factor in the two-dimensional projection detection frame is greater than a third preset confidence threshold;

[0195] The projection detection model determines a three-dimensional projection detection frame corresponding to the two-dimensional projection detection frame as a projection processing result corresponding to the projection data, based on the first dimension as an image expansion direction and in combination with two image layering boundary lines corresponding to the weighted processing result;

[0196] The two image layering boundary lines are two image layering boundary lines with the longest straight line distance among all the image layering boundary lines included in the weighted processing result.

[0197] Further, in the second projection processing module 304, based on the second image size, a size adjustment operation is performed on the two-dimensional axis-projection region according to a preset size adjustment algorithm to obtain a size adjustment result corresponding to the two-dimensional axis-projection region.

[0198] The interpolation processing is performed on the two-dimensional axis-projection region based on the second image size, and an interpolation processing result corresponding to the two-dimensional axis-projection region is obtained, wherein the image size of the two-dimensional axis-projection region in the interpolation processing result is a third image size;

[0199] The interpolation processing result is input into a preset target processing layer to obtain a target processing result with the second image size, and the target processing result is converted into a weighting value in a preset value interval by using a preset activation function, and the weighting value is taken as a size adjustment result corresponding to the two-dimensional axis-projection region.

[0200] The target processing layer includes one or more of a plurality of preset convolution layers, normalization layers and activation layers.

[0201] It can be seen that, in the optional embodiment, when the projection detection model performs the projection detection processing operation on the input projection data, the P lesion The comprehensive operation is performed, the feature weighting processing is performed, the Enface projection detection model is fully focused on the region where the target detection factor is located, and the accuracy of the target detection factor, the region where the target detection factor is located and the probability corresponding to the target detection factor that is detected is improved.

[0202] Embodiment four

[0203] Please refer to Figure 4 , Figure 4 is another structure schematic diagram of the OCT image detection device based on multi-directional image fusion disclosed by the embodiment of the application. As shown in Figure 4 , the OCT image detection device based on multi-directional image fusion can include:

[0204] The memory 401 stores executable program codes.

[0205] The processor 402 is coupled with the memory 401.

[0206] The processor 402 invokes the executable program codes stored in the memory 401 to execute the steps in the OCT image detection method based on multi-directional image fusion described in the embodiment one or the embodiment two of the application.

[0207] Embodiment five

[0208] The embodiment of the application discloses a computer storage medium, which stores computer instructions. When the computer instructions are invoked, the steps in the OCT image detection method based on multi-directional image fusion described in the embodiment one or the embodiment two of the application are executed.

[0209] Embodiment six

[0210] The embodiment of the present application discloses a computer program product, which comprises a non-transitory computer storage medium storing a computer program, and the computer program is operable to make a computer execute steps in the OCT image detection method based on multi-directional image fusion described in the embodiment one or the embodiment two.

[0211] The device embodiments described above are only schematic, wherein the modules illustrated as separate components can or can not be physically separated, and the components illustrated as modules can or can not be physical modules, i.e., can be located in one place or distributed on multiple network modules. Part or all of the modules can be selected to achieve the purpose of the embodiment scheme according to actual needs. Those skilled in the art can understand and implement it without creative labor.

[0212] Through the specific description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be realized by means of software and the necessary general hardware platform, and of course, it can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of software products, and the computer software product can be stored in a computer storage medium, and the storage medium includes a read-only memory (Read-Only Memory, ROM), a random access memory (Random Access Memory, RAM), a programmable read-only memory (Programmable Read-only Memory, PROM), an erasable programmable read-only memory (Erasable Programmable Read Only Memory, EPROM), a one-time programmable read-only memory (One-time Programmable Read-Only Memory, OTPROM), an electrically erasable programmable read-only memory (Electrically-Erasable Programmable Read-Only Memory, EEPROM), a compact disc read-only memory (Compact Disc Read-Only Memory, CD-ROM) or other optical disk storage, a magnetic disk storage, a magnetic tape storage, or any other computer readable medium that can be used to carry or store data.

[0213] It should be noted that the OCT image detection method and device based on multi-directional image fusion disclosed in the embodiments of the present application are only the preferred embodiments of the present application, and are used to illustrate the technical solutions of the present application, but not to limit the present application; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalent ones; and these modifications or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. An OCT image detection method based on multi-directional image fusion, characterized in that, The method is applied to a target detection model, and the method comprises: When it is determined that the target detection model has inputted multiple frames of OCT images, all the OCT images are transmitted to a sub-detection model included in the target detection model, the sub-detection model at least comprising an axis detection model and a projection detection model, the target detection model being used to determine a target detection factor from all the OCT images; An axis detection processing operation is performed on all the OCT images by the axis detection model to obtain an axis detection processing result corresponding to all the OCT images, the axis detection processing operation at least comprising a target detection operation, a neighboring frame detection operation based on a confidence value, a multi-frame fusion operation and a confidence value determination operation; the axis detection processing result comprising at least one three-dimensional axis detection box and an axis confidence value corresponding to the target detection factor in each three-dimensional axis detection box; An image layering and first projection processing operation is performed on all the OCT images by the projection detection model to obtain projection data of all the OCT images, and a second projection processing operation is performed on the axis detection processing result and the projection data to obtain a projection processing result corresponding to the projection data, the projection processing result comprising a three-dimensional projection detection box and a projection confidence value corresponding to the target detection factor in the three-dimensional projection detection box; A detection box fusion operation is performed on the projection processing result and the axis detection processing result to obtain target detection information corresponding to the target detection factor, the target detection information comprising a prediction box region where the target detection factor is located and a target confidence value of the target detection factor in the prediction box region.

2. The multi-directional image fusion-based OCT image detection method according to claim 1, characterized in that, The axis detection model comprises a fast-axis detection model and / or a slow-axis detection model, the fast-axis detection model being used to process data in a first dimension and a second dimension of the OCT image, and the slow-axis detection model being used to process data in the first dimension and a third dimension of the OCT image; the projection detection model being used to process data in the second dimension and the third dimension of the OCT image; The three-dimensional axis detection box corresponding to the target detection factor is a three-dimensional axis detection box matching a processing dimension corresponding to the fast-axis detection model and / or the slow-axis detection model; When the axis detection model performing the axis detection processing operation on all the OCT images comprises the fast-axis detection model and the slow-axis detection model, the three-dimensional axis detection box corresponding to the target detection factor is a three-dimensional axis detection box obtained by performing a fusion operation on the three-dimensional axis detection boxes respectively corresponding to the fast-axis detection model and the slow-axis detection model.

3. The multi-directional image fusion-based OCT image detection method according to claim 2, characterized in that, The axis detection model performing the axis detection processing operation on all the OCT images to obtain the axis detection processing result corresponding to all the OCT images comprises: performing feature detection operations on all the OCT images based on the feature information corresponding to the target detection factor, to obtain at least one frame of first OCT image, the first OCT image comprising a first detection box, and a confidence value of the target detection factor in the first detection box being greater than a first preset confidence threshold; performing the feature detection operations on the preceding and subsequent frames of the first OCT image based on the first OCT image as a reference image, a coordinate interval of a region where the first detection box is located as a reference, and a preset coordinate interval error value, to obtain a plurality of frames of second OCT image, each frame of the second OCT image comprising a second detection box, and a confidence value of the target detection factor in the second detection box being greater than the first preset confidence threshold; determining, by the axis detection model, a maximum overlap region of the first detection box and all the second detection boxes as a two-dimensional axis detection box corresponding to the target detection factor; determining, based on a total number of the first detection box and all the second detection boxes as a third dimension of the two-dimensional axis detection box, a three-dimensional axis detection box corresponding to the target detection factor based on a second preset confidence threshold as a screening condition, the three-dimensional axis detection box comprising an axis confidence value of the target detection factor in the three-dimensional axis detection box, and the axis confidence value being greater than the second preset confidence threshold; projecting the three-dimensional axis detection box on a projection plane formed by a second dimension and a third dimension in any of the OCT images to obtain a two-dimensional axis-projection region of the three-dimensional axis detection box on the projection plane, the two-dimensional axis-projection region comprising a confidence value of the target detection factor in the three-dimensional axis detection box projected on the projection plane; determining the three-dimensional axis detection box and the two-dimensional axis-projection region as an axis detection processing result corresponding to all the OCT images.

4. The multi-directional image fusion-based OCT image detection method according to claim 3, characterized in that, The image layering and first projection processing operations performed by the projection detection model on all the OCT images to obtain projection data of all the OCT images include: determining, by the projection detection model, an inter-layer boundary of each of the OCT images according to a preset inter-layer boundary determination algorithm, and performing segmentation and layering processing on each of the OCT images to obtain a segmentation and layering processing result of each of the OCT images; performing, by the projection detection model, projection processing operations on the segmentation and layering processing result of each of the OCT images according to a preset projection processing algorithm to obtain a projection processing result corresponding to each of the OCT images as projection data of the frame of OCT image, the projection processing algorithm comprising an average value projection or a maximum / minimum value projection processing algorithm; performing, by the projection detection model, image splicing on the projection data of a preset number of frames of the OCT images in a channel direction corresponding to the projection detection model to obtain an input projection image for an input projection feature extraction network of the projection detection model as projection data of all the OCT images.

5. The multi-directional image fusion-based OCT image detection method according to claim 4, characterized in that, The projection feature extraction network comprises a plurality of projection feature extraction layers; The second projection processing operation is performed on the shaft detection processing result and the projection data to obtain a projection processing result corresponding to the projection data, and the second projection processing operation comprises: The projection feature extraction network performs a feature extraction operation on the input projection image to obtain a feature extraction result corresponding to the input projection image input into any projection feature extraction layer; For any projection feature extraction layer, the projection feature extraction network determines a first image size corresponding to the two-dimensional shaft-projection region and a second image size of the feature extraction result corresponding to the projection feature extraction layer; and the second image size is taken as a reference to perform a size adjustment operation on the two-dimensional shaft-projection region according to a preset size adjustment algorithm to obtain a size adjustment result corresponding to the two-dimensional shaft-projection region, and the size adjustment operation comprises at least one of an interpolation scaling processing, a convolution processing, a normalization processing, and a weighted value conversion processing. The projection detection model performs a weighted processing on the feature extraction result and the size adjustment result to obtain a weighted processing result corresponding to the feature extraction result, and the weighted processing result comprises a two-dimensional projection detection box corresponding to the target detection factor, and a confidence of the target detection factor in the two-dimensional projection detection box is greater than a third preset confidence threshold. The projection detection model determines a three-dimensional projection detection box corresponding to the two-dimensional projection detection box as a projection processing result corresponding to the projection data, taking the first dimension as an image expansion direction and combining two image hierarchical boundary lines corresponding to the weighted processing result. Of all the hierarchical boundary lines included in the weighted processing result, the two image hierarchical boundary lines are two hierarchical boundary lines with the longest straight line distance.

6. The multi-directional image fusion-based OCT image detection method according to claim 5, characterized in that, The second image size is taken as a reference to perform a size adjustment operation on the two-dimensional shaft-projection region according to a preset size adjustment algorithm to obtain a size adjustment result corresponding to the two-dimensional shaft-projection region, and the size adjustment operation comprises: The second image size is taken as a reference to perform an interpolation processing on the two-dimensional shaft-projection region to obtain an interpolation processing result corresponding to the two-dimensional shaft-projection region, and an image size of the two-dimensional shaft-projection region in the interpolation processing result is a third image size; The interpolation processing result is input into a preset target processing layer to obtain a target processing result with an image size of the second image size, and the target processing result is converted into a weighted value in a preset value interval through a preset activation function, serving as a size adjustment result corresponding to the two-dimensional shaft-projection region. The target processing layer comprises one or more of a plurality of preset convolution layers, normalization layers, and activation layers.

7. The multi-directional image fusion-based OCT image detection method according to claim 5 or 6, characterized in that, The detection box fusion operation is performed on the projection processing result and the shaft detection processing result to obtain target detection information corresponding to the target detection factor, and the detection box fusion operation comprises: An intersection region between the three-dimensional projection detection box and the three-dimensional shaft detection box is determined as a prediction box region. An average value of a projection confidence value corresponding to the three-dimensional projection bounding box and an axis confidence value corresponding to the three-dimensional axis bounding box is calculated to obtain a target confidence value as a confidence value corresponding to the target detection factor in the prediction bounding box region; The prediction bounding box region and the target confidence value are determined as target detection information corresponding to the target detection factor.

8. An OCT image detecting apparatus based on multi-directional image fusion, characterized by The device is applied to a target detection model, and the device comprises: A transmission module is configured to, when it is determined that the target detection model has input multiple frames of OCT images, transmit all the OCT images to a sub-detection model included in the target detection model, the sub-detection model at least comprising an axis detection model and a projection detection model, the target detection model being configured to determine a target detection factor from all the OCT images; An axis detection module is configured to perform an axis detection processing operation on all the OCT images by the axis detection model to obtain an axis detection processing result corresponding to all the OCT images, the axis detection processing operation at least comprising a target detection operation, a confidence value-based adjacent frame detection operation, a multi-frame fusion operation, and a confidence value determination operation; the axis detection processing result comprising at least one three-dimensional axis bounding box and an axis confidence value corresponding to the target detection factor in each three-dimensional axis bounding box; A first projection processing module is configured to perform image layering and a first projection processing operation on all the OCT images by the projection detection model to obtain projection data of all the OCT images; A second projection processing module is configured to perform a second projection processing operation on the axis detection processing result and the projection data by the projection detection model to obtain a projection processing result corresponding to the projection data, the projection processing result comprising a three-dimensional projection bounding box and a projection confidence value corresponding to the target detection factor in the three-dimensional projection bounding box; A fusion processing module is configured to perform a bounding box fusion operation on the projection processing result and the axis detection processing result to obtain target detection information corresponding to the target detection factor, the target detection information comprising a prediction bounding box region in which the target detection factor is located and a target confidence value of the target detection factor in the prediction bounding box region.

9. An OCT image detecting apparatus based on multi-directional image fusion, characterized by, The device comprises: a memory storing executable program codes; a processor coupled with the memory; the processor invokes the executable program codes stored in the memory to execute the OCT image detection method based on multi-directional image fusion according to any one of claims 1-7.

10. A computer storage medium, characterized in that, The computer storage medium stores computer instructions, which are invoked to execute the OCT image detection method based on multi-directional image fusion according to any one of claims 1-7.

Citation Information

Patent Citations

  • Method and System for Evaluating Progression of Age-Related Macular Degeneration

    US20160174830A1

  • Early Prediction Of Age Related Macular Degeneration By Image Reconstruction

    US20180084988A1