Determining target viewing angle
Through the comparison of object detection algorithm and rule sets, the target perspective of fetal screening images is automatically identified, which solves the problem of manual navigation and freezing images in fetal screening, and realizes efficient automation of scanning processes and automation of bioassays.
Patent Information
- Application Number
- CN202380083377.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-19
- Filing Date
- 2023-11-29
- Publication Date
- 2025-07-11
AI Technical Summary
In fetal screening, prior art requires manual navigation to the target viewing angle and freezing the image to perform bioassay measurements, resulting in a disruption in scanning and affecting efficiency.
An object detection algorithm is used to identify image features and compare them with a predetermined set of rules. The confidence score and weighting process are used to determine whether the image corresponds to the target viewing angle, and automatic viewing angle recognition is achieved.
Automatically determine the target viewing angle during real-time image acquisition, avoid image freezing, improve scanning process efficiency and support automation of biometric measurements.
Smart Images

Figure CN120303706A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of determining whether an image has been acquired from a target view. A particular application of the present invention relates to determining whether an ultrasound image corresponds to a standard view. Background Art
[0002] Ultrasound is the modality of choice for fetal screening because it shows fetal anatomy in sufficient detail while being cost-effective without known side effects. Fetal screening allows for the detection of abnormalities at an early gestational age, enabling the planning and execution of therapeutically appropriate interventions as needed. Most examinations are performed at a gestational age of 18 to 22 weeks with specific recommended criteria measurements.
[0003] These measurements are related to the sizes of certain bones and structures and provide insight into fetal growth. For these measurements, a standard view is required. For example, the abdominal circumference is measured in an axial plane containing the spine, stomach, and umbilical vein, and the femur length is measured using a plane containing the entire femur.
[0004] In the clinical workflow of a fetal screening examination, multiple standard views need to be captured, and for some of these views, one or more biometric measurements need to be performed. To this end, the user needs to navigate to the target plane, "freeze" the plane (i.e., store the 2D image from the incoming stream of 2D images) and perform the biometric measurement. This means that the ultrasound scan is repeatedly interrupted to perform this "freezing" and biometric measurement.
[0005] Therefore, there is a need to improve the method of obtaining a standard view (or more generally, a target view). Summary of the Invention
[0006] The present invention is defined by the independent claims. The dependent claims define advantageous embodiments.
[0007] According to an example of one aspect of the present invention, there is provided a computer-implemented method for determining whether an image corresponds to a target view, the method comprising:
[0008] Applying an object detection algorithm to the image to identify one or more features in the image;
[0009] Comparing the identified features with a predetermined set of rules for the target view, wherein the set of rules defines one or more features, and wherein comparing the identified features with the predetermined set of rules comprises:
[0010] Obtaining a confidence score for each of the identified features;
[0011] weighting one or more of the confidence scores based on features defined in the rule set corresponding to the target viewpoint; and
[0012] combining the weighted confidence scores; and
[0013] A determination is made based on the combined weighted confidence scores whether the image corresponds to a target viewing angle.
[0014] It has been recognized that identifying features in an image can provide sufficient information to determine whether the image corresponds to a target view. This is based on comparing the identified features to a set of rules. Therefore, object detection algorithms that are typically used for object detection, rather than scene / target view recognition, can be used to identify the target view.
[0015] It will be appreciated that the image can be compared to a plurality of rule sets to determine whether the image corresponds to one of a plurality of target perspectives.
[0016] The combined confidence score essentially weights the identified features based on the confidence on their identification. Thus, when the object recognition algorithm is unreliable about the identification of a feature, identified features with low confidence will be weighted lower than identified features with higher confidence scores.
[0017] Further weighting the confidence score based on the features defined in the rule set can improve the accuracy of the combined confidence score to determine whether the image belongs to the target view. For example, for features that are not defined in the rule set and / or are explicitly defined as not present in the target view, the confidence score can be negative.
[0018] This allows the target viewing angle to be automatically determined. Thus, the target viewing angle can be obtained during real-time acquisition of the image without the user having to "freeze" the real-time capture of the image to identify the target viewing angle.
[0019] Determining whether an image corresponds to a target view may be based on obtaining a confidence score for each identified feature, combining the confidence scores for the identified features defined in a rule set, and comparing the combined confidence score to a threshold score.
[0020] For example, an object detection algorithm may output each identified feature and its confidence score.
[0021] Typically, the recognition features will mean recognizing anatomical objects (e.g., organs, parts of organs, or multiple organs). However, when working in an intraoperative setting, the method can also be extended to non-anatomical structures, such as surgical instruments. For example, if imaging is used for surgical guidance, or if positioning clamps, stents, etc., the instruments to be reported may need to be present in the image. Additionally, features may include image artifacts, such as shadows (e.g., due to an ultrasound beam passing through bone), and such artifacts can render the image unusable for a given (e.g., biometric) task.
[0022] In one example, the recognition features can include recognizing at least one anatomical object.
[0023] The rule set can also define one or more features that need to be present in the target view and one or more features that do not need to be present in the target view.
[0024] Thus, the rule set defines which features are expected to be recognized in the image and / or the relationships between the recognized features in the image such that the image corresponds to the target view. Additionally, the rule set can define which features are not expected in the target view.
[0025] This further allows for the automatic determination of the target view. Thus, the target view can be obtained during the real-time acquisition of the image, without the user having to "freeze" the real-time capture of the image to identify the target view.
[0026] The method can also include determining weights for weighting one or more of the confidence scores, where determining the weights includes receiving a labeled image in the target view that includes one or more recognized features and a confidence score for each recognized feature in the labeled image, and determining weights for one or more of the recognized features in the labeled image that, when combined, maximize the combined confidence score for the labeled image.
[0027] Maximizing the combined confidence score of a labeled image known to correspond to the target view makes the weights more accurate for images of unknown views.
[0028] Of course, it should be understood that it may not always be possible to fully maximize the combined confidence score for each labeled image, and it may generally not be possible at all. Additionally, achieving the theoretically optimal set of weights that maximizes the combined confidence score of the labeled image may require an impractical amount of time and an unrealistic amount of processing resources. Therefore, maximizing the combined confidence score of the labeled image should be understood as an attempt to maximize the combined weights within a limited amount of time and using a limited amount of processing resources.
[0029] The method may further include determining or receiving a sensitivity score corresponding to the image, wherein comparing the identified features to a set of predefined rules is further based on the sensitivity score.
[0030] The sensitivity score may be determined based on information of the image or information of the scene being imaged. For example, in ultrasound imaging, an ultrasound image of a subject with a higher body mass index (BMI) may require a lower sensitivity score because some features in the set of rules may not be clearly visible (and thus have a lower confidence score) when compared to a subject with a lower BMI.
[0031] The set of rules may also define the presence of one or more present features in the image and the absence of one or more absent features in the image.
[0032] The absence of a feature in the image can be used to determine which images do not correspond to the target view, regardless of the presence of all other expected features. For example, in ultrasound imaging, if the femur is present in an abdominal ultrasound image, that abdominal image can be discarded as not being a standard view, regardless of what other features have been identified.
[0033] Determining whether the image corresponds to the target view may include determining that the image does not correspond to the target view in cases where the identified features in the image correspond to absent features in the set of rules.
[0034] The set of rules may define the relative geometry between two or more features in the image.
[0035] It has been recognized that the relative geometry between features in the target view is quite consistent. Thus, the relative geometry (e.g., size, position, etc.) can be used to confirm whether an image corresponds to the target view.
[0036] The relative geometry may include the relative position and / or relative size between two or more features in the image.
[0037] The method may further include receiving an image stream in real time and selecting an image from the image stream to apply an object detection algorithm.
[0038] For example, the latest available image from the stream may be selected. Then the object detection algorithm can be applied to a later image periodically or when needed.
[0039] The method may further include displaying an image with a bounding box and / or a confidence score output by the object detection algorithm.
[0040] In some cases, the image may be displayed without a bounding box and / or a confidence score.
[0041] The image can be an ultrasound image, and the target view can be a standard ultrasound view. However, the image can also be different types of medical images, such as but not limited to X-ray images, CT images, or MR images.
[0042] Additionally, the method can further include measuring one or more target biometric measurements based on determining that the ultrasound image corresponds to the standard view.
[0043] The present invention also provides a computer program product comprising computer program code which, when executed on a processor, causes the processor to perform the above method. The product can be a carrier such as a (non-transitory) medium, or software that can be downloaded from a server (e.g., via the Internet).
[0044] The present invention also provides a system for determining whether an image corresponds to a target view, the system comprising a processor configured to perform the method as claimed or described herein. The processor can be configured to:
[0045] Apply an object detection algorithm to the image to identify one or more features in the image;
[0046] Compare the identified features with a predefined set of rules for the target view, wherein the set of rules defines one or more features, and wherein the processor is configured to compare the identified features with the predefined set of rules by:
[0047] Obtaining a confidence score for each of the identified features among the identified features;
[0048] Weighting one or more of the confidence scores based on the features defined in the set of rules corresponding to the target view; and
[0049] Combining the weighted confidence scores; and
[0050] Determining whether the image corresponds to the target view based on the combined weighted confidence scores.
[0051] The processor can also be configured to determine or receive a sensitivity score corresponding to the image, wherein the processor is configured to compare the identified features with the set of rules based on the sensitivity score.
[0052] The set of rules can also define at least one or more features that need to be present in the target view and one or more features that do not need to be present in the target view.
[0053] The processor may further be configured to determine weights for weighting one or more of the confidence scores, wherein determining the weights includes receiving a labeled image at the target perspective that includes one or more identified features and a confidence score for each identified feature in the labeled image, and determining weights for the one or more identified features in the labeled image, the weights maximizing a combined weighted confidence score for the labeled image when combined with the confidence scores.
[0054] The set of rules may define the presence of one or more present features in the image and the absence of one or more absent features in the image.
[0055] The processor may be configured to determine whether the image corresponds to the target perspective by determining that the image does not correspond to the target perspective if an identified feature in the image corresponds to an absent feature in the set of rules.
[0056] The set of rules may define the relative geometry between two or more features in the image. The relative geometry may include the relative position between two or more features in the image. Additionally or alternatively, the relative geometry may include the relative size between two or more features in the image.
[0057] These and other aspects of the invention will be apparent from, and will be elucidated with reference to, the embodiments described hereinafter. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] For a better understanding of the present invention, and to more clearly show how the present invention may be implemented, reference will now be made, by way of example only, to the following drawings, in which:
[0059] Figure 1 A method for determining whether an image corresponds to a target perspective is shown; and
[0060] Figure 2 An ultrasound image 200 for measuring abdominal circumference is shown. DETAILED DESCRIPTION
[0061] The present invention will be described with reference to the drawings.
[0062] It should be understood that the detailed description and specific examples, while indicating exemplary embodiments of the apparatus, system, and method, are only for illustrative purposes and are not intended to limit the scope of the present invention. These and other features, aspects, and advantages of the apparatus, system, and method of the present invention will be better understood from the following description, the appended claims, and the drawings. It should be understood that the drawings are merely schematic and not drawn to scale. It should also be understood that the same reference numerals are used throughout the drawings to indicate the same or similar components.
[0063] The present invention provides a computer-implemented method for determining whether an image corresponds to a target view. The method includes applying an object detection algorithm to the image to identify one or more features in the image and comparing the identified features with a predefined set of rules for the target view. The set of rules defines one or more features. Comparing the identified features with the predefined set of rules includes: obtaining a confidence score for each of the identified features; weighting one or more of the confidence scores based on the features defined in the set of rules corresponding to the target view; and combining the weighted confidence scores. Determining whether the image corresponds to the target view is based on the combined weighted confidence scores. The present invention stems from the recognition that if automatic view recognition that can identify and store target frames during scanning is used, the ultrasound fetal screening workflow can be substantially accelerated. In addition, an automated version of biometric measurements can be invoked in a further step to also automate the measurements performed. However, those skilled in the art will recognize that the methods disclosed herein are not limited to fetal ultrasound screening or ultrasound images and can also be applied to, for example, chest X-ray imaging.
[0064] Ultimately, this can result in an uninterrupted scanning session. While the user is manipulating the ultrasound probe to ultimately cover all the target views required for the examination, in the background, the views can be analyzed, identified, stored, and fed to the appropriate biometric measurement functions.
[0065] Figure 1 A method for determining whether an image corresponds to a target view is shown. In step 102, an object detection algorithm is applied to the image to detect features within the image (e.g., objects, groups of objects, planar orientation of the image, etc.). For example, the object detection algorithm can detect the presence / absence of features in the image, the relative positions between the detected features, and / or the relative sizes of the detected features. Additionally, the object detection algorithm can provide a confidence score for the detected features, which indicates the likelihood that the detected features given by the object detection algorithm are correct.
[0066] For example, during fetal screening, the anatomical structure detection network can be invoked. An example of the anatomical structure detection network can be a YOLO network, which provides bounding boxes and confidence scores for a list of anatomical objects, and the anatomical structure detection network is trained with respect to the list of anatomical objects.
[0067] In step 104, the detected features are compared with a predefined set of rules for the target perspective. The set of rules typically defines the expected features in the target perspective and / or the expected geometric relationships between the expected features.
[0068] In the case of fetal screening, the set of rules can be derived from clinical guidelines. The set of rules can define the required anatomical content for each target perspective and the associated relationships between the required anatomical content. The rules can include mandatory features (i.e., features that must be present), optional features (i.e., features that may be present but are not required), and prohibited features (features that should not be present).
[0069] For example, for the transventricular (TV) brain plane, the posterior lateral ventricle (PLV) can be a mandatory feature, the Falx can be an optional feature, and the cerebellum can be a prohibited feature.
[0070] Additionally, if the underlying object detection algorithm can provide confidence score thresholds, the set of rules can include them. Furthermore, it can include the object size and its relative position.
[0071] Moreover, the set of rules for the target perspective can provide a method for checking the combined confidence score for that particular target perspective. Specifically, the confidence scores can be weighted based on the significance of each feature for the specific perspective. For example, mandatory features can have a greater weight than optional features, and prohibited features can have a negative weight.
[0072] Typically, the set of rules for the target perspective is predefined. To determine an optimized set of rules, any weights between the expected features can be optimized by testing the weights with pre-labeled images and maximizing, for example, the combined weighted confidence scores over all labeled images.
[0073] It should be understood that in most cases, there will be multiple sets of rules corresponding to multiple target perspectives. Thus, an image can be compared with multiple sets of rules to determine whether the image corresponds to any of the target perspectives, and if so, which target perspective the image corresponds to.
[0074] In step 106, the detected features are compared with a set of rules to determine whether the image corresponds to a target view. This comparison can include ensuring the presence of all required / mandatory features. This comparison can also include ensuring the absence of all prohibited features. This comparison can also include combining all the confidence scores of the detected features in the image, where each confidence can be weighted according to the set of rules with which it is compared.
[0075] For example, in the case of real-time ultrasound screening, for each new frame (or every second frame, third frame, etc., depending on the frame rate and / or processing power), the detected anatomical structures are compared with a set of rules. If all the required rules are met for one of the target views, the frame can be automatically processed according to the type of target view.
[0076] Different types of automatic processing can be implemented on the frame. These range from merely making the user aware of the recognition (e.g., displaying a corresponding icon), asking the user for confirmation before further or automatically performing any relevant measurements, and optionally filling in the examination report.
[0077] Thus, the method essentially provides automatic view recognition for images. This is based on the recognition that view recognition can be achieved by using an object detection algorithm that is typically used for object recognition / detection and cannot itself be used for view recognition.
[0078] The following exemplary view types are typical criteria (target views) in a fetal screening examination:
[0079]
[0080] Each feature provided in the "Visible Anatomical Structure" column can be a mandatory or optional feature in the set of rules for the corresponding target view. The "Biometric Measurement" column presents the measurements typically performed from the corresponding target view.
[0081] Figure 2 An ultrasound image 200 for measuring the abdominal circumference is shown. The ultrasound image 200 is covered with four bounding boxes 202, 204, 206, and 208, each corresponding to a feature detected by an object detection algorithm in the ultrasound image 200. In this case, the object detection algorithm is the yolo network. The expression "yolo" means: you only look once.
[0082] The first maximum bounding box 202 corresponds to the abdominal box and has a confidence score of 91.05%. The second bounding box 204 corresponds to the stomach and has a confidence score of 66.45%. The third bounding box 206 corresponds to the spine and has a confidence score of 61.58%. The fourth bounding box 208 corresponds to the umbilical vein and has a confidence score of 30.76%.
[0083] The set of rules for the abdominal circumference plane may require that the stomach, spine, and umbilical vein all be present and within the abdominal box. As Figure 2 shown, this is the case in the ultrasound image 200. Thus, the rules can specify that, for example, the confidence scores of the spine, stomach, and umbilical vein are combined to determine a combined confidence score. The combined confidence score can then be compared with a confidence threshold for the abdominal circumference plane to determine whether the ultrasound image 202 is in the abdominal circumference plane.
[0084] In some cases, the specific nature of the environment in which the image is taken may affect the confidence score. For example, in ultrasound imaging generally, subjects with a high body mass index (BMI) are typically not easily imaged, and the resulting ultrasound frames may not generally provide high confidence scores. In these cases, the sensitivity score can be reduced. The reduction of the sensitivity score can relax the requirements for the set of rules (e.g., it can lower the confidence score threshold or reduce the number of mandatory features).
[0085] A person skilled in the art will be able to easily develop a processor for performing any of the methods described herein. Thus, each step of the flowchart can represent a different action performed by the processor and can be executed by the corresponding module of the processing processor.
[0086] The set of rules for identifying the target image can be manually set by the designer of the system (e.g., a clinician). However, the rules themselves can be optimized using a machine learning algorithm by training the machine learning algorithm using standard perspectives and perspectives that are not standard. The system can indicate to the user the rules that have been applied to identify different target perspectives.
[0087] The user can have the option to adapt the rules, for example, if the ultrasound examination does not find the target perspective based on the existing set of rules, to relax the rules (e.g., only a lower confidence score is required). As described above, this may be required, for example, in cases that are difficult to image, such as obese patients.
[0088] As described above, the system utilizes a processor to perform data processing. The processor can be implemented in various ways using software and / or hardware to perform the various required functions. The processor typically employs one or more microprocessors, which can be programmed using software (e.g., microcode) to perform the required functions. The processor can be implemented as a combination of dedicated hardware for performing some functions and one or more programmed microprocessors and associated circuitry for performing other functions.
[0089] Examples of circuits that can be employed in various embodiments of the present disclosure include, but are not limited to, conventional microprocessors, application specific integrated circuits (ASICs), and field programmable gate arrays (FPGAs).
[0090] In various embodiments, the processor can be associated with one or more storage media, such as volatile and non-volatile computer memories, such as RAM, PROM, EPROM, and EEPROM. The storage media can be encoded with one or more programs that, when executed on one or more processors and / or controllers, perform the required functions. The various storage media can be fixed within the processor or controller, or can be portable, such that the one or more programs stored thereon can be loaded into the processor.
[0091] By studying the drawings, the disclosure, and the appended claims, those skilled in the art can understand and implement variations of the disclosed embodiments when practicing the claimed invention. In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality.
[0092] A single processor or other unit can implement the functions recited in several of the claims.
[0093] The measures recited in mutually different dependent claims can be advantageously combined.
[0094] A computer program can be stored / distributed on a suitable medium, such as an optical storage medium or a solid-state medium supplied together with or as part of other hardware, but can also be distributed in other forms, such as via the Internet or other wired or wireless telecommunication systems.
[0095] If the word "adapted" is used in the claims or the specification, it should be noted that the word "adapted" is intended to be equivalent to the word "configured to".
[0096] Any reference signs in the claims should not be construed as limiting the scope.
Claims
1. A computer-implemented method for determining whether an image (200) corresponds to a target perspective, the method comprising: Applying (102) an object detection algorithm to the image to identify one or more features in the image; Comparing (104) the identified features with a predefined set of rules for the target perspective, wherein the set of rules defines one or more features, and wherein comparing the identified features with the predefined set of rules includes: Obtaining a confidence score for each of the identified features among the identified features; Weighting one or more of the confidence scores based on features defined in the set of rules corresponding to the target perspective; and Combining the weighted confidence scores; and Determining (106) whether the image corresponds to the target perspective based on the combined weighted confidence scores.
2. The method according to claim 1, wherein The set of rules further defines at least one or more features that need to be present in the target perspective and one or more features that do not need to be present in the target perspective.
3. The method according to claim 1 or 2, further comprising determining weights for weighting one or more of the confidence scores, wherein, Determining the weights includes: Receiving a labeled image at the target perspective including one or more identified features and a confidence score for each of the identified features in the labeled image; and Determining weights for the one or more identified features in the labeled image, the weights maximizing the combined confidence score for the labeled image when combined.
4. The method according to any one of claims 1 to 3, further comprising determining or receiving a sensitivity score corresponding to the image, wherein, Comparing the identified features with the predefined set of rules is also based on the sensitivity score.
5. The method according to any one of claims 1 to 4, wherein The set of rules defines the presence of one or more present features in the image and the absence of one or more absent features in the image.
6. The method according to claim 5, wherein Determining whether the image corresponds to the target perspective includes determining that the image does not correspond to the target perspective in the case where an identified feature in the image corresponds to an absent feature in the set of rules.
7. The method according to any one of claims 1 to 6, wherein, The set of rules defines the relative geometry between two or more features in the image.
8. The method according to claim 7, wherein The relative geometry includes the relative position and / or relative size between two or more features in the image.
9. The method according to any one of claims 1-8, further comprising receiving an image stream in real time and selecting an image from the image stream to apply the object detection algorithm.
10. The method according to any one of claims 1 to 9, further comprising displaying the image with bounding boxes (202, 204, 206, 208) and / or the confidence scores, the bounding boxes being output by the object detection algorithm.
11. The method according to any one of claims 1 to 10, wherein The image is an ultrasound image, and the target perspective is a standard ultrasound perspective.
12. A computer program product comprising computer program code which, when executed on a processor, causes the processor to perform the method according to any one of the preceding claims.
13. A system for determining whether an image corresponds to a target perspective, the system comprising a processor configured to perform the method according to any one of claims 1-11.