System and method for reconstructing 3D medical representations based on screen images

By extracting medical images from screen images and reconstructing 3D representations using machine learning models, the problem that patients cannot directly analyze medical images is solved, and autonomous diagnosis and report generation is achieved, which improves diagnostic efficiency and accuracy.

CN120412930APending Publication Date: 2025-08-01SHANGHAI UNITED IMAGING INTELLIGENCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510484557.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-12-18
Filing Date
2025-04-17
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

Patients lack the ability to directly analyze and diagnose abnormalities when viewing medical images, and the prior art is difficult to effectively use screen images to reconstruct 3D medical representations.

Method used

By extracting medical images from multiple screen images, a 3D representation of the anatomy is reconstructed using machine learning models, detecting abnormalities and providing diagnostic indications.

Benefits of technology

It enables patients to independently analyze medical images, identify abnormalities and generate diagnostic reports, improving the accuracy and efficiency of diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120412930A_ABST
    Figure CN120412930A_ABST
Patent Text Reader

Abstract

Systems, methods, and tools related to reconstructing a 3D representation of an anatomical structure based on screen recording of a display when displaying a medical image of the anatomical structure on the display are disclosed herein. A plurality of medical images of the anatomical structure may be extracted from the screen record, and one or more parameters may be determined for reconstructing a 3D representation of the anatomical structure based on the extracted medical images. A 3D representation of the anatomical structure may then be reconstructed based on the extracted medical image and the one or more determined parameters. From the 3D representation of the anatomical structure, anomalies may be detected, and medical reports may be performed using a pre-trained machine learning model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of medical imaging, and more particularly to the modeling of the human body. Background Art

[0002] During a clinical visit or via an online portal provided by a medical facility, medical images, such as two-dimensional (2D) or three-dimensional (3D) medical scans of a patient's anatomy, are often presented to the patient via a display (e.g., a computer screen). Although patients can view medical images on a display, they may not have a means to directly access those images, and if patients could access the images, their ability to analyze the images and obtain alerts, indications, or diagnoses regarding abnormalities in those anatomical structures would be much smaller. Summary of the Invention

[0003] Disclosed herein are systems, methods, and means associated with reconstructing a 3D representation of an anatomical structure (e.g., a human organ) based on multiple screen images that can be obtained via screen recording. According to various embodiments of the present disclosure, a device may be configured to obtain multiple screen images of a display, including screenshots or photographs (e.g., included in a video recording of the display), which may be captured (including possible photographing or screenshotting) when one or more medical representations of the anatomical structure (e.g., one or more medical scan images) are shown on the display. The device may be further configured to extract multiple medical images of the anatomical structure from the multiple screen images, and determine one or more parameters for reconstructing a three-dimensional (3D) representation of the anatomical structure based on the multiple extracted medical images. Then, the device may use the extracted medical images and the one or more determined parameters to reconstruct a 3D representation of the anatomical structure.

[0004] In an example, multiple screen images may be captured from the display of a desktop computer, laptop computer, tablet computer, or mobile phone. In an example, the device may receive multiple screen images from a tablet computer or a mobile phone. In an example, the device may be a tablet computer or a mobile phone itself.

[0005] In an instance, the device may be further configured to detect an abnormality associated with the anatomical structure based on one or more pre-trained machine learning (ML) models and the 3D representation of the anatomical structure, and provide an indication of the abnormality on the 3D representation of the anatomical structure. In an example, the device may be further configured to generate a report associated with the anatomical structure based on one or more pre-trained ML models and the 3D representation of the anatomical structure.

[0006] In an example, the one or more determined parameters may include the distance between two medical images in the extracted medical images and / or the voxel size of the 3D representation. In an example, the device may use a pre-trained ML model to predict the distance between two extracted medical images and determine the voxel size of the 3D representation based on the distance. In an example, a device configured to reconstruct a 3D representation of an anatomical structure based on the extracted medical images may include a device configured to identify one or more duplicate data in the extracted medical images and exclude the one or more duplicate data from the reconstruction of the 3D representation. In an example, a device configured to reconstruct a 3D representation of an anatomical structure based on the extracted medical images may include a device configured to adjust the size, orientation, or aspect ratio of at least one of the extracted medical images. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] A more detailed understanding of the examples disclosed herein can be obtained from the following description given by way of example in conjunction with the accompanying drawings.

[0008] Figure 1 is a simplified block diagram showing an example of reconstructing a 3D representation of an anatomical structure based on multiple screen images of a display.

[0009] Figure 2 is a simplified block diagram showing an example of extracting medical images from multiple screen images.

[0010] Figure 3 is a simplified block diagram showing an example of determining one or more 3D reconstruction parameters based on medical images extracted from multiple screen images.

[0011] Figure 4 is a flowchart showing an example operation associated with reconstructing a 3D representation of an anatomical structure based on a set of screen images.

[0012] Figure 5 is a flowchart showing an example operation associated with training an artificial neural network to perform one or more tasks described in embodiments of the present disclosure.

[0013] Figure 6 is a simplified block diagram showing an example device that can be configured to perform one or more tasks described in embodiments of the present disclosure. DETAILED DESCRIPTION

[0014] The present disclosure is shown in the drawings by way of example and not limitation. A detailed description of illustrative embodiments will be provided with reference to these drawings. Although the embodiments may be described with certain details, it should be noted that the details are not intended to limit the scope of the present disclosure.

[0015] Figure 1Shows an example of reconstructing a 3D representation of an anatomical structure (e.g., a human heart) based on multiple screen images of a display. As Figure 1 shown, the display (e.g., Figure 1 102) can be a monitor or a screen, such as a monitor of a desktop computer or a laptop computer or a screen of a tablet computer or a mobile phone. The multiple screen images 104 can be obtained when a medical representation of the anatomical structure (e.g., a 2D or 3D computed tomography (CT) scan, a magnetic resonance imaging (MRI) scan, or an ultrasound scan) is presented on the display. For example, the multiple screen images 104 can be a part of a screen recording (e.g., a video) of a computer taken by a patient using a mobile device (e.g., a smart phone or a tablet computer) when the medical scan images (e.g., CT, MRI, etc.) of the anatomical structure are shown to the patient using the computer during a patient's visit to a doctor's office. As another example, the multiple screen images 104 can be computer-recorded or screenshot when a person is browsing medical scan images of the anatomical structure on a computer (e.g., via a patient portal accessible through a web browser installed on the computer). Although the examples provided in the present invention may regard the multiple screen images 104 as part of a video recording, those skilled in the art will understand that the multiple screen images 104 can also be individually taken as, for example, photos rather than video recordings.

[0016] According to this embodiment, the multiple screen images 104 can be processed to extract multiple medical images 106 (e.g., which may also be referred to herein as slices) of the anatomical structure based on the multiple screen images. The processing can be performed by a computing device including, for example, a mobile device or a server. For example, the processing can be performed by the mobile device (e.g., a tablet computer or a mobile phone) used to capture the multiple screen images 104. The processing can also be performed by a server device (e.g., in a computing cloud) that can receive the multiple screen images 104 from the mobile device used to capture the screen images. The processing can also be performed by the computer on which the multiple screen images 104 are recorded. As will be described in more detail below, the processing can include identifying duplicate medical images from the multiple screen images 104 and excluding those duplicate medical images from the multiple medical images 106. The processing can also include adjusting the size, orientation, and / or aspect ratio of at least a subset of the multiple medical images 106 (e.g., such that the multiple medical images can be aligned for subsequent processing).

[0017] Because multiple medical images 106 are extracted from multiple screen images 104 of the display 102, the multiple medical images 106 can correspond to those medical images shown on the display 102 when the multiple screen images 104 are captured. Once obtained, the multiple medical images 106 can be used, together with one or more determined parameters 108, to reconstruct a 3D representation 110 of the anatomical structure. The one or more parameters 108 can include, for example, slice thickness, pixel pitch, physical size covered by the slice, voxel size of the 3D representation, distance between two (e.g., any two) of the extracted medical images 106, and / or similar parameters. As will be described in more detail below, the one or more parameters 108 can be determined in different ways, including, for example, extracting the parameters from the multiple screen images 104 (e.g., via optical character recognition (OCR)) or using one or more pre-trained machine learning (ML) models to predict the parameters. It can also be identified not through screen images (which can also be called photos or screenshots), but known from other sources, such as doctor instructions.

[0018] The 3D representation of the anatomical structure 110 reconstructed based on the extracted medical images 106 and the one or more determined parameters 108 can correspond to the representation shown on the display 102 (e.g., 2D or 3D medical scan), and can be used for diagnostic purposes, treatment planning, and / or surgical navigation. For example, the 3D representation 110 can be used to detect anomalies associated with the anatomical structure and provide an indication 112 (e.g., bounding box, segmentation mask, etc.) of the anomaly (e.g., on the 3D representation 110). As another example, the 3D representation 110 can be used to generate a diagnostic report associated with the anatomical structure based on features extracted from the 3D representation. As yet another example, the 3D representation 110 can be used to generate a treatment plan associated with the anatomical structure based on the features of the extracted 3D representation and / or the medical history of the relevant patient. As described in more detail below, one or more of these tasks can be completed using pre-trained ML models.

[0019] Figure 2 An example of extracting medical images from multiple screen images is shown. As described above, a mobile device can be used to take screen pictures or capture screenshots (e.g., Figure 2 204) while a medical representation (e.g., 2D or 3D medical scan) of the anatomical structure is being displayed on the display device. Due to the nature of such screen captures, the images 204 may not be directly suitable for reconstructing a 3D representation of the anatomical structure (e.g., the images 204 can include duplicate images, images of different sizes or aspect ratios, low-quality images, etc.). Therefore, the multiple screen images 204 can be processed at 202 to extract multiple eligible medical images 206 that can be used for 3D reconstruction, and as part of the extraction process, one or more of the following can be performed.

[0020] Operations at 202 may include image preprocessing. For example, from an obtained screen image containing a medical image (e.g., a 2D medical image), four corners of the medical image may be determined (e.g., using a machine learning model trained to detect visual features associated with the corners), and a bounding box may be derived based on the four corners and used to crop the medical image from the screenshot. In this way, only the medical image can be extracted from the screen image, while the rest of the screen image (e.g., unrelated GUI components displayed on the screen) can be ignored. As another example, normalization may be applied to the medical images extracted from screen image 204 to ensure they have consistent illumination and / or color correction to reduce differences. As yet another example, resizing, upsampling, or downsampling may be performed to speed up processing while maintaining sufficient detail. As yet another example, one or more filters (e.g., Gaussian filter, median filter, etc.) may be applied to reduce noise in the image and improve image quality.

[0021] Operations at 202 may include aligning the medical images extracted from screen image 204. Alignment may involve adjusting the geometric properties (e.g., size, aspect ratio, etc.) of the extracted medical images, and / or translating / rotating them to match corresponding points or features across multiple images to establish their relative positions and / or orientations. For example, feature detection algorithms (e.g., SIFT (Scale-Invariant Feature Transform), SURF (Speeded-Up Robust Features), and / or ORB (Oriented FAST and Rotated BRIEF)) may be used to identify significant points of interest (e.g., key points) in each of the extracted images. For each detected key point, a descriptor may be computed to represent the features (e.g., local appearance) around the key point (e.g., the descriptor may contain a vector representing the unique characteristics of the feature). Descriptors between the images may then be matched to find corresponding points, and the closest match may be determined based on one or more distance metrics (e.g., Euclidean distance).

[0022] In some instances, due to noise, occlusion, repetitive textures, etc., not all of the matched key points may be correct, so geometric verification may be performed to improve the alignment of the images. For example, a filter matching technique based on RANSAC (Random Sample Consensus) may be used to estimate a robust transformation (e.g., a transformation matrix) that can align the images while discarding outliers. In some instances, if the extracted medical images are not yet aligned on the same plane, correction may be performed to align the corresponding points. This may be done, for example, in the case where two screen images are captured from slightly different viewpoints.

[0023] Operations at 202 may include detecting and excluding duplicate medical images extracted from the screen image 204 (e.g., including medical images that are not exactly the same but are substantially similar). This can be achieved by comparing the extracted medical images based on the visual content of the medical images (e.g., rather than their file names or metadata). For example, a hashing method can be used to identify duplicate medical images, by which a hash representing the visual content of the extracted medical image can be calculated, and duplicate medical images can be detected as having similar hashes. As another example, duplicate medical images can be identified using a feature-based method, by which key points and / or descriptors in the image can be identified and compared to find similar or duplicate images. As yet another example, a deep learning-based method that can utilize an artificial neural network (such as, for example, a convolutional neural network CNN) can be used to identify duplicate medical images. The neural network can be pre-trained to extract deep features from medical images, and then the features can be compared using one or more similarity measures (e.g., cosine similarity, Euclidean distance, etc.) to detect duplicates in the medical images.

[0024] As Figure 1 shown, certain parameters 108 may be required to reconstruct a 3D representation of the anatomical structure (e.g., based on the set of medical images 206 extracted from Figure 2 the multiple screen images 204 shown). The parameters may include, for example, the number of medical images included in the extracted image set, the distance between consecutive images or slices in the extracted image set, the physical size represented by the voxels in the 3D reconstruction, etc. In some examples, one or more parameters can be obtained based on an external input (e.g., a user can provide one or more parameters), while in other examples, one or more parameters can be determined based on a set of medical images 206 or screen images 204 (e.g., by extracting information from the images or screen images via OCR, or by using a pre-trained machine learning model to predict one or more parameters).

[0025] Figure 3 shows a method for determining one or more 3D reconstruction parameters (e.g., Figure 3 306) based on medical images (e.g., Figure 3Example process of (308). Parameter determination can be performed at 302 and different techniques can be used, including conventional or deep learning-based techniques. Using the determination of the voxel size of 3D reconstruction as an example, such parameters can be estimated based on an image 306 (e.g., a 2D image), such as the physical size of the image, the image resolution (e.g., the number of pixels in a 2D image represented by width × height), and / or the slice thickness that can be represented by the distance between consecutive 2D images or segments. This estimation can involve, for example, determining the pixel size in a physical unit and further determining the voxel size in the in-plane dimensions (x, y) and the z dimension (e.g., the z dimension measurement can be determined based on the above slice thickness).

[0026] In an example, a pre-trained ML model (such as a deep learning model or a neural network model) can be used to predict the voxel size. Such an ML model can be implemented using different neural network architectures (including, for example, CNN, ResNet, etc.). The ML model can be trained through a training process that can involve providing multiple 2D medical images to the ML model, having the ML model make predictions about the voxel size, comparing the predicted voxel size with the corresponding gold standard, and adjusting the parameters of the ML model based on the loss between the predicted voxel size and the gold standard.

[0027] Although the parameter determination at 302 is described using the voxel size as an example, those skilled in the art will recognize that similar techniques can be applied to derive other parameters (e.g., the distance between consecutive images or slices) that may be required for 3D reconstruction.

[0028] Using one or more parameters estimated according to Figure 3 and a series of medical images (e.g., CT or MRI images) extracted from a screen recording (e.g., as Figure 2 shown), 3D medical image reconstruction can be performed to create a 3D representation of an anatomical structure (e.g., the anatomical structure of a patient), as Figure 1As shown in 110. One or more of the following operations can be performed as part of a 3D reconstruction process. Other operations described above (e.g., image resizing / normalization, image registration or alignment, image denoising, etc.) can also be performed as part of the 3D reconstruction process. For example, the 3D reconstruction process can include identifying regions of interest (ROIs) in a series of input images. This can be achieved using various segmentation techniques, such as thresholding, region growing, or deep learning methods, by which the boundaries of organs, tissues, or other structures within each input image can be delineated. Then a 3D mesh (e.g., a voxel mesh) can be created to represent the volume including the input images or slices. The mesh can be defined, for example, based on the dimensions estimated from the 2D images and the voxel size (e.g., as discussed previously). Then, for example, one or more interpolation techniques (e.g., linear or cubic spline) can be used to estimate the density values of the voxels between consecutive 2D slices. As a result of the interpolation, a continuous 3D volume can be created, and various volume rendering techniques such as ray casting, maximum intensity projection (MIP), or surface rendering can be used to visualize the continuous 3D volume.

[0029] In an example, the 3D reconstruction can also include one or more post-processing steps to enhance or improve the reconstructed 3D representation. For example, one or more smoothing filters can be applied to reduce any artifacts from the reconstruction and enhance important features such as the edges of anatomical structures. Additionally, anatomical knowledge or machine learning-based correction methods can be used to adjust for errors in segmentation or interpolation.

[0030] For diagnostic purposes, treatment planning, and / or surgical navigation, 3D images of anatomical structures reconstructed using the techniques described herein (e.g., Figure 1 of 110) can be analyzed. The analysis can be performed using conventional methods (such as by performing volume measurements, distance calculations, morphological studies, or edge detection on the reconstructed 3D data). Machine learning (e.g., deep learning)-based techniques can also be used to perform the analysis, such as by training a machine learning model to detect anomalies in anatomical structures and generate an indication, report, or treatment plan associated with the detected anomalies (e.g., as shown in Figure 1 of 112).

[0031] Machine learning (ML) models for anomaly detection and / or report generation can include CNN-based ML models, transformer-based ML models, large language models (LLMs), and / or other types of deep learning models. The ML models can be trained on actual patient medical images (e.g., 2D or 3D medical scans) to learn and extract features from those medical images and identify features that may be associated with anomalies. As will be described in more detail below, the ML models can be trained based on a large amount of clinical data, and more specifically, by splitting the data into training, validation, and test sets, training the model on the training set and validating it (e.g., the ML model) on the validation set. Using appropriate loss functions, such as binary cross-entropy, wafer loss, etc.).

[0032] In an example, anomaly detection and / or report generation can be done using a reconstructed 3D representation in combination with other types of information (e.g., the data used to train the ML model and generate the diagnosis can be multimodal). These other types of information can include, for example, textual information such as descriptions of symptoms experienced by the patient, the patient's laboratory reports, the patient's medical history, etc. The information can also include other medical images of the patient, such as the patient's MRI or CT images. The information can further include audio information such as recordings of conversations between the patient and the doctor, descriptions of the patient's own narrative of their health condition, etc.

[0033] The ML models can be trained using multimodal patient data and, once trained, deployed to generate outputs based on multimodal patient data including the reconstructed 3D representation (e.g., Figure 1 of 112). One or more of the ML models can be implemented using an artificial neural network that can include multiple encoders and decoders. Each of the encoders can be configured to receive a corresponding type of patient data and produce an encoded representation of the type of patient data (e.g., in the form of one or more vectors). The decoder can be configured to receive an encoded representation of the multimodal patient data (e.g., a concatenation of the encoded representations) and predict an output based on the encoded representation and / or a query (e.g., a question posed by the patient). As described herein, the predicted output can include medical decisions such as medical procedures recommended for the patient (e.g., an MRI or CT scan), an indication of whether a tumor region has been detected in the 3D representation, etc. The predicted output can also include a medical summary of the patient's health condition (e.g., a text summary) generated based on the encoded representation.

[0034] In an example, one or more encoders in the encoder can include a CNN, which includes one or more convolutional layers, one or more pooling layers, and / or one or more fully connected layers. Each convolutional layer can include a plurality of kernels or filters with respective weights, and the weights can be configured to extract features from an input (e.g., a text input or an image-based input). After the convolutional operation, batch normalization and / or an activation function (e.g., such as a rectified linear unit (ReLu) activation function) can be performed, and the features extracted by the convolutional layer can be downsampled via one or more pooling layers and / or fully connected layers to obtain a representation of the extracted features, e.g., in the form of a feature vector. In an example, the network can adopt a recursive architecture to store a hidden state associated with the input and feed the hidden state back into the convolutional layer of the encoder (e.g., via one or more recursive connections). In this way, the encoder can utilize not only the current set of data samples passing through the network during feature encoding but also the previous data samples represented by the hidden state to derive a more accurate representation of the input data.

[0035] In an example, one or more encoders in the encoder can include a Transformer neural network having a built-in attention (e.g., self-attention) mechanism (e.g., including one or more self-attention layers), which is configured to detect relationships between different parts of an input data sequence and learn the context of the input data (and thus learn its meaning). These tasks can be implemented, for example, based on query, key, and value vectors or matrices.

[0036] In an example, the ML model described herein can include a vision-language model, which can be trained to learn a mapping between visual embeddings and text embeddings (e.g., between visual features and text features) from a dataset including paired images and text descriptions. Training data can be obtained from different sources, including, for example, the Internet (e.g., websites that can include images and descriptions of image content), publicly accessible databases (e.g., figures and captions from repositories of academic publications), hospital records (e.g., radiology reports), etc. The training data can be preprocessed, for example, to ensure its suitability for training. For example, preprocessing can include resizing images, tokenizing text, creating image-text input pairs, etc. Preprocessing can also include augmenting the training data (e.g., by varying the text descriptions to increase the diversity of the training dataset) to improve the robustness and accuracy of the vision-language model.

[0037] A vision-language model may include a vision encoding part (e.g., implemented via a vision encoder) and a text encoding part (e.g., implemented via a text encoder). In an example, the vision encoder may utilize a vision transformer architecture designed to extract image features from an input image, while the text encoder may be implemented using a conventional transformer architecture designed to extract text features from a text description. The image features and text features may then be aligned (e.g., mapped to each other) in a joint embedding space (e.g., via concatenation or some other suitable fusion technique) to capture the relationship between visual information and text information. In an example, the vision encoder and text encoder may first be trained (e.g., separately) on a large number of images and text descriptions respectively, and then fine-tuned using an application-specific dataset (e.g., a certain type of medical scan image) and / or based on a specific downstream task (e.g., medical image classification).

[0038] This training may allow the vision-language model to acquire an understanding of the relationships between certain visual embeddings or features, such that when given an image (e.g., Figure 1 a 3D representation 110) as input, the vision-language model can extract visual features from the input and generate a coherent and informative interpretation (e.g., a diagnostic report) of the visual information contained in the input by relating the extracted visual features to the corresponding text features in the learned joint embedding space.

[0039] Figure 4 An example program 400 is shown that may include one or more of the operations described herein. Process 400 may be performed independently or collaboratively by different devices. For example, process 400 may be performed by a server (e.g., in a computing cloud) configured to receive a screen recording of a display from another device (e.g., a mobile device). Program 400 may also be performed by a device (e.g., a mobile device or a desktop computer) for capturing screen images.

[0040] As Figure 4 shown, program 400 may include obtaining a plurality of screen images of a display at 402, where the plurality of screen images may be captured while a medical representation of an anatomical structure is shown on the display. Program 400 may further include extracting a plurality of medical images (e.g., 2D medical images) of the anatomical structure from the plurality of screen images at 404, and determining one or more parameters for reconstructing a 3D representation of the anatomical structure using the plurality of extracted medical images at 406. Program 400 may also include reconstructing a 3D representation of the anatomical structure at 408 based on the plurality of medical images extracted from the screen images and the one or more determined parameters.

[0041] Figure 5Shows an example operation 500, which may be associated with training an artificial neural network (e.g., the artificial neural network may be configured to implement one or more of the ML models described herein) to perform one or more of the tasks described herein. As Figure 5 shown, the training operation 500 may include an operation of initializing the operation parameters (e.g., weights associated with the respective layers of the neural network) of the neural network at 502, for example, by sampling from a probability distribution or by copying the parameters of another neural network having a similar structure. The training operation may further include providing an input (e.g., a reconstructed 3D medical image) to the neural network at 504, and causing the neural network to make a prediction (e.g., regarding a classification label, a segmentation mask, etc.) using the currently assigned network parameters at 506. At 508, the training operation may include determining a loss associated with the prediction, for example, based on the difference between the prediction and the corresponding ground truth. At 510, the training operation may further include determining whether one or more training termination criteria have been met. For example, if the difference between the prediction and the ground truth falls below a predetermined threshold, it may be determined that the training termination criteria have been met. If the determination at 510 is that the training termination criteria are met, the training may end. Otherwise, the currently assigned network parameters may be adjusted at 512, for example, by gradient descent of the loss backpropagated through the network before the training returns to 506.

[0042] For simplicity of explanation, the training operations are depicted and described herein in a particular order. However, it should be understood that the training operations may occur in a different order, simultaneously, and / or together with other operations not presented or described herein. Additionally, it should be noted that not all operations depicted and described herein may be included in the training process, and not all of the shown operations need to be performed.

[0043] The systems, methods, and / or means described herein may be implemented using one or more processors, one or more storage devices, and / or other suitable accessory devices (such as display devices, communication devices, input / output devices, etc.). Figure 6is a block diagram showing an exemplary device 600 that can be configured to perform the tasks described herein. As shown, device 600 may include a processor (e.g., one or more processors) 602, which can be a central processing unit (CPU), a graphics processing unit (GPU), a microcontroller, a reduced instruction set computer (RISC) processor, an application specific integrated circuit (ASIC), an application specific instruction set processor (ASIP), a physics processing unit (PPU), a digital signal processor (DSP), a field programmable gate array (FPGA), or any other circuit or processor capable of performing the functions described herein. Device 600 may further include a communication circuit 604, a memory 606, a mass storage device 608, an input device 610, and / or a communication link 612 (e.g., a communication bus), through which one or more of the components shown in the figure can exchange information.

[0044] The communication circuit 604 can be configured to transmit and receive information using one or more communication protocols (e.g., TCP / IP) and one or more communication networks including a local area network (LAN), a wide area network (WAN), the Internet, a wireless data network (e.g., Wi-Fi, 3G, 4G / LTE, or 5G network). The memory 606 may contain a storage medium (e.g., a non-transitory storage medium) configured to store machine-readable instructions that, when executed, cause the processor 602 to perform one or more of the functions described herein. Examples of machine-readable media may include volatile or non-volatile memory, including but not limited to semiconductor memory (e.g., electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM)), flash memory, etc. The mass storage device 608 may include one or more disks, such as one or more internal hard disks, one or more removable disks, one or more magneto-optical disks, one or more CD-ROM or DVD-ROM disks, etc., on which instructions and / or data can be stored to facilitate the operation of the processor 602. The input device 610 may include a keyboard, a mouse, a voice-controlled input device, a touch-sensitive input device (e.g., a touch screen), etc. for receiving user input to the device 600.

[0045] It should be noted that device 600 can operate as a stand-alone device or can be connected (e.g., networked or clustered) with other computing devices to perform the functions described herein. And even though Figure 6 only one instance of each component is shown in the figure, those skilled in the art will understand that device 600 may include multiple instances of one or more of the components shown in the figure.

[0046] Although the present disclosure has been described in terms of certain embodiments and the methods generally associated therewith, changes and permutations of the embodiments and methods will be apparent to those of ordinary skill in the art. Accordingly, the foregoing description of example embodiments does not limit the disclosure. Other variations, substitutions, and alterations are possible without departing from the spirit and scope of the disclosure. Additionally, unless specifically stated otherwise, discussions using terms such as "analyze", "determine", "enable", "identify", "modify", etc. refer to the actions and processes of a computer system or similar electronic computing device that manipulates and transforms data represented as physical (e.g., electronic) quantities within the registers and memories of the computer system into other data represented as physical quantities within the memories or other such information storage, transmission, or display devices of the computer system.

[0047] It should be understood that the foregoing description is intended to be illustrative and not restrictive. After reading and understanding the foregoing description, many other implementations will be apparent to those skilled in the art. Accordingly, the scope of the disclosure should be determined with reference to the appended claims and the full scope of equivalents to which those claims are entitled.

[0048] As used herein, the term "computer-readable storage medium" may include any tangible medium that can store or encode a set of instructions executable by a computer that cause the computer to perform any one or more of the methods described herein. The term "computer-readable storage medium" as used herein may include, but is not limited to, solid-state memory, optical media, and magnetic media.

Claims

1. A method for reconstructing a three-dimensional (3D) representation of an anatomical structure: Obtain a plurality of screen images of a display, wherein the plurality of screen images are captured while a medical representation of the anatomical structure is being displayed on the display; Extract a plurality of medical images of the anatomical structure from the plurality of screen images; Determine one or more parameters for reconstructing the 3D representation of the anatomical structure; And Reconstruct the 3D representation of the anatomical structure based on the plurality of extracted medical images and the one or more determined parameters.

2. The method according to claim 1, wherein The plurality of screen images are captured from the display of a desktop computer, a laptop computer, a tablet computer, or a mobile phone.

3. The device according to claim 1, wherein, The plurality of screen images are captured as a screen recording of the display.

4. The method according to claim 1, wherein Further comprising detecting an abnormality associated with the anatomical structure based on one or more pre-trained machine learning (ML) models and the 3D representation of the anatomical structure, and providing an indication of the abnormality on the 3D representation of the anatomical structure.

5. The method according to claim 1, wherein Further comprising generating a report associated with the anatomical structure based on one or more pre-trained machine learning (ML) models and the 3D representation of the anatomical structure.

6. The method according to claim 1, wherein The one or more determined parameters include a distance between two of the plurality of extracted medical images or a voxel size of the 3D representation.

7. The method according to claim 8, wherein Further comprising predicting the distance between two of the plurality of extracted medical images using a pre-trained machine learning (ML) model, and determining the voxel size based on the distance.

8. The method according to claim 1, wherein Reconstructing the 3D representation of the anatomical structure based on the plurality of extracted medical images and the one or more determined parameters includes: identifying one or more duplicates among the plurality of extracted medical images and excluding the one or more duplicates from the reconstruction of the 3D representation.

9. The method according to claim 1, wherein, Reconstructing the 3D representation of the anatomical structure based on the plurality of extracted medical images and the one or more determined parameters includes that one or more processors are configured to adjust the size, orientation, or aspect ratio of at least one of the plurality of extracted medical images.

10. A computer program product that, when run on a computer, causes the computer to perform the method according to any one of claims 1-9.