Deep learning image analysis with increased modularization and reduced footprint
By introducing a parallel architecture of public backbone networks and task-specific backbone networks into deep learning neural networks, the problem of high computing resources consumption in multi-task processing in the existing technology is solved, and efficient and modular image analysis capabilities are achieved.
Patent Information
- Application Number
- CN202380072966.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-13
- Filing Date
- 2023-08-21
- Publication Date
- 2025-05-30
AI Technical Summary
Existing deep learning neural networks consume a lot of computing resources and lack modularity when performing multiple medical image inference tasks, resulting in large space and reduced performance.
A deep learning neural network architecture is designed, including a public backbone network parallel to multiple task-specific backbone networks. Through parallel processing of public backbone networks and task-specific backbone networks, the consumption of computing resources is reduced, and a white box structure is realized through a modular architecture.
With less computing resources, efficient processing of performing multiple inference tasks on medical images is achieved, and the modularity and transparency of the system is improved, avoiding the redundancy and complexity of the fully connected architecture.
Smart Images

Figure CN120077387A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This subject application claims priority to U.S. Non - Provisional Patent Application Serial No. 18 / 046,347, filed on October 13, 2022, entitled "DEEP LEARNING IMAGE ANALYSIS WITH INCREASED MODULARITY AND REDUCED FOOTPRINT", the entire content of which is incorporated herein by reference. Technical Field
[0003] This disclosure generally relates to deep learning, and more particularly to deep learning image analysis with increased modularity and reduced footprint. Background Art
[0004] Trainable deep learning neural networks can perform inference tasks on medical images. Unfortunately, achieving high - performance deep learning neural networks often consumes excessive computing resources.
[0005] Therefore, a system or technique that can solve one or more of these technical problems may be desirable. Summary of the Invention
[0006] The following presents a summary of the invention to provide a basic understanding of one or more embodiments of the invention. This summary is not intended to identify key or important elements, nor to delineate any scope of particular embodiments or any scope of the claims. Its sole purpose is to present concepts in a simplified form as a prelude to the more detailed description that is presented later. In one or more embodiments described herein, devices, systems, computer - implemented methods, apparatuses, or computer program products that facilitate deep learning image analysis with increased modularity and reduced footprint are described.
[0007] According to one or more embodiments, a system is provided. The system may include a non - transitory computer - readable memory that may store computer - executable components. The system may also include a processor that may be operably coupled to the non - transitory computer - readable memory and may execute the computer - executable components stored in the non - transitory computer - readable memory. In various embodiments, the computer - executable components may include an access component capable of accessing medical imaging data. In various aspects, the computer - executable components may further include an inference component that may perform a plurality of inference tasks on the medical imaging data via the execution of a deep learning neural network. In various cases, the deep learning neural network may include a common backbone network in parallel with a plurality of task - specific backbone networks. In various cases, the plurality of task - specific backbone networks may respectively correspond to the plurality of inference tasks.
[0008] According to one or more embodiments, the above system may be implemented as a computer-implemented method or a computer program product. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 A block diagram illustrating an exemplary non-limiting system in accordance with one or more embodiments described herein, the system facilitating deep learning image analysis with increased modularity and reduced footprint.
[0010] Figure 2 A block diagram illustrating an exemplary non-limiting system in accordance with one or more embodiments described herein, the system including a deep learning neural network and multiple inference outputs, the system facilitating deep learning image analysis with increased modularity and reduced footprint.
[0011] Figure 3 An exemplary non-limiting block diagram in accordance with one or more embodiments described herein, showing how a deep learning neural network may include a backbone network portion, a splicing portion, and a head portion.
[0012] Figure 4 An exemplary non-limiting block diagram of a backbone network portion of a deep learning neural network in accordance with one or more embodiments described herein.
[0013] Figure 5 An exemplary non-limiting block diagram of a splicing portion of a deep learning neural network in accordance with one or more embodiments described herein.
[0014] Figure 6 An exemplary non-limiting block diagram of a head portion of a deep learning neural network in accordance with one or more embodiments described herein.
[0015] Figure 7 A block diagram illustrating an exemplary non-limiting system including training components and a training data set in accordance with one or more embodiments described herein, the system facilitating deep learning image analysis with increased modularity and reduced footprint.
[0016] Figure 8 An exemplary non-limiting block diagram of a training data set in accordance with one or more embodiments described herein.
[0017] Figure 9 An exemplary non-limiting block diagram in accordance with one or more embodiments described herein, showing how a deep learning neural network may be trained.
[0018] Figures 10 to 14A flow chart illustrating an exemplary non-limiting computer-implemented method for training a deep learning neural network according to one or more embodiments described herein.
[0019] Figure 15 A flow chart illustrating an exemplary non-limiting computer-implemented method for facilitating deep learning image analysis with increased modularity and reduced footprint according to one or more embodiments described herein.
[0020] Figure 16 A block diagram illustrating an example non-limiting operating environment in which one or more embodiments described herein may be facilitated.
[0021] Figure 17 An example networking environment is illustrated that is operable to perform various implementations described herein. DETAILED DESCRIPTION
[0022] The following detailed description is merely illustrative and is not intended to limit the embodiments or the application / use of the embodiments. In addition, it is not intended to be bound by any express or implied information set forth in the aforementioned "background technology" or "invention content" section or "detailed description" section.
[0023] One or more embodiments are now described with reference to the accompanying drawings, wherein the same reference numerals are used to represent the same elements throughout. In the following description, for the purpose of explanation, many specific details are set forth in order to provide a more thorough understanding of one or more embodiments. However, it is apparent that in various cases, one or more embodiments may be practiced without these specific details.
[0024] Deep learning neural networks can be trained (e.g., via supervised training, unsupervised training, reinforcement learning) to perform inference tasks (e.g., image quality enhancement, image denoising, image kernel transformation, image segmentation, image classification) on medical images (e.g., scanned / reconstructed images generated by a computed tomography (CT) scanner, scanned / reconstructed images generated by a magnetic resonance imaging (MRI) scanner, scanned / reconstructed images generated by a positron emission tomography (PET) scanner, scanned / reconstructed images generated by an X-ray scanner, scanned / reconstructed images generated by an ultrasound scanner).
[0025] Unfortunately, implementing high-performance deep learning neural networks often consumes a large amount of computing resources. For example, when performing a given inference task on any given medical image, a deep learning neural network that has achieved high accuracy may consume a large amount of time (e.g., it may consume tens of milliseconds of inference time). Another example is that a deep learning neural network that has achieved high accuracy for performing a given inference task may have tens of millions of internal parameters, and the electronic storage or execution of these tens of millions of internal parameters may consume a large amount of computer memory or processing power (e.g., random access memory (RAM) capacity, graphics processing unit (GPU) capacity).
[0026] In the context of an operation that expects to perform multiple different inference tasks (e.g., pneumothorax classification or segmentation, endotracheal tube positioning, brightness contrast enhancement) on any given medical image, this increased consumption of computing resources can be complex. In particular, for each of such multiple different inference tasks, different or unique deep learning neural networks can be trained to perform such an inference task on the input medical image. This can result in multiple different deep learning neural networks, each of which is capable of performing a corresponding one of the multiple different inference tasks. Thus, when a medical image is obtained, each of such multiple different deep learning neural networks can be executed on the medical image in order to jointly perform multiple different inference tasks on the medical image. Unfortunately, since storing or executing a single deep learning neural network may already consume a large amount of inference time or computer memory, storing or executing such multiple different deep learning neural networks may consume a large amount of inference time or computer memory.
[0027] Some existing technologies attempt to solve these problems by training a single deep learning neural network to perform all such multiple different inference tasks simultaneously. However, in such existing technologies, a single deep learning neural network typically implements a fully connected internal architecture. The inventors of the various embodiments described herein recognize that such a fully connected internal architecture can be disadvantageous. In particular, the inventors recognize that such a fully connected internal architecture may globally average (e.g., globally degrade) the performance of a single deep learning neural network across such multiple different inference tasks. In other words, a single deep learning neural network can perform each of such multiple different inference tasks, but the performance of each task is less than satisfactory. Additionally, the inventors recognize that such a fully connected internal architecture can be regarded as a black box that vaguely intertwines multiple different inference tasks with each other, which may unnecessarily complicate subsequent retraining or revalidation of a single deep learning neural network. For example, assume that a single deep learning neural network has achieved sufficient accuracy for all but a few of the multiple different inference tasks (e.g., as required or mandated by applicable regulatory restrictions). In such a case, because the fully connected internal architecture may obscure which internal parameters of the single deep learning neural network affect which inference tasks and in what way, the overall single deep learning neural network may have to undergo additional training or validation, even though the single deep learning neural network has been able to perform the vast majority of the multiple different inference tasks with sufficient accuracy. In other words, retraining a single deep learning neural network to perform those inference tasks for which it has already performed well enough may waste additional time and resources.
[0028] Accordingly, when implementing various existing technologies, a deep learning neural network may have an excessive footprint (e.g., may consume an excessive amount of computing resources) and may lack modularity (e.g., a fully connected internal architecture can be regarded as a black box and may thus require complete retraining or revalidation, even when a deep learning neural network performs most of its inference tasks accurately enough). Such a deep learning neural network can be disadvantageous, especially when deployed via resource-constrained computing devices (e.g., medical diagnostic devices, smartphones, autonomous vehicles).
[0029] Accordingly, a system or technology that can solve one or more of these technical problems may be desirable.
[0030] The various embodiments described herein can solve one or more of these technical problems. One or more embodiments described herein can include a system, a computer-implemented method, an apparatus, or a computer program product that can facilitate deep learning image analysis with increased modularity and reduced footprint. More specifically, the inventors have designed a deep learning neural network architecture that can perform multiple different inference tasks on an input medical image, which can perform this operation while consuming fewer computing resources compared to various prior arts, and is capable of performing this operation in a white-box manner rather than in the black-box manner of various prior arts.
[0031] In particular, such a deep learning neural network architecture can achieve these benefits by including a common backbone network (also referred to herein as a shared backbone network) in parallel with various task-specific backbone networks, where the various task-specific backbone networks can each correspond to (e.g., in a one-to-one manner) multiple different inference tasks. That is, each unique inference task can have a unique task-specific backbone network. In various cases, the common backbone network can be an order of magnitude larger, or even larger (e.g., can contain approximately ten times more internal parameters), than any of the various task-specific backbone networks. As described herein, the common backbone network and each of the various task-specific backbone networks can be trained to independently analyze a given medical image. This can cause the common backbone network to produce a first output based on the given medical image, and this can also cause each of the various task-specific backbone networks to produce a corresponding second output based on the given medical image. Also as described herein, each of such second outputs can be separately concatenated with the first output, thereby forming various concatenation results (e.g., one concatenation result per task-specific backbone network). As further described herein, the deep learning neural network architecture can include various task-specific heads corresponding to the various task-specific backbones (e.g., one task-specific head per task-specific backbone network). In various cases, each of the various task-specific heads can be configured to perform a corresponding inference task among multiple different inference tasks by receiving the corresponding concatenation result among the various concatenation results as input.
[0032] As explained herein, implementing the common backbone network in parallel with the various task-specific backbone networks can reduce the consumption of computing resources. In fact, since the common backbone network and each of the various task-specific backbone networks can be parallel to each other, they can each analyze a given medical image simultaneously or in a chronologically overlapping manner. Similarly, the various task-specific heads can be parallel to each other, which means they can analyze their corresponding concatenation results of the input simultaneously or in a chronologically overlapping manner. This can save inference time compared to the prior art of implementing separate, different deep learning neural networks for each inference task.
[0033] In addition, since the various second outputs generated by the various task-specific backbone networks can each be separately concatenated with the first output generated by the common backbone network, the common backbone network can be regarded as contributing its analysis to all of the plurality of different inference tasks without having to be repeatedly stored or repeatedly executed. Compared with the prior art that implements separate and different deep learning neural networks for each inference task, this can save computer memory or processing power. In fact, the present inventors have recognized that when a plurality of different inference tasks are all performed on medical images, such a plurality of different inference tasks may all involve at least some amount of common analysis (for example, even if the tasks are different, some amount of analysis between the tasks may be common because they are all performed on common imaging data). Therefore, the present inventors have recognized that when a plurality of different deep learning neural networks are separately trained to perform a plurality of different inference tasks, at least some parts of those plurality of different deep learning neural networks can all be regarded as performing common analysis with each other. In other words, such parts can be regarded as redundant (for example, being repeatedly stored or repeatedly executed). Therefore, the present inventors have recognized that computer memory or processing power can be saved by replacing such redundant parts of the plurality of different deep learning neural networks with a common backbone network.
[0034] In addition, since the various task-specific backbone networks may be different from each other or separated from the common backbone network, it is possible to know which specific internal parameters of such a deep learning neural network architecture affect which inference task among the plurality of different inference tasks and in what way. Therefore, any specific task-specific backbone network or task-specific head can be retrained or re-validated without affecting the common backbone network or any other task-specific backbone or task-specific head. In other words, such a deep learning neural network architecture can be regarded as a modular white box structure, as opposed to an opaque black box structure. This can be advantageous compared with the prior art of training a single deep learning neural network to perform a plurality of different inference tasks via a fully connected internal architecture.
[0035] The various embodiments described herein can be regarded as computerized tools (e.g., any suitable combination of computer-executable hardware or computer-executable software) that can facilitate deep learning image analysis with increased modularity and reduced footprint. In various aspects, such a computerized tool can include an access component, an inference component, or a result component.
[0036] In various embodiments, there may be medical imaging data. In various aspects, the medical imaging data can be any suitable electronic data, which may include any suitable number of medical images related to a medical patient (e.g., a human, an animal, or others). For example, in some cases, the medical imaging data may include one medical image related to a medical patient, where this one medical image may depict one or more anatomical structures (e.g., tissues, organs, body parts, or portions thereof) of the medical patient according to an imaging modality (e.g., this one medical image may be generated or captured by a CT scanner). As another example, in other cases, the medical imaging data may include multiple medical images related to a medical patient, and each of such multiple medical images depicts one or more anatomical structures of the medical patient according to a corresponding imaging modality (e.g., the first medical image in the medical imaging data may be generated or captured by a CT scanner, the second medical image in the medical imaging data may be generated or captured by an MRI scanner, the third medical image in the medical imaging data may be generated or captured by an X-ray scanner, the fourth medical image in the medical imaging data may be generated or captured by a PET scanner, the fifth medical image in the medical imaging data may be generated or captured by an ultrasound scanner). In any case, the medical images may exhibit any suitable format, size, or dimension (e.g., the medical images may be two-dimensional pixel arrays or three-dimensional voxel arrays).
[0037] In various aspects, it may be desirable to perform multiple inference tasks on the medical imaging data. In various cases, non-limiting examples of the inference tasks may include image quality enhancement, image denoising, image kernel transformation, image deconvolution, image segmentation, or image classification. In any case, as described herein, computerized tools may facilitate multiple inference tasks for the medical imaging data.
[0038] In various embodiments, an access component of the computerized tool may electronically receive or otherwise electronically access the medical imaging data. In some aspects, the access component may electronically retrieve the medical imaging data from any suitable centralized or decentralized data structure (e.g., a graphical data structure, a relational data structure, a hybrid data structure), whether remote from or local to the access component. In other aspects, the access component may electronically retrieve the medical imaging data from any one of the medical imaging devices that generate or capture the medical imaging data (e.g., a CT scanner, an MRI scanner, an X-ray scanner, a PET scanner, an ultrasound scanner). In any case, the access component may electronically obtain or access the medical images such that other components of the computerized tool may electronically interact with the medical imaging data (e.g., read, write, edit, copy, manipulate).
[0039] In various embodiments, the inference component of the computerized tool may electronically store, maintain, control, or otherwise access a deep learning neural network. In various aspects, as described herein, the deep learning neural network may be configured or trained to perform multiple inference tasks on input medical images. Thus, in various cases, the inference component may electronically execute the deep learning neural network on medical imaging data such that the deep learning neural network produces multiple inference results (e.g., one inference result per inference task) respectively corresponding to the multiple inference tasks.
[0040] In various cases, the inference result may be any suitable electronic data, the format, size, or dimension of which may depend on the inference task to which the inference result corresponds. As an example, if a particular inference task is image quality enhancement, the inference result corresponding to that particular inference task may be an inferred / predicted quality-enhanced version of the medical image in the medical imaging data. As another example, if a particular inference task is image denoising, the inference result corresponding to that particular inference task may be an inferred / predicted denoised version of the medical image in the medical imaging data. As yet another example, if a particular inference task is image kernel transformation, the inference result corresponding to that particular inference task may be an inferred / predicted kernel-transformed version of the medical image in the medical imaging data. As yet another example, if a particular inference task is image segmentation, the inference result corresponding to that particular inference task may be an inferred / predicted segmentation mask of the medical image in the medical imaging data. As yet another example, if a particular inference task is image classification, the inference result corresponding to that particular inference task may be an inferred / predicted classification label of the medical image in the medical imaging data.
[0041] As explained herein, the internal architecture of the deep learning neural network may be configured such that the deep learning neural network may consume fewer computational resources compared to multiple deep learning neural networks that are co-trained to perform multiple inference tasks. As also explained herein, the internal architecture of the deep learning neural network may be configured such that the deep learning neural network exhibits improved modularity or less internal opacity compared to deep learning neural networks that perform multiple inference tasks via a fully connected internal architecture.
[0042] In particular, the deep learning neural network may include a backbone network portion, a splicing portion, or a head portion.
[0043] In various aspects, the backbone network portion may include a common backbone network, multiple modality-specific backbone networks, or multiple task-specific backbone networks, all of which may be parallel to each other. In various cases, the common backbone network may include any suitable number of any suitable type of neural network layers (e.g., an input layer, one or more hidden layers, an output layer, where any layer may be a convolutional layer, a dense layer, a non-linear layer, a pooling layer, a batch normalization layer, or a padding layer), may include any suitable number of neurons in various layers (e.g., different layers may have the same or different numbers of neurons from each other), may include any suitable activation function in various neurons (e.g., softmax, sigmoid, hyperbolic tangent, rectified linear unit) (e.g., different neurons may have different or the same activation functions from each other), or may include any suitable inter-neuron connections or inter-layer connections (e.g., forward connections, skip connections, recurrent connections).
[0044] In various aspects, the multiple modality-specific backbone networks may include any suitable number of modality-specific backbone networks. In various cases, a modality-specific backbone network may include any suitable number of any suitable type of layers (e.g., an input layer, one or more hidden layers, an output layer, where any layer may be a convolutional layer, a dense layer, a non-linear layer, a pooling layer, a batch normalization layer, or a padding layer), may include any suitable number of neurons in various layers (e.g., different layers may have the same or different numbers of neurons from each other), may include any suitable activation function in various neurons (e.g., softmax, sigmoid, hyperbolic tangent, rectified linear unit) (e.g., different neurons may have different or the same activation functions from each other), or may include any suitable inter-neuron connections or inter-layer connections (e.g., forward connections, skip connections, recurrent connections). In various cases, a modality-specific backbone network may be an order of magnitude smaller than the common backbone network (e.g., may contain an order of magnitude fewer internal parameters). As described above, each modality-specific backbone network may be parallel to the common backbone network. Additionally, in various aspects, each modality-specific backbone network may be isolated from the common backbone network (e.g., there may be no forward connections or skip connections between the common backbone network and any modality-specific backbone network). Similarly, in various cases, each modality-specific backbone network may be isolated from each other modality-specific backbone network (e.g., there may be no forward connections or skip connections between any two modality-specific backbone networks).
[0045] In various aspects, the multiple task-specific backbone networks may each correspond to (e.g., in a one-to-one manner) the multiple inference tasks. In other words, each unique inference task may have a unique task-specific backbone network. In various cases, a task-specific backbone network may include any suitable number of any suitable types of layers (e.g., an input layer, one or more hidden layers, an output layer, where any layer may be a convolutional layer, a dense layer, a non-linear layer, a pooling layer, a batch normalization layer, or a padding layer), may include any suitable number of neurons in various layers (e.g., different layers may have the same or different numbers of neurons from each other), may include any suitable activation functions in various neurons (e.g., softmax, sigmoid, hyperbolic tangent, rectified linear unit) (e.g., different neurons may have different or the same activation functions from each other), or may include any suitable inter-neuron connections or inter-layer connections (e.g., forward connections, skip connections, recurrent connections). In various cases, a task-specific backbone network may be an order of magnitude smaller than a modality-specific backbone network or a common backbone network (e.g., may contain an order of magnitude fewer internal parameters). As described above, each task-specific backbone network may be parallel to the common backbone network. Additionally, in various aspects, each task-specific backbone network may be isolated from the common backbone network (e.g., there may be no forward connections or skip connections between the common backbone network and any modality-specific backbone network). Similarly, in various cases, each task-specific backbone network may be isolated from each other task-specific backbone network (e.g., there may be no forward connections or skip connections between any two task-specific backbone networks). Additionally, in various cases, each task-specific backbone network may be isolated from each modality-specific backbone network (e.g., for any task-specific backbone network and modality-specific backbone network, there may be no forward connections or skip connections between them).
[0046] In various aspects, the corresponding backbone network in the set of modality-specific backbone networks can correspond to a respective subset of the plurality of task-specific backbone networks. More specifically, each modality-specific backbone network in the plurality of modality-specific backbone networks can be considered to be associated with a respective medical imaging modality (e.g., the first modality-specific backbone network can be associated with the CT medical imaging modality, the second modality-specific backbone network can be associated with the MRI medical imaging modality, the third modality-specific backbone network can be associated with the X-ray medical imaging modality, the fourth modality-specific backbone network can be associated with the PET medical imaging modality, the fifth modality-specific backbone network can be associated with the ultrasound medical imaging modality). Additionally, in various cases, the plurality of task-specific backbone networks can be considered to be composed of a plurality of non-overlapping groupings (e.g., a plurality of non-overlapping subsets), which can be sorted by medical imaging modality (e.g., the first grouping of task-specific backbone networks can be associated with the CT medical imaging modality, the second grouping of task-specific backbone networks can be associated with the MRI medical imaging modality, the third grouping of task-specific backbone networks can be associated with the X-ray medical imaging modality, the fourth grouping of task-specific backbone networks can be associated with the PET medical imaging modality, the fifth grouping of task-specific backbone networks can be associated with the ultrasound medical imaging modality). In various cases, if a particular modality-specific backbone network is associated with a particular medical imaging modality and if a particular grouping of task-specific backbone networks is also associated with the same particular medical imaging modality, then such a particular modality-specific backbone network can be considered to correspond to such a particular grouping of task-specific backbone networks (e.g., a modality-specific backbone network associated with the CT medical imaging modality can be considered to correspond to any one of the groupings of task-specific backbone networks that are also associated with the CT medical imaging modality; a modality-specific backbone network associated with the MRI medical imaging modality can be considered to correspond to any one of the groupings of task-specific backbone networks that are also associated with the MRI medical imaging modality). In this way, each grouping of task-specific backbone networks can correspond to a respective one of the plurality of modality-specific backbone networks (e.g., any given task-specific backbone network can correspond to a corresponding modality-specific backbone network; any given modality-specific backbone network can correspond to one or more task-specific backbone networks).
[0047] In various aspects, the common backbone network can be configured to receive medical imaging data as input and produce a first intermediate output. More specifically, the input layer of the common backbone network can receive medical imaging data as input, the medical imaging data can complete a forward pass through one or more hidden layers of the common backbone network, and the output layer of the common backbone network can calculate the first intermediate output based on the activations provided by the one or more hidden layers of the common backbone network. In various cases, the first intermediate output can be any suitable electronic data having any suitable format, size, or dimension (e.g., it can be one or more scalars, one or more vectors, one or more matrices, one or more tensors, or one or more strings).
[0048] In various aspects, the plurality of modality-specific backbone networks can be configured to receive medical imaging data (or any suitable portion thereof) as input and produce a plurality of second intermediate outputs. More specifically, for any given modality-specific backbone network, the input layer of the given modality-specific backbone network can receive medical imaging data (or any suitable portion thereof) as input, the medical imaging data (or any suitable portion thereof) can complete a forward pass through one or more hidden layers of the given modality-specific backbone network, and the output layer of the given modality-specific backbone network can calculate the corresponding one of the plurality of second intermediate outputs based on the activations provided by the one or more hidden layers of the given modality-specific backbone network. In various cases, the second intermediate output can be any suitable electronic data having any suitable format, size, or dimension (e.g., it can be one or more scalars, one or more vectors, one or more matrices, one or more tensors, or one or more strings). In various cases, different modality-specific backbone networks can receive the same or different portions of the medical imaging data as input (e.g., the modality-specific backbone network associated with the CT medical imaging modality can receive the CT scan image in the medical imaging data as input; the modality-specific backbone network associated with the X-ray medical imaging modality can receive the X-ray scan image in the medical imaging data as input).
[0049] In various aspects, the plurality of task-specific backbone networks can be configured to receive medical imaging data (or any suitable portion thereof) as input and generate a plurality of third intermediate outputs. More specifically, for any given task-specific backbone network, the input layer of the given task-specific backbone network can receive medical imaging data (or any suitable portion thereof) as input, the medical imaging data (or any suitable portion thereof) can complete a forward pass through one or more hidden layers of the given task-specific backbone network, and the output layer of the given task-specific backbone network can calculate a corresponding one of the plurality of third intermediate outputs based on the activations provided by one or more hidden layers of the given task-specific backbone network. In various cases, the third intermediate output can be any suitable electronic data having any suitable format, size, or dimension (e.g., it can be one or more scalars, one or more vectors, one or more matrices, one or more tensors, or one or more strings). In various cases, different task-specific backbone networks can receive the same or different portions of the medical imaging data as input (e.g., a task-specific backbone network associated with the PET medical imaging modality can receive a PET scan image in the medical imaging data as input; a task-specific backbone network associated with the ultrasound medical imaging modality can receive an ultrasound scan image in the medical imaging data as input).
[0050] In various aspects, the concatenation portion of the deep learning neural network can include a plurality of concatenation layers, and all the concatenation layers can be parallel to each other. In various cases, the plurality of concatenation layers can be respectively connected in series with the plurality of task-specific backbone networks. That is, for any given task-specific backbone network, there can be a corresponding concatenation layer connected in series downstream of the given task-specific backbone network. In any case, the concatenation layer can be any suitable neural network layer capable of concatenating two or more inputs together.
[0051] In various aspects, a corresponding concatenation layer among the plurality of concatenation layers can concatenate a corresponding third intermediate output among the plurality of third intermediate outputs with the first intermediate output or a corresponding second intermediate output among the plurality of second intermediate outputs. As a non-limiting example, for any given task-specific backbone network, such a given task-specific backbone network can be considered to correspond to a given modality-specific backbone network. Additionally, such a given task-specific backbone network can be continuously upstream of a given concatenation layer. In various cases, the given concatenation layer can receive as input any one of the plurality of third intermediate outputs generated by the given task-specific backbone network; any one of the plurality of second intermediate outputs generated by the given modality-specific backbone network; and the first intermediate output generated by the common backbone network. Thus, in such a case, the given concatenation layer can concatenate these inputs together, thereby generating a concatenation result. In this way, the plurality of concatenation layers can jointly generate a plurality of concatenation results.
[0052] In various aspects, the head portion of the deep learning neural network may include a plurality of task-specific heads, all of which may be parallel to each other. In various cases, the plurality of task-specific heads may be serially connected to the plurality of splicing layers, respectively. That is, for any given splicing layer, there may be a corresponding task-specific head that is serially connected downstream of the given splicing layer. In any case, the task-specific head may include any suitable number of any suitable type of layers (e.g., an input layer, one or more hidden layers, an output layer, where any of the layers may be a convolutional layer, a dense layer, a non-linear layer, a pooling layer, a batch normalization layer, or a padding layer), may include any suitable number of neurons in various layers (e.g., different layers may have the same or different numbers of neurons from each other), may include any suitable activation function in various neurons (e.g., softmax, sigmoid, hyperbolic tangent, rectified linear unit) (e.g., different neurons may have different or the same activation functions from each other), or may include any suitable inter-neuron connections or inter-layer connections (e.g., forward connections, skip connections, recurrent connections). In various aspects, the plurality of task-specific heads may be parallel to each other. Additionally, in various cases, each task-specific head may be isolated from each other task-specific head (e.g., there may be no forward connection or skip connection between any two task-specific heads).
[0053] In various aspects, the plurality of task-specific heads may respectively generate the plurality of inference results based on the plurality of splicing results. As a non-limiting example, for any given splicing layer, such given splicing layer may be serially connected upstream of a given task-specific head. In various cases, the given task-specific head may receive any one of the plurality of splicing results generated by the given splicing layer as an input, and may generate a corresponding one of the plurality of inference results as an output. More specifically, the input layer of the given task-specific head may receive any one splicing generated by the given splicing layer, such splicing may complete a forward pass through one or more hidden layers of the given task-specific head, and the output layer of the given task-specific head may calculate a corresponding one of the plurality of inference results based on the activations provided by one or more hidden layers of the given task-specific head. In this way, the plurality of task-specific heads may jointly generate the plurality of inference results.
[0054] In various embodiments, the result component of the computerized tool may electronically initiate any suitable electronic action based on the plurality of inference results. As an example, the result component may electronically send one or more of the plurality of inference results to any suitable computing device. As another example, the result component may electronically present one or more of the plurality of inference results on any suitable computer screen, monitor, display, or graphical user interface. As yet another example, the result component may electronically generate any suitable warning or alert based on the plurality of inference results.
[0055] To help make the plurality of inference results accurate, the deep learning neural network may first undergo any suitable type or paradigm of training (e.g., supervised training, unsupervised training, reinforcement learning). Accordingly, in various aspects, the access component may receive, retrieve, or otherwise access a training data set, and the computerized tool may include a training component that can train the deep learning neural network on the training data set.
[0056] In some cases, the training data set may be an annotated training data set. In such a case, the training data set may include a set of training inputs and a set of multiple ground truth annotations respectively corresponding to the set of training inputs. In various aspects, any given training input may have the same format, size, or dimension as the medical imaging data described above. In various cases, since it may be desirable to train the deep learning neural network to perform the plurality of inference tasks, each training input may correspond to multiple ground truth annotations (e.g., one ground truth annotation per inference task). In various cases, the ground truth annotations may be considered known or be considered the correct or accurate inference results corresponding to the respective training inputs. In various aspects, the ground truth annotations may be manually made by a technician. In various other aspects, the ground truth annotations may be generated by a teacher network (e.g., a pre-trained neural network may have been trained to perform a given inference task on input medical imaging data; thus, such a pre-trained neural network may be fed the training inputs, which may cause the pre-trained neural network to produce inference results based on the training inputs, and such inference results may be considered or regarded as the ground truth annotations for training the deep learning neural network).
[0057] If the training data set is annotated, the training component may perform supervised training on the deep learning neural network in various aspects. Before starting such supervised training, the internal parameters (e.g., weights, biases, convolutional kernels) of the deep learning neural network (e.g., a common backbone network, a backbone network dedicated to each modality, a backbone network dedicated to each task, a head dedicated to each task) may be randomly initialized.
[0058] In various aspects, the training component can select any suitable training input from the training dataset and any suitable plurality of ground truth annotations corresponding to such selected training input. In various cases, the training component can feed the selected training input into a deep learning neural network, which can cause the deep learning neural network to produce a plurality of outputs. For example, the training input can complete a forward pass through the backbone network portion of the deep learning neural network, through the concatenation portion of the deep learning neural network, and through the head portion of the deep learning neural network, such that each task-specific head can produce a corresponding one of the plurality of outputs.
[0059] In various aspects, the plurality of outputs can be regarded as predictions or inferences that the deep learning neural network believes should correspond to the selected training input (e.g., predicted / inferred quality-enhanced image, predicted / inferred kernel-transformed image, predicted / inferred denoised image, predicted / inferred segmentation mask, predicted / inferred classification label). Conversely, the selected plurality of ground truth annotations can be regarded as the correct or accurate results known or believed to correspond to the selected training input (e.g., correct / accurate quality-enhanced image, correct / accurate kernel-transformed image, correct / accurate denoised image, correct / accurate segmentation mask, correct / accurate classification label). Note that if the deep learning neural network has not or has hardly been trained so far, the plurality of outputs may be very inaccurate (e.g., the plurality of outputs may be very different from the selected plurality of ground truth annotations).
[0060] In any case, the training component can compute one or more errors or losses (e.g., mean absolute error (MAE), mean squared error (MSE), cross entropy) between the plurality of outputs and the selected plurality of ground truth annotations. In various aspects, the training component can update the internal parameters of the deep learning neural network by performing backpropagation (e.g., stochastic gradient descent) driven by such computed errors or losses.
[0061] In various cases, this supervised training process can be repeated for each training input in the training dataset, with the result that the internal parameters of the deep learning neural network can become iteratively optimized to accurately generate predictions / inferences based on the input medical imaging data. In various cases, the training component can implement any suitable training batch size, any suitable training termination criterion, or any suitable error, loss, or objective function.
[0062] In various aspects, such training can be carried out in two stages. In the first stage of such training, the training component can feed the selected training input into the deep learning neural network and update the internal parameters via backpropagation, as described above. However, during this first stage, the training component can apply any suitable regularization term (e.g., L1 regularization, L2 regularization) to the plurality of task-specific backbone networks, and the training component can inhibit the application of this regularization term to the common backbone network. This use of regularization can enable the common backbone network to learn more than the plurality of task-specific backbone networks during this first stage of training. In the second stage of training, the training component can feed the selected training input into the deep learning neural network and update the internal parameters via backpropagation, as described above. However, during this second stage, the training component can remove this regularization term from the plurality of task-specific backbone networks, and the training component can freeze the internal parameters of the common backbone network. In other words, during this second stage of training, the common backbone network can remain unchanged, and the task-specific backbone networks can be considered to be fine-tuned. In some cases, the training component can apply any suitable regularization term to the plurality of modality-specific backbone networks or the plurality of task-specific heads during the first stage, and the training component can remove this regularization term during the second stage. However, in other cases, the training component can completely inhibit the application of the regularization term to the plurality of modality-specific backbone networks or the plurality of task-specific heads.
[0063] Note that if the deep learning neural network performs all but a few of the plurality of inference tasks with sufficient accuracy, the entire deep learning neural network does not need to undergo a full retraining or revalidation. Instead, the task-specific backbone networks or task-specific heads corresponding to such a few inference tasks can be identified, those identified task-specific backbone networks or task-specific heads can be retrained and revalidated, and the remaining part of the deep learning neural network (e.g., the common backbone network, the plurality of modality-specific backbone networks, the remaining part of the plurality of task-specific backbone networks, the remaining part of the plurality of task-specific heads) can be frozen (e.g., can remain unchanged) during such retraining. This is in contrast to the prior art that utilizes a fully connected architecture, in which case a full retraining is required.
[0064] In addition, it should be noted that if it is desired to teach a deep learning neural network how to perform a new inference task, the entire deep learning neural network does not need to undergo a complete retraining or revalidation. Instead, in various aspects, a new task-specific backbone network can be inserted into the plurality of task-specific backbone networks, a new concatenation layer can be inserted into the plurality of concatenation layers to be in series with the new task-specific backbone network, a new task-specific head can be inserted into the plurality of task-specific heads to be in series with the new concatenation layer, and new ground truth annotations corresponding to the new inference task can be associated with the training inputs (e.g., such new ground truth annotations can be generated by a new pre-trained teacher network). In this case, the training component can train the new task-specific backbone network and the new task-specific head while freezing the remaining part of the deep learning neural network (e.g., while keeping the common backbone network, the remaining parts of the plurality of modality-specific backbone networks, the remaining parts of the plurality of task-specific backbone networks, and the remaining parts of the plurality of task-specific heads unchanged). Similarly, this is in contrast to the prior art that uses a fully connected architecture, in which case a full retraining is required.
[0065] In addition, it should be noted that one or more of the plurality of inference tasks can be removed from the library of the deep learning neural network without affecting the remaining parts of the plurality of inference tasks. In particular, assume that it is desired to prevent the deep learning neural network from performing a given one of the plurality of inference tasks. In this case, it can be known which specific task-specific backbone network in the plurality of task-specific backbone networks, which specific concatenation layer in the plurality of concatenation layers, and which specific task-specific head in the plurality of task-specific heads correspond to the given inference task. Therefore, such a specific task-specific backbone network, such a specific concatenation layer, and such a specific task-specific head can be deleted, removed, or deactivated in other ways, so that the deep learning neural network no longer performs the given inference task. However, such deletion, removal, or deactivation may inhibit from affecting the remaining part of the deep learning neural network (e.g., may inhibit from affecting the common backbone network, the remaining parts of the plurality of modality-specific backbone networks, the remaining parts of the plurality of task-specific backbone networks, the remaining parts of the plurality of concatenation layers, and the remaining parts of the plurality of task-specific heads). In this way, the inference tasks can be selectively removed from the library of the deep learning neural network without affecting other parts of the deep learning neural network.
[0066] In any case, the common backbone network can be considered to perform analytical work that is useful (e.g., common) for all of the plurality of inference tasks, the modality-specific backbone network can be considered to perform analytical work that is useful (e.g., common) for a subset of the plurality of inference tasks (e.g., modality collation grouping), and the task-specific backbone network can be considered to perform analytical work that is useful only for a corresponding one of the plurality of inference tasks. Thus, compared with the prior art, by implementing the common backbone network or the plurality of modality-specific backbone networks, the consumption of computing resources can be reduced. Additionally, because the common backbone network, the plurality of modality-specific backbone networks, and the plurality of task-specific backbones can all be isolated or different from each other, the deep learning neural network can be considered to have a transparent, modular, white-box architecture (e.g., changes to the common backbone network do not affect any modality-specific backbone network or any task-specific backbone network; changes to a modality-specific backbone network do not affect any other modality-specific backbone network, the common backbone network, or any task-specific backbone network; changes to a task-specific backbone network do not affect any other task-specific backbone network, the common backbone network, or any modality-specific backbone network). Thus, if it is desired to improve the performance of the deep learning neural network with respect to a specific inference task, the discrete task-specific backbone network corresponding to such a specific inference task can be retrained or revalidated without affecting other parts of the deep learning neural network. Similarly, if it is desired to teach the deep learning neural network how to perform a new inference task, a new task-specific backbone network (as well as a new stitching layer and a new task-specific head) can be added to the deep learning neural network and can be trained or validated without affecting other parts of the deep learning neural network. Thus, compared with the prior art, by implementing the common backbone network, the plurality of modality-specific backbone networks, or the plurality of task-specific backbone networks as described herein, the modularity or transparency of the deep learning neural network can be increased.
[0067] The various embodiments described herein can be used to solve highly technical problems using hardware or software (e.g., to facilitate deep learning image analysis with increased modularity and reduced footprint), which are not abstract and cannot be performed as a set of mental acts of a human. Additionally, some of the processes performed can be carried out by a specialized computer (e.g., a deep learning neural network with internal parameter convolutional kernels) to implement defined tasks related to deep learning image analysis with increased modularity and reduced footprint. For example, such defined tasks can include: accessing medical imaging data through a device operatively coupled to a processor; and performing a plurality of inference tasks on the medical imaging data through the device and via the execution of a deep learning neural network, wherein the deep learning neural network includes a common backbone network parallel to a plurality of task-specific backbone networks, and wherein the plurality of task-specific backbone networks respectively correspond to the plurality of inference tasks. In various aspects, the deep learning neural network can further include a plurality of modality-specific backbone networks parallel to the common backbone network and parallel to the plurality of task-specific backbone networks, wherein a corresponding modality-specific backbone network among the plurality of modality-specific backbone networks corresponds to a respective subset of the plurality of task-specific backbone networks. In various cases, the deep learning neural network can further include a plurality of splicing layers respectively in series with the plurality of task-specific backbone networks, wherein the common backbone network can receive the medical imaging data as an input and can produce a first intermediate output, wherein the plurality of task-specific backbone networks can receive the medical imaging data as an input and can produce a plurality of second intermediate outputs, wherein the plurality of task-specific backbone networks can receive the medical imaging data as an input and can produce a plurality of third intermediate outputs, and wherein the plurality of splicing layers can splice a corresponding third intermediate output among the plurality of third intermediate outputs with the first intermediate output and with a corresponding second intermediate output among the plurality of second intermediate outputs, thereby producing a plurality of splicing results. In various cases, the deep learning neural network can further include a plurality of task-specific heads respectively in series with the plurality of splicing layers, wherein the plurality of task-specific heads can receive the plurality of splicing results as inputs and produce a plurality of inference outputs respectively corresponding to the plurality of inference tasks.
[0068] Such defined tasks are not performed manually by a human. In fact, neither a human brain nor a person with pen and paper can electronically access medical images (e.g., two-dimensional pixel arrays, three-dimensional voxel arrays) or electronically perform a deep learning neural network on these medical images. Similarly, a deep learning neural network is an inherently computerized construct that cannot be realized mentally by a human in any way without a computer. Accordingly, computerized tools that can train or execute a deep learning neural network are also inherently computerized and cannot be realized in any sensible, practical, or reasonable way without a computer.
[0069] In addition, the various embodiments described herein can integrate various teachings related to deep learning image analysis with increased modularity and reduced footprint into practical applications. As described above, some prior art trains multiple different deep learning neural networks to jointly perform multiple different inference tasks on input medical images. However, such techniques may consume excessive computational resources. Also as described above, other prior art trains a single deep learning neural network to perform all such multiple different inference tasks, where such a single deep learning neural network utilizes a fully connected internal architecture. Unfortunately, such techniques can be regarded as opaque black boxes lacking modularity.
[0070] The various embodiments described herein can address these technical problems. Specifically, the inventors have designed a deep learning neural network architecture which, as described herein, can consume fewer computational resources or exhibit increased modularity compared to the prior art. In particular, such a deep learning neural network architecture can include a common backbone network in parallel with multiple task-specific backbone networks. In various aspects, each task-specific backbone network can be regarded as performing analytical work useful for a single corresponding inference task, while the common backbone network can be regarded as performing analytical work useful for multiple different inference tasks. By implementing such an internal architecture, compared to prior art that trains multiple different deep learning neural networks to jointly perform multiple different inference tasks (e.g., such prior art can be regarded as redundantly storing or executing neural network layers that perform the same analytical work as each other), the deep learning neural network can consume less inference time, less computer memory, or less computer processing power. In addition, by implementing such an internal architecture, it can be known which specific internal parameters of the deep learning neural network affect or act on which specific inference tasks. That is, the deep learning neural network can be regarded as a transparent white box rather than an opaque black box, which is contrary to the technique of training a single deep learning neural network with a fully connected architecture to perform multiple different inference tasks. Therefore, the various embodiments described herein undoubtedly constitute a specific and tangible technical improvement in the field of deep learning. Thus, the various embodiments described herein are clearly eligible as useful and practical applications for a computer.
[0071] In addition, the various embodiments described herein can control a tangible device in the real world based on the disclosed teachings. For example, the various embodiments described herein can electronically execute (or train) a real-world deep learning neural network on real-world medical images (e.g., CT images, MRI images, X-ray images, PET images, ultrasound images), and can electronically render any results produced by such a real-world deep learning neural network on a real-world computer screen.
[0072] It should be understood that the figures and descriptions herein provide non-limiting examples of various embodiments and are not necessarily drawn to scale.
[0073] Figure 1 A block diagram of an exemplary non-limiting system 100 is illustrated in accordance with one or more embodiments described herein, which system can facilitate deep learning image analysis with increased modularity and reduced footprint. As shown, the image analysis system 102 can be electronically integrated with the medical imaging data 104 via any suitable wired or wireless electronic connection.
[0074] In various embodiments, the medical imaging data 104 can include any suitable number of medical images associated with a medical patient. As a non-limiting example, the medical imaging data 104 can include one medical image depicting any suitable anatomical structure of a medical patient according to one medical imaging modality. For example, the medical imaging data 104 can be an X-ray scan image depicting the anatomical structure of a medical patient. As another non-limiting example, the medical imaging data 104 can include multiple medical images, each depicting the anatomical structure of the medical image according to a corresponding medical imaging modality. For example, the first medical image in the medical imaging data 104 can be an X-ray scan image depicting the anatomical structure of a medical patient, the second medical image in the medical imaging data 104 can be a CT scan image depicting the same anatomical structure of the same medical patient, or the third medical image in the medical imaging data 104 can be an MRI scan image depicting the same anatomical structure of the same medical patient.
[0075] In any case, the medical images of the medical imaging data 104 can exhibit any suitable format, size, or dimension. For example, a medical image can be a pixel array with Hounsfield unit values of x times y pixels, where x and y are any suitable positive integers. As another example, a medical image can be a voxel array with Hounsfield unit values of x times y times z voxels, where x, y, and z are any suitable positive integers. In various cases, different medical images in the medical imaging data 104 can have the same or different formats, sizes, or dimensions from each other.
[0076] In various embodiments, it may be desirable to perform multiple inference tasks on medical imaging data 104. Non-limiting examples of such inference tasks may include: quality enhancement of CT scan images; quality enhancement of X-ray scan images; quality enhancement of MRI scan images; quality enhancement of PET scan images; quality enhancement of ultrasound scan images; denoising of CT scan images; denoising of X-ray scan images; denoising of MRI scan images; denoising of PET scan images; denoising of ultrasound scan images; kernel transformation of CT scan images; kernel transformation of X-ray scan images; kernel transformation of MRI scan images; kernel transformation of PET scan images; kernel transformation of ultrasound scan images; segmentation of CT scan images; segmentation of X-ray scan images; segmentation of MRI scan images; segmentation of PET scan images; segmentation of ultrasound scan images; classification of CT scan images; classification of X-ray scan images; classification of MRI scan images; classification of PET scan images; classification of ultrasound scan images; object localization within CT scan images; object localization within X-ray scan images; object localization within MRI scan images; object localization within PET scan images; object localization within ultrasound scan images. In any case, the image analysis system 102 may perform such multiple inference tasks on the medical imaging data 104 as described herein.
[0077] In various embodiments, the image analysis system 102 may include a processor 106 (e.g., a computer processing unit, a microprocessor) and a non-transitory computer-readable memory 108 that is operably or operationally or communicatively connected or coupled to the processor 106. The non-transitory computer-readable memory 108 may store computer-executable instructions that, when executed by the processor 106, may cause the processor 106 or other components of the image analysis system 102 (e.g., the access component 110, the inference component 112, the result component 114) to perform one or more actions. In various embodiments, the non-transitory computer-readable memory 108 may store computer-executable components (e.g., the access component 110, the inference component 112, the result component 114), and the processor 106 may execute the computer-executable components.
[0078] In various embodiments, the image analysis system 102 may include an access component 110. In various aspects, the access component 110 may electronically receive or otherwise electronically access medical imaging data 104. In various cases, the access component 110 may electronically retrieve the medical imaging data 104 from any suitable centralized or decentralized data structure (not shown) or from any suitable centralized or decentralized computing device (not shown). As a non-limiting example, any medical imaging device (e.g., CT scanner, MRI scanner, X-ray scanner, PET scanner, ultrasound scanner) that generates or captures the medical imaging data 104 may send the medical imaging data 104 to the access component 110. In any case, the access component 110 may electronically obtain or access the medical imaging data 104 such that other components of the image analysis system 102 may electronically interact with the medical imaging data 104.
[0079] In various embodiments, the image analysis system 102 may include an inference component 112. In various aspects, as described herein, the inference component 112 may execute a deep learning neural network on the medical imaging data 104, thereby generating a plurality of inference outputs respectively corresponding to the plurality of inference tasks.
[0080] In various embodiments, the image analysis system 102 may include a result component 114. In various cases, as described herein, the result component 114 may send any of the plurality of inference outputs to any suitable computing device or may present any of the plurality of inference outputs on any suitable computer display.
[0081] Figure 2 A block diagram of an exemplary non-limiting system 200 in accordance with one or more embodiments described herein is illustrated. The system includes a deep learning neural network and a plurality of inference outputs, and the system may facilitate deep learning image analysis with increased modularity and reduced footprint. As shown, the system 200 may include the same components as the system 100 in some cases, and may further include a deep learning neural network 202 or a plurality of inference outputs 204.
[0082] In various embodiments, the inference component 112 may electronically store, electronically maintain, electronically control, or otherwise electronically access the deep learning neural network 202. In various aspects, as described herein, the deep learning neural network 202 may be configured to perform a plurality of inference tasks on an input medical image. Thus, in various cases, the inference component 112 may electronically execute the deep learning neural network 202 on the medical imaging data 104, thereby generating a plurality of inference outputs 204. With respect to Figures 3 to 6 Further non-limiting aspects are described.
[0083] Figure 3An exemplary non-limiting block diagram is illustrated showing how a deep learning neural network may include a backbone network portion, a concatenation portion, and a head portion according to one or more embodiments described herein. In other words, Figure 3 An exemplary non-limiting block diagram of the internal architecture of deep learning neural network 202 is shown.
[0084] As shown, the deep learning neural network 202 may include a backbone network portion 302, a concatenated portion 304, or a head portion 306. In various aspects, the concatenated portion 304 may be connected in series downstream of the backbone network portion 302 (e.g., located downstream and in series therewith), and the head portion 306 may be connected in series downstream of the concatenated portion 304.
[0085] In various cases, as shown, the inference component 112 can perform the deep learning neural network 202 on the medical imaging data 104. That is, the medical imaging data 104 can complete a forward pass through the backbone network portion 302, through the concatenation portion 304, and through the head portion 306, which can cause the head portion 306 to generate a plurality of inference outputs 204.
[0086] In various aspects, the plurality of inference outputs 204 may correspond to (e.g., in a one-to-one manner) the plurality of inference tasks, respectively. In other words, the plurality of inference outputs 204 may include a unique inference output for each unique inference task that is desired to be performed on the medical imaging data 104. In various cases, the inference output may be any suitable electronic data in any suitable format, size, or dimension. For example, the inference output may be one or more scalars, one or more vectors, one or more matrices, one or more tensors, one or more strings, or any suitable combination thereof. In some cases, the format, size, or dimension of the inference output may depend on which inference task the inference output corresponds to. For example, if the inference output corresponds to a CT quality enhancement inference task, the inference output may be considered to be a quality enhanced version of a CT image included or otherwise specified in the medical imaging data 104. For another example, if the inference output corresponds to an MRI denoising inference task, the inference output may be considered to be a denoised version of an MRI image included or otherwise specified in the medical imaging data 104. For another example, if the inference output corresponds to an X-ray segmentation inference task, the inference output may be considered as a segmentation mask of an X-ray image included or otherwise specified in the medical imaging data 104. For another example, if the inference output corresponds to an ultrasound classification inference task, the inference output may be considered as a classification label of an ultrasound image included or otherwise specified in the medical imaging data 104.
[0087] Various non-limiting aspects of the backbone network portion 302, the splicing portion 304, and the header portion 306 are described with reference to Figures 4 to 6 Give a description.
[0088] Figure 4 Exemplary non - limiting block diagram 400 of the backbone network portion 302 of the deep learning neural network 202 in accordance with one or more embodiments described herein is illustrated.
[0089] As shown, the backbone network portion 302 may include a common backbone network 402. In various aspects, the common backbone network 402 may include any suitable number of any suitable type of neural network layers arranged in any suitable manner. For example, the common backbone network 402 may include an input layer, any suitable number of hidden layers, and an output layer. In various cases, any layer in such a layer may implement any suitable type of trainable internal parameters (e.g., any layer in such a layer may be a convolutional layer, and the trainable internal parameter of the convolutional layer may be a convolutional kernel; any layer in such a layer may be a dense layer, and the trainable internal parameter of the dense layer may be a weight matrix or a bias value; any layer in such a layer may be a batch normalization layer, and the trainable internal parameter of the batch normalization layer may be a shift factor or a scaling factor). In various cases, any layer in such a layer may implement any suitable type of non - trainable internal parameters (e.g., any layer in such a layer may be a pooling layer, a padding layer, or a non - linear layer, and the internal parameters of these layers may be considered fixed or otherwise non - trainable). Additionally, in various aspects, the common backbone network 402 may include any suitable number of any suitable type of intermediate neuron connections or inter - layer connections, such as forward connections, skip connections, or recurrent connections.
[0090] In various aspects, as shown, the backbone network portion 302 can include a plurality of modality-specific backbone networks 404. In various cases, the plurality of modality-specific backbone networks 404 can include p backbone networks, where p is any suitable positive integer: from the modality-specific backbone network 404(1) to the modality-specific backbone network 404(p). In various cases, the modality-specific backbone network can include any suitable number of any suitable type of neural network layers arranged in any suitable manner. For example, the modality-specific backbone network can include an input layer, any suitable number of hidden layers, and an output layer. In various aspects, any layer in such layers can implement any suitable type of trainable internal parameters (e.g., any layer in such layers can be a convolutional layer, and the trainable internal parameter of the convolutional layer can be a convolutional kernel; any layer in such layers can be a dense layer, and the trainable internal parameter of the dense layer can be a weight matrix or a deviation value; any layer in such layers can be a batch normalization layer, and the trainable internal parameter of the batch normalization layer can be a shift factor or a scaling factor). In various cases, any layer in such layers can implement any suitable type of non-trainable internal parameters (e.g., any layer in such layers can be a pooling layer, a padding layer, or a non-linear layer, and the internal parameters of these layers can be regarded as fixed or otherwise non-trainable). Additionally, in various cases, the modality-specific backbone network can include any suitable number of any suitable type of intermediate neuron connections or inter-layer connections, such as forward connections, skip connections, or recurrent connections. In various aspects, different modality-specific backbone networks among the plurality of modality-specific backbone networks 404 can exhibit the same or different architectures from each other.
[0091] In various cases, as shown, the plurality of modality-specific backbone networks 404 can be parallel (rather than serial) to the common backbone network 402. That is, each modality-specific backbone network among the plurality of modality-specific backbone networks 404 can be parallel to the common backbone network 402, and thus each modality-specific backbone network among the plurality of modality-specific backbone networks 404 can be parallel to each other modality-specific backbone network among the plurality of modality-specific backbone networks 404. Additionally, in various cases and as shown, the plurality of modality-specific backbone networks 404 can be isolated from the common backbone network 402. In other words, for any given modality-specific backbone network, there may be no intermediate neuron connections or inter-layer connections between the given modality-specific backbone network and the common backbone network 402. Additionally, in various aspects and as shown, the plurality of modality-specific backbone networks 404 can be isolated from each other. In other words, for any given modality-specific backbone network, there may be no intermediate neuron connections or inter-layer connections between the given modality-specific backbone network and each other modality-specific backbone network among the plurality of modality-specific backbone networks 404.
[0092] In various cases, the multiple modality-specific backbone networks 404 can be considered to be curated by medical imaging modalities. In other words, each modality-specific backbone network among the multiple modality-specific backbone networks 404 can be considered to be associated with a corresponding, unique medical imaging modality. As a non-limiting example, the modality-specific backbone network 404(1) can be considered to be associated with a first medical imaging modality, which means that the modality-specific backbone network 404(1) can be configured to receive as input medical images generated or captured according to such first medical imaging modality (e.g., the modality-specific backbone network 404(1) can be associated with a CT medical imaging modality, which means that the modality-specific backbone network 404(1) can be configured to receive CT scan images as input). As another non-limiting example, the modality-specific backbone network 404(p) can be considered to be associated with the p-th medical imaging modality, which means that the modality-specific backbone network 404(p) can be configured to receive as input medical images generated or captured according to such p-th medical imaging modality (e.g., the modality-specific backbone network 404(p) can be associated with a PET medical imaging modality, which means that the modality-specific backbone network 404(p) can be configured to receive PET scan images as input). In various cases, different modalities among the multiple modality-specific backbone networks 404 can be associated with different medical imaging modalities from one another.
[0093] In various aspects, as shown, the backbone network portion 302 can include a plurality of task-specific backbone networks 406. In various cases, the plurality of task-specific backbone networks 406 can each correspond to (e.g., in a one-to-one manner) a plurality of inference tasks desired to be performed on the medical imaging data 104. That is, the plurality of task-specific backbone networks 406 can include one task-specific backbone network per unique inference task. In various cases, a task-specific backbone network can include any suitable number of any suitable type of neural network layers arranged in any suitable manner. For example, a task-specific backbone network can include an input layer, any suitable number of hidden layers, and an output layer. In various aspects, any layer in such a layer can implement any suitable type of trainable internal parameters (e.g., any layer in such a layer can be a convolutional layer, and the trainable internal parameter of the convolutional layer can be a convolutional kernel; any layer in such a layer can be a dense layer, and the trainable internal parameter of the dense layer can be a weight matrix or a deviation value; any layer in such a layer can be a batch normalization layer, and the trainable internal parameter of the batch normalization layer can be a shift factor or a scaling factor). In various cases, any layer in such a layer can implement any suitable type of non-trainable internal parameters (e.g., any layer in such a layer can be a pooling layer, a padding layer, or a non-linear layer, and the internal parameters of these layers can be considered fixed or otherwise non-trainable). Additionally, in various cases, a task-specific backbone network can include any suitable number of any suitable type of intermediate neuron connections or inter-layer connections, such as forward connections, skip connections, or recurrent connections. In various aspects, different task-specific backbone networks among the plurality of task-specific backbone networks 406 can exhibit the same or different architectures from each other.
[0094] In various cases, as shown in the figure, multiple task-specific backbone networks 406 can be parallel (rather than in series) with a common backbone network 402 or with multiple modality-specific backbone networks 404. That is, each task-specific backbone network among the multiple task-specific backbone networks 406 can be parallel with the common backbone network 402, and thus each task-specific backbone network among the multiple task-specific backbone networks 406 can be parallel with each other task-specific backbone network among the multiple task-specific backbone networks 406 and with each modality-specific backbone network among the multiple modality-specific backbone networks 404. Additionally, in various cases and as shown in the figure, the multiple task-specific backbone networks 406 can be isolated from the common backbone network 402. In other words, for any given task-specific backbone network, there may be no intermediate neuron connections or inter-layer connections between the given task-specific backbone network and the common backbone network 402. Furthermore, in various aspects and as shown in the figure, the multiple task-specific backbone networks 406 can be isolated from each other. In other words, for any given task-specific backbone network, there may be no intermediate neuron connections or inter-layer connections between the given task-specific backbone network and each other task-specific backbone network among the multiple task-specific backbone networks 406. Even further, in various cases and as shown in the figure, the multiple task-specific backbone networks 406 can be isolated from the multiple modality-specific backbone networks 404. In other words, for any given task-specific backbone network and any given modality-specific backbone network, there may be no intermediate neuron connections or inter-layer connections between them.
[0095] In various aspects, corresponding subsets of the multiple task-specific backbone networks 406 can correspond to the corresponding modality-specific backbone networks of the multiple modality-specific backbone networks 404. In particular, as described above, the multiple modality-specific backbone networks 404 can be considered to be organized according to p medical imaging modalities (e.g., there may be p total modality-specific backbone networks, and each of such p total modality-specific backbone networks can be associated with a different or unique medical imaging modality). Thus, the multiple task-specific backbone networks 406 can be considered to have p subsets: subset 406(1) to subset 406(p). In various cases, each of such p subsets can have any suitable number of task-specific backbone networks. For example, subset 406(1) can have q backbone networks, where q is any suitable positive integer: task-specific backbone network 406(1)(1) to task-specific backbone network 406(1)(q). Similarly, subset 406(p) can have q backbone networks, where q is any suitable positive integer: task-specific backbone network 406(p)(1) to task-specific backbone network 406(p)(q).
[0096] In Figure 4In a non - limiting example, there may be a total of (p)(q) task - specific backbone networks among the multiple task - specific backbone networks 406. Thus, this may imply that it is desired to perform a total of (p)(q) inference tasks on the medical imaging data 104 (e.g., one task - specific backbone network per inference task).
[0097] Although Figure 4 Sub - set 406(1) and sub - set 406(p) are depicted as having the same number of task - specific backbone networks (e.g., q) with respect to each other, this is merely a non - limiting example for ease of illustration. In various cases, different sub - sets among the p sub - sets of the multiple task - specific backbone networks 406 may have the same or different numbers of task - specific backbone networks with respect to each other.
[0098] In various aspects, each of these p sub - sets of the multiple task - specific backbone networks 406 may be associated with a corresponding unique medical imaging modality, since each of these p sub - sets and each of the modality - specific backbone networks 404 among the multiple modality - specific backbone networks 404 may be associated with a corresponding unique medical imaging modality.
[0099] For example, as described above, the modality - specific backbone network 404(1) may be associated with a first medical imaging modality, which means that the modality - specific backbone network 404(1) may be configured to receive medical images generated or captured according to this first medical imaging modality as input. In various cases, sub - set 406(1) may also be associated with this first medical imaging modality, which means that each task - specific backbone network in sub - set 406(1) may be configured to receive medical images generated or captured according to this first medical imaging modality as input. Thus, sub - set 406(1) (e.g., task - specific backbone networks 406(1)(1) to task - specific backbone networks 406(1)(q)) may be considered to correspond to the modality - specific backbone network 404(1), since they are all associated with the first medical imaging modality (e.g., if the modality - specific backbone network 404(1) is configured to receive CT scan images as input, then the task - specific backbone networks 406(1)(1) to task - specific backbone networks 406(1)(q) may also be configured to receive CT scan images as input).
[0100] For another example, as mentioned above, the modality-specific backbone network 404(p) can be associated with the p-th medical imaging modality, which means that the modality-specific backbone network 404(p) can be configured to receive medical images generated or captured according to this p-th medical imaging modality as input. In various aspects, the subset 406(p) can also be associated with the p-th medical imaging modality, which means that each task-specific backbone network in the subset 406(p) can be configured to receive medical images generated or captured according to this p-th medical imaging modality as input. Thus, the subset 406(p) (e.g., the task-specific backbone networks 406(p)(1) to the task-specific backbone networks 406(p)(q)) can be regarded as corresponding to the modality-specific backbone network 404(p) because they are all associated with the p-th medical imaging modality (e.g., if the modality-specific backbone network 404(p) is configured to receive PET scan images as input, then the task-specific backbone networks 406(p)(1) to the task-specific backbone networks 406(p)(q) can also be configured to receive PET scan images as input).
[0101] In various aspects, the common backbone network 402 can receive the medical imaging data 104 as input, and the common backbone network 402 can produce features 408 as output. More specifically, the input layer of the common backbone network 402 can receive the medical imaging data 104, the medical imaging data 104 can complete a forward pass through one or more hidden layers of the common backbone network 402, and the output layer of the common backbone network 402 can calculate the features 408 based on the activation maps generated by the one or more hidden layers of the common backbone network 402. In any case, the features 408 can be any suitable electronic data exhibiting any suitable format, size, or dimension. For example, the features 408 can include one or more scalars, one or more vectors, one or more matrices, one or more tensors, one or more strings, or any suitable combination thereof.
[0102] In various cases, the multiple modality-specific backbone networks 404 can receive the medical imaging data 104 (or any suitable portion thereof) as input, and the multiple modality-specific backbone networks 404 can produce multiple features 410 as output. In particular, for any given modality-specific backbone network, the medical imaging data 104 (or any suitable portion thereof) can complete a forward pass through the given modality-specific backbone network, which can cause the given modality-specific backbone network to produce the corresponding feature among the multiple features 410. Thus, since the multiple modality-specific backbone networks 404 can include p backbone networks, the multiple features 410 can include p features: features 410(1) to features 410(p).
[0103] For example, the modality-specific backbone network 404(1) can receive medical imaging data 104 (or any suitable portion thereof) as input, and the modality-specific backbone network 404(1) can produce features 410(1) as output. More specifically, the input layer of the modality-specific backbone network 404(1) can receive medical imaging data 104 (or any suitable portion thereof), the medical imaging data 104 (or any suitable portion thereof) can complete a forward pass through one or more hidden layers of the modality-specific backbone network 404(1), and the output layer of the modality-specific backbone network 404(1) can calculate the features 410(1) based on activation maps generated by one or more hidden layers of the modality-specific backbone network 404(1). In some cases, the modality-specific backbone network 404(1) can receive all of the medical imaging data 104 as input. In other cases, the modality-specific backbone network 404(1) can receive as input any portion of the medical imaging data 104 that is associated with the same medical imaging modality as the modality-specific backbone network 404(1) (e.g., if the modality-specific backbone network 404(1) is associated with the CT medical imaging modality, the modality-specific backbone network 404(1) can receive one or more CT scan images in the medical imaging data 104). In any case, the features 410(1) can be any suitable electronic data that exhibits any suitable format, size, or dimension. For example, the features 410(1) can include one or more scalars, one or more vectors, one or more matrices, one or more tensors, one or more strings, or any suitable combination thereof.
[0104] For another example, the modality-specific backbone network 404(p) can receive medical imaging data 104 (or any suitable portion thereof) as input, and the modality-specific backbone network 404(p) can produce features 410(p) as output. In particular, the input layer of the modality-specific backbone network 404(p) can receive medical imaging data 104 (or any suitable portion thereof), the medical imaging data 104 (or any suitable portion thereof) can complete a forward pass through one or more hidden layers of the modality-specific backbone network 404(p), and the output layer of the modality-specific backbone network 404(p) can calculate the features 410(p) based on the activation maps generated by one or more hidden layers of the modality-specific backbone network 404(p). In some aspects, the modality-specific backbone network 404(p) can receive all of the medical imaging data 104 as input. In other aspects, the modality-specific backbone network 404(p) can receive as input any portion of the medical imaging data 104 that is associated with the same medical imaging modality as the modality-specific backbone network 404(p) (e.g., if the modality-specific backbone network 404(p) is associated with the PET medical imaging modality, the modality-specific backbone network 404(p) can receive one or more PET scan images in the medical imaging data 104). In any case, the features 410(p) can be any suitable electronic data that exhibits any suitable format, size, or dimension. For example, the features 410(p) can include one or more scalars, one or more vectors, one or more matrices, one or more tensors, one or more strings, or any suitable combination thereof.
[0105] In various cases, different features among the multiple features 410 can have the same or different formats, sizes, or dimensions from each other.
[0106] In various aspects, the multiple task-specific backbone networks 406 can receive medical imaging data 104 (or any suitable portion thereof) as input, and the multiple task-specific backbone networks 406 can produce multiple features 412 as output. In particular, for any given task-specific backbone network, the medical imaging data 104 (or any suitable portion thereof) can complete a forward pass through the given task-specific backbone network, which can cause the given task-specific backbone network to produce the corresponding feature among the multiple features 412.
[0107] For example, the task-specific backbone network 406(1)(1) may receive medical imaging data 104 (or any suitable portion thereof) as input, and the task-specific backbone network 406(1)(1) may produce features 412(1)(1) as output. More specifically, the input layer of the task-specific backbone network 406(1)(1) may receive medical imaging data 104 (or any suitable portion thereof), the medical imaging data 104 (or any suitable portion thereof) may complete a forward pass through one or more hidden layers of the task-specific backbone network 406(1)(1), and the output layer of the task-specific backbone network 406(1)(1) may calculate the features 412(1)(1) based on activation maps generated by one or more hidden layers of the task-specific backbone network 406(1)(1). In some cases, the task-specific backbone network 406(1)(1) may receive all of the medical imaging data 104 as input. In other cases, the task-specific backbone network 406(1)(1) may receive as input any portion of the medical imaging data 104 that is associated with the same medical imaging modality as the task-specific backbone network 406(1)(1) (e.g., if the task-specific backbone network 406(1)(1) is associated with the CT medical imaging modality, the task-specific backbone network 406(1)(1) may receive one or more CT scan images in the medical imaging data 104). In any case, the features 412(1)(1) may be any suitable electronic data that exhibits any suitable format, size, or dimension. That is, the features 412(1)(1) may include one or more scalars, one or more vectors, one or more matrices, one or more tensors, one or more strings, or any suitable combination thereof.
[0108] For another example, the task-specific backbone network 406(1)(q) can receive medical imaging data 104 (or any suitable portion thereof) as input, and the task-specific backbone network 406(1)(q) can produce features 412(1)(q) as output. In particular, the input layer of the task-specific backbone network 406(1)(q) can receive medical imaging data 104 (or any suitable portion thereof), the medical imaging data 104 (or any suitable portion thereof) can complete a forward pass through one or more hidden layers of the task-specific backbone network 406(1)(q), and the output layer of the task-specific backbone network 406(1)(q) can calculate the features 412(1)(q) based on the activation maps generated by one or more hidden layers of the task-specific backbone network 406(1)(q). In some aspects, as described above, the task-specific backbone network 406(1)(q) can receive all of the medical imaging data 104 as input. In other aspects, the task-specific backbone network 406(1)(q) can receive as input any portion of the medical imaging data 104 that is associated with the same medical imaging modality as the task-specific backbone network 406(1)(q) (e.g., if the task-specific backbone network 406(1)(q) is associated with the CT medical imaging modality, the task-specific backbone network 406(1)(q) can receive one or more CT scan images in the medical imaging data 104). In any case, the features 412(1)(q) can be any suitable electronic data that exhibits any suitable format, size, or dimension (e.g., the features 412(1)(1) can include one or more scalars, one or more vectors, one or more matrices, one or more tensors, one or more strings, or any suitable combination thereof).
[0109] For another example, the task-specific backbone network 406(p)(1) can receive medical imaging data 104 (or any suitable portion thereof) as input, and the task-specific backbone network 406(p)(1) can produce features 412(p)(1) as output. For example, the input layer of the task-specific backbone network 406(p)(1) can receive medical imaging data 104 (or any suitable portion thereof), the medical imaging data 104 (or any suitable portion thereof) can complete a forward pass through one or more hidden layers of the task-specific backbone network 406(p)(1), and the output layer of the task-specific backbone network 406(p)(1) can calculate the features 412(p)(1) based on activation maps generated by one or more hidden layers of the task-specific backbone network 406(p)(1). In some aspects, the task-specific backbone network 406(p)(1) can receive all of the medical imaging data 104 as input. In other aspects, the task-specific backbone network 406(p)(1) can receive as input any portion of the medical imaging data 104 that is associated with the same medical imaging modality as the task-specific backbone network 406(p)(1) (e.g., if the task-specific backbone network 406(p)(1) is associated with the PET medical imaging modality, the task-specific backbone network 406(p)(1) can receive one or more PET scan images in the medical imaging data 104). In any case, the features 412(p)(1) can be any suitable electronic data that exhibits any suitable format, size, or dimension (e.g., the features 412(p)(1) can include one or more scalars, one or more vectors, one or more matrices, one or more tensors, one or more strings, or any suitable combination thereof).
[0110] For another example, the task-specific backbone network 406(p)(q) can receive medical imaging data 104 (or any suitable portion thereof) as input, and the task-specific backbone network 406(p)(q) can produce features 412(p)(q) as output. For example, the input layer of the task-specific backbone network 406(p)(q) can receive medical imaging data 104 (or any suitable portion thereof), the medical imaging data 104 (or any suitable portion thereof) can complete a forward pass through one or more hidden layers of the task-specific backbone network 406(p)(q), and the output layer of the task-specific backbone network 406(p)(q) can calculate the features 412(p)(q) based on the activation maps generated by one or more hidden layers of the task-specific backbone network 406(p)(q). In some cases, the task-specific backbone network 406(p)(q) can receive all of the medical imaging data 104 as input. In other cases, the task-specific backbone network 406(p)(q) can receive as input any portion of the medical imaging data 104 that is associated with the same medical imaging modality as the task-specific backbone network 406(p)(q) (e.g., if the task-specific backbone network 406(p)(q) is associated with the PET medical imaging modality, the task-specific backbone network 406(p)(q) can receive one or more PET scan images in the medical imaging data 104). In any case, the features 412(p)(q) can be any suitable electronic data that exhibits any suitable format, size, or dimension (e.g., the features 412(p)(q) can include one or more scalars, one or more vectors, one or more matrices, one or more tensors, one or more strings, or any suitable combination thereof).
[0111] In various cases, the features 412(1)(1) to features 412(1)(q) can be collectively regarded as a subset 412(1) of the multiple features 412 (e.g., the subset 412(1) can be generated by the subset 406(1)). Similarly, in various cases, the features 412(p)(1) to features 412(p)(q) can be collectively regarded as a subset 412(p) of the multiple features 412 (e.g., the subset 412(p) can be generated by the subset 406(p)).
[0112] Note that because the multiple modality-specific backbone networks 404 can be parallel to the common backbone network 402, when the common backbone network 402 generates the features 408, the multiple modality-specific backbone networks 404 can generate multiple features 410 simultaneously or in a chronologically overlapping manner. Similarly, because the multiple task-specific backbone networks 406 can be parallel to the common backbone network 402, when the common backbone network 402 generates the features 408, the multiple task-specific backbone networks 406 can generate multiple features 412 simultaneously or in a chronologically overlapping manner.
[0113] Figure 5 Exemplary non - limiting block diagram 500 illustrating the stitching portion 304 of the deep learning neural network 202 in accordance with one or more embodiments described herein.
[0114] In various aspects, the stitching portion 304 may include a plurality of stitching layers 502. In various cases, the plurality of stitching layers 502 may respectively correspond to (e.g., in a one - to - one manner) a plurality of task - specific backbone networks 406. In particular, the plurality of stitching layers 502 may be in series with the plurality of task - specific backbone networks 406 respectively. In other words, for any given task - specific backbone network, the corresponding one of the plurality of stitching layers 502 may be downstream of and in series with the given task - specific backbone network. For example, the stitching layer 502(1)(1) may be downstream of and in series with the task - specific backbone network 406(1)(1). Also, for example, the stitching layer 502(1)(q) may be downstream of and in series with the task - specific backbone network 406(1)(q). Also, for example, the stitching layer 502(p)(1) may be downstream of and in series with the task - specific backbone network 406(p)(1). Also, for example, the stitching layer 502(p)(q) may be downstream of and in series with the task - specific backbone network 406(p)(q).
[0115] In various cases, the stitching layers 502(1)(1) to 502(1)(q) may be collectively regarded as a subset 502(1) of the plurality of stitching layers 502 (e.g., the subset 502(1) may be in series with the subset 406(1) respectively). Similarly, in various cases, the stitching layers 502(p)(1) to 502(p)(q) may be collectively regarded as a subset 502(p) of the plurality of stitching layers 502 (e.g., the subset 502(p) may be in series with the subset 406(p) respectively).
[0116] In any case, a stitching layer may be any suitable neural network layer capable of stitching together two or more inputs.
[0117] In various aspects, a corresponding concatenation layer in the plurality of concatenation layers 502 may concatenate corresponding features in the plurality of features 412 with corresponding features in the plurality of features 410 or with features 408, thereby generating a plurality of concatenation results 504. More specifically, for any given concatenation layer, the given concatenation layer may be connected in series with a corresponding task-specific backbone network, and the corresponding task-specific backbone network may correspond to a corresponding modality-specific backbone network. Therefore, such a given concatenation layer may receive any one of the plurality of features 412 generated by the corresponding task-specific backbone network, any one of the plurality of features 410 generated by the corresponding modality-specific backbone network, and the feature 408 generated by the common backbone network 402 as input, and such a given concatenation layer may generate a corresponding one of the plurality of concatenation results 504 based on such input.
[0118] For example, concatenation layer 502(1)(1) may be connected in series with task-specific backbone network 406(1)(1), and task-specific backbone network 406(1)(1) may correspond to modality-specific backbone network 404(1). Thus, concatenation layer 502(1)(1) may receive as input: features 412(1)(1) generated by task-specific backbone network 406(1)(1); features 410(1) generated by modality-specific backbone network 404(1); and features 408 generated by common backbone network 402. In various cases, concatenation layer 502(1)(1) may concatenate such inputs together (e.g., may concatenate features 412(1)(1) with features 410(1) and with features 408) to produce concatenation result 504(1)(1).
[0119] As another example, concatenation layer 502(1)(q) may be connected in series with task-specific backbone network 406(1)(q), and task-specific backbone network 406(1)(q) may correspond to modality-specific backbone network 404(1). Thus, concatenation layer 502(1)(q) may receive as input: features 412(1)(q) generated by task-specific backbone network 406(1)(q); features 410(1) generated by modality-specific backbone network 404(1); and features 408 generated by common backbone network 402. In various cases, concatenation layer 502(1)(q) may concatenate such inputs together (e.g., may concatenate features 412(1)(q) with features 410(1) and with features 408), thereby generating concatenation result 504(1)(q).
[0120] For another example, the splicing layer 502(p)(1) can be connected in series with the task-specific backbone network 406(p)(1), and the task-specific backbone network 406(p)(1) can correspond to the modality-specific backbone network 404(p). Therefore, the splicing layer 502(p)(1) can receive the following as inputs: the feature 412(p)(1) generated by the task-specific backbone network 406(p)(1); the feature 410(p) generated by the modality-specific backbone network 404(p); and the feature 408 generated by the common backbone network 402. In various cases, the splicing layer 502(p)(1) can splice this kind of input together (for example, the feature 412(p)(1) can be spliced with the feature 410(p) and with the feature 408), thereby generating the splicing result 504(p)(1).
[0121] For another example, the splicing layer 502(p)(q) can be connected in series with the task-specific backbone network 406(p)(q), and the task-specific backbone network 406(p)(q) can correspond to the modality-specific backbone network 404(p). Therefore, the splicing layer 502(p)(q) can receive the following as inputs: the feature 412(p)(q) generated by the task-specific backbone network 406(p)(q); the feature 410(p) generated by the modality-specific backbone network 404(p); and the feature 408 generated by the common backbone network 402. In various cases, the splicing layer 502(p)(q) can splice this kind of input together (for example, the feature 412(p)(q) can be spliced with the feature 410(p) and with the feature 408), thereby generating the splicing result 504(p)(q).
[0122] In various cases, the splicing results 504(1)(1) to 504(1)(q) can be jointly regarded as a subset 504(1) of the multiple splicing results 504 (for example, the subset 504(1) can be respectively generated by the subset 502(1)). Similarly, in various cases, the splicing results 504(p)(1) to 504(p)(q) can be jointly regarded as a subset 504(p) of the multiple splicing results 504 (for example, the subset 504(p) can be respectively generated by the subset 502(p)).
[0123] It should be understood that Figure 5 the shown splicing paths are only non-limiting examples. In all aspects, any other suitable splicing paths can be implemented. That is, in various cases, the splicing layer in the multiple splicing layers 502 can splice one or more features among the multiple features 412 with one or more features among the multiple features 410 or with the feature 408.
[0124] In addition, it should be understood that the concatenation part 304 of the deep learning neural network 202 is only a non-limiting example. In various aspects, the concatenation part 304 can be more generally regarded as a "combination part" that can combine the various outputs generated by the backbone network part 302 in any suitable manner. In this case, such a combination part can be more generally regarded as including a plurality of combination layers rather than including a plurality of concatenation layers 502.
[0125] As a non-limiting example, the concatenation layer 502(1)(1) can alternatively be more generally regarded as a "combination" layer 502(1)(1), and the concatenation result 504(1)(1) can alternatively be more generally regarded as a "combination" 504(1)(1). In this case, the "combination" layer 502(1)(1) can combine the feature 412(1)(1) with the feature 410(1) and with the feature 408 in any suitable manner. In some cases, this can be facilitated via concatenation as described above. However, in other cases, this can be facilitated without using concatenation. For example, in some aspects, the "combination" layer 502(1)(1) can project the feature 412(1)(1), the feature 410(1), and the feature 408 into the same dimensional space with each other via any suitable dimensionality reduction or projection technique, and the "combination" layer 502(1)(1) can add these projections together to produce the "combination" 504(1)(1).
[0126] As another non-limiting example, the concatenation layer 502(1)(q) can alternatively be more generally regarded as a "combination" layer 502(1)(q), and the concatenation result 504(1)(q) can alternatively be more generally regarded as a "combination" 504(1)(q). In this case, the "combination" layer 502(1)(q) can combine the feature 412(1)(q) with the feature 410(1) and with the feature 408 in any suitable manner. In some cases, this can be facilitated via concatenation as described above. However, in other cases, this can be facilitated without using concatenation. For example, in some aspects, the "combination" layer 502(1)(q) can project the feature 412(1)(q), the feature 410(1), and the feature 408 into the same dimensional space with each other via any suitable dimensionality reduction or projection technique, and the "combination" layer 502(1)(q) can add these projections together to produce the "combination" 504(1)(q).
[0127] As yet another non - limiting example, the concatenation layer 502(p)(1) may alternatively and more generally be regarded as a "combination" layer 502(p)(1), and the concatenation result 504(p)(1) may alternatively and more generally be regarded as a "combination" 504(p)(1). In this case, the "combination" layer 502(p)(1) can combine the feature 412(p)(1) with the feature 410(p) and with the feature 408 in any suitable manner. In some cases, this may be facilitated via concatenation, as described above. However, in other cases, this may be facilitated without using concatenation. For example, in some aspects, the "combination" layer 502(p)(1) can project the feature 412(p)(1), the feature 410(p), and the feature 408 into the same dimensional space with each other via any suitable dimensionality reduction or projection technique, and the "combination" layer 502(p)(1) can add these projections together, thereby producing the "combination" 504(p)(1).
[0128] As yet another non - limiting example, the concatenation layer 502(p)(q) may alternatively and more generally be regarded as a "combination" layer 502(p)(q), and the concatenation result 504(p)(q) may alternatively and more generally be regarded as a "combination" 504(p)(q). In this case, the "combination" layer 502(p)(q) can combine the feature 412(p)(q) with the feature 410(p) and with the feature 408 in any suitable manner. In some cases, this may be facilitated via concatenation, as described above. However, in other cases, this may be facilitated without using concatenation. For example, in some aspects, the "combination" layer 502(p)(q) can project the feature 412(p)(q), the feature 410(p), and the feature 408 into the same dimensional space with each other via any suitable dimensionality reduction or projection technique, and the "combination" layer 502(p)(q) can add these projections together, thereby producing the "combination" 504(p)(q).
[0129] That is, in various embodiments, the outputs of the backbone network portion 302 (e.g., 408, 410, 412) can be combined in any suitable manner (e.g., not limited to being combined only via concatenation).
[0130] Figure 6 Exemplary non - limiting block diagram 600 of the head portion 306 of the deep - learning neural network 202 in accordance with one or more embodiments described herein is illustrated.
[0131] In various aspects, the head portion 306 may include a plurality of task-specific heads 602. In various cases, the plurality of task-specific heads 602 may respectively correspond to (e.g., in a one-to-one manner) a plurality of splicing layers 502. In particular, the plurality of task-specific heads 602 may be respectively in series with the plurality of splicing layers 502. In other words, for any given splicing layer, the corresponding one of the plurality of task-specific heads 602 may be downstream of and in series with the given splicing layer. For example, the task-specific head 602(1)(1) may be downstream of and in series with the splicing layer 502(1)(1). As another example, the task-specific head 602(1)(q) may be downstream of and in series with the splicing layer 502(1)(q). As another example, the task-specific head 602(p)(1) may be downstream of and in series with the splicing layer 502(p)(1). As another example, the task-specific head 602(p)(q) may be downstream of and in series with the splicing layer 502(p)(q).
[0132] In various cases, the task-specific heads 602(1)(1) to the task-specific head 602(1)(q) may be collectively regarded as a subset 602(1) of the plurality of task-specific heads 602 (e.g., the subset 602(1) may be respectively in series with the subset 502(1)). Similarly, in various cases, the task-specific heads 602(p)(1) to the task-specific head 602(p)(q) may be collectively regarded as a subset 602(p) of the plurality of task-specific heads 602 (e.g., the subset 602(p) may be respectively in series with the subset 502(p)).
[0133] In various aspects, the task-specific head may include any suitable number of any suitable type of neural network layers arranged in any suitable manner. For example, the task-specific head may include an input layer, any suitable number of hidden layers, and an output layer. In various cases, any of such layers may implement any suitable type of trainable internal parameters (e.g., any of such layers may be a convolutional layer, and the trainable internal parameters of this convolutional layer may be convolutional kernels; any of such layers may be a dense layer, and the trainable internal parameters of this dense layer may be a weight matrix or a bias value; any of such layers may be a batch normalization layer, and the trainable internal parameters of this batch normalization layer may be a shift factor or a scaling factor). In various cases, any of such layers may implement any suitable type of non-trainable internal parameters (e.g., any of such layers may be a pooling layer, a padding layer, or a non-linear layer, and the internal parameters of these layers may be regarded as fixed or otherwise non-trainable). Additionally, in various aspects, the task-specific head may include any suitable number of any suitable type of intermediate neuron connections or inter-layer connections, such as forward connections, skip connections, or recurrent connections. In various cases, different task-specific heads among the multiple task-specific heads 602 may exhibit the same or different architectures from each other.
[0134] In various aspects, the multiple task-specific heads 602 may generate multiple inference outputs 204 based on the multiple splicing results 504. More specifically, for any given task-specific head, this given task-specific head may be connected in series with the corresponding splicing layer. Thus, this given task-specific head may receive any one of the multiple splicing results 504 generated by this corresponding splicing layer as an input, and this given task-specific head may generate a corresponding one of the multiple inference outputs 204 based on such a splicing result.
[0135] For example, the task-specific head 602(1)(1) may be connected in series with the splicing layer 502(1)(1), and the splicing layer 502(1)(1) may generate the splicing result 504(1)(1). Thus, the task-specific head 602(1)(1) may receive the splicing result 504(1)(1) as an input and may produce the inference output 204(1)(1). Specifically, the input layer of the task-specific head 602(1)(1) may receive the splicing result 504(1)(1), the splicing result 504(1)(1) may complete the forward pass through one or more hidden layers of the task-specific head 602(1)(1), and the output layer of the task-specific head 602(1)(1) may calculate the inference output 204(1)(1) based on the activation map generated by one or more hidden layers of the task-specific head 602(1)(1).
[0136] For another example, the task-specific head 602(1)(q) can be connected in series with the splicing layer 502(1)(q), and the splicing layer 502(1)(q) can generate a splicing result 504(1)(q). Therefore, the task-specific head 602(1)(q) can receive the splicing result 504(1)(q) as input and can produce an inference output 204(1)(q). More specifically, the input layer of the task-specific head 602(1)(q) can receive the splicing result 504(1)(q), the splicing result 504(1)(q) can complete a forward pass through one or more hidden layers of the task-specific head 602(1)(q), and the output layer of the task-specific head 602(1)(q) can calculate the inference output 204(1)(q) based on the activation map generated by one or more hidden layers of the task-specific head 602(1)(q).
[0137] For another example, the task-specific head 602(p)(1) can be connected in series with the splicing layer 502(p)(1), and the splicing layer 502(p)(1) can generate a splicing result 504(p)(1). Therefore, the task-specific head 602(p)(1) can receive the splicing result 504(p)(1) as input and can produce an inference output 204(p)(1). That is, the input layer of the task-specific head 602(p)(1) can receive the splicing result 504(p)(1), the splicing result 504(p)(1) can complete a forward pass through one or more hidden layers of the task-specific head 602(p)(1), and the output layer of the task-specific head 602(p)(1) can calculate the inference output 204(p)(1) based on the activation map generated by one or more hidden layers of the task-specific head 602(p)(1).
[0138] For another example, the task-specific head 602(p)(q) can be connected in series with the splicing layer 502(p)(q), and the splicing layer 502(p)(q) can generate a splicing result 504(p)(q). Therefore, the task-specific head 602(p)(q) can receive the splicing result 504(p)(q) as input and can produce an inference output 204(p)(q). In other words, the input layer of the task-specific head 602(p)(q) can receive the splicing result 504(p)(q), the splicing result 504(p)(q) can complete a forward pass through one or more hidden layers of the task-specific head 602(p)(q), and the output layer of the task-specific head 602(p)(q) can calculate the inference output 204(p)(q) based on the activation map generated by one or more hidden layers of the task-specific head 602(p)(q).
[0139] In various cases, the inference outputs 204(1)(1) through 204(1)(q) can be collectively regarded as a subset 204(1) of the multiple inference outputs 204 (e.g., the subset 204(1) can be generated by the subsets 602(1) respectively). Similarly, in various cases, the inference outputs 204(p)(1) through 204(p)(q) can be collectively regarded as the 204(p) of the multiple inference outputs 204 (e.g., the subset 204(p) can be generated by the subsets 602(p) respectively).
[0140] Although not explicitly shown in Figures 4 to 6 the deep learning neural network 202 can include any suitable gating layer that can determine or otherwise control which of the multiple modality-specific backbone networks 404, which of the multiple task-specific backbone networks 406, which of the multiple splicing layers 502, or which of the multiple task-specific heads 602 are activated during any given forward pass.
[0141] For example, as described herein, the medical imaging data 104 can complete a forward pass through the deep learning neural network 202, thereby generating multiple inference outputs 204. During such a forward pass, as described above, the medical imaging data 104 (or any suitable portion thereof) can be analyzed by the multiple modality-specific backbone networks 404. However, in some cases, the medical imaging data 104 (or any suitable portion thereof) can be analyzed by fewer than all of the multiple modality-specific backbone networks 404. In particular, any suitable neural network gating layer can be parallel to the multiple modality-specific backbone networks 404, and such a neural network gating layer can determine or otherwise control which of the multiple modality-specific backbone networks 404 should receive the medical imaging data 104 (or any suitable portion thereof) and which of the multiple modality-specific backbone networks 404 should not receive the medical imaging data.
[0142] As another example, as described herein, the medical imaging data 104 can complete a forward pass through the deep learning neural network 202, thereby generating multiple inference outputs 204. During such a forward pass, as described above, the medical imaging data 104 (or any suitable portion thereof) can be analyzed by the multiple task-specific backbone networks 406. However, in some cases, the medical imaging data 104 (or any suitable portion thereof) can be analyzed by fewer than all of the multiple task-specific backbone networks 406. In particular, any suitable neural network gating layer can be parallel to the multiple task-specific backbone networks 406, and such a neural network gating layer can determine or otherwise control which of the multiple task-specific backbone networks 406 should receive the medical imaging data 104 (or any suitable portion thereof) and which of the multiple task-specific backbone networks 406 should not receive the medical imaging data.
[0143] As another example, as described herein, the medical imaging data 104 may complete a forward pass through the deep learning neural network 202, thereby generating a plurality of inference outputs 204. During such a forward pass, as described above, the medical imaging data 104 (or any suitable portion thereof) may be analyzed by a plurality of stitching layers 502. However, in some cases, the medical imaging data 104 (or any suitable portion thereof) may be analyzed by less than all of the plurality of stitching layers 502. More specifically, any suitable neural network gating layer may be parallel to the plurality of stitching layers 502, and such a neural network gating layer may determine or otherwise control which of the plurality of stitching layers 502 should receive the medical imaging data 104 (or any suitable portion thereof) and which of the plurality of stitching layers 502 should not receive the medical imaging data.
[0144] As another example, as described herein, the medical imaging data 104 may complete a forward pass through the deep learning neural network 202, thereby generating a plurality of inference outputs 204. During such a forward pass, as described above, the medical imaging data 104 (or any suitable portion thereof) may be analyzed by a plurality of task-specific heads 602. However, in some cases, the medical imaging data 104 (or any suitable portion thereof) may be analyzed by less than all of the plurality of task-specific heads 602. More specifically, any suitable neural network gating layer may be parallel to the plurality of task-specific heads 602, and such a neural network gating layer may determine or otherwise control which of the plurality of task-specific heads 602 should receive the medical imaging data 104 (or any suitable portion thereof) and which of the plurality of task-specific heads 602 should not receive the medical imaging data.
[0145] In this way, any suitable gating layer may be implemented to dynamically change which parts of the deep learning neural network 202 can be executed during any suitable forward pass (e.g., which of the plurality of modality-specific backbone networks 404, which of the plurality of task-specific backbone networks 406, which of the plurality of stitching layers 502, or which of the plurality of task-specific heads 602).
[0146] In any case, the inference component 112 may execute the deep learning neural network 202 on the medical imaging data 104, thereby generating a plurality of inference outputs 204.
[0147] Return reference Figure 2, in various embodiments, the result component 114 may initiate any suitable electronic action based on the plurality of inference outputs 204. As a non-limiting example, the result component 114 may electronically send one or more of the plurality of inference outputs 204 to any suitable computing device (not shown) to notify a user or technician that one or more of the plurality of inference outputs 204 have been generated. As another non-limiting example, the result component 114 may electronically present one or more of the plurality of inference outputs 204 on any suitable computer screen, computer monitor, computer display, or graphical user interface (not shown) to allow a user or technician to visually inspect one or more of the plurality of inference outputs 204. As a non-limiting example, such a computer screen, computer monitor, computer display, or graphical user interface may be integrated into or otherwise associated with any medical imaging device that generates the medical imaging data 104.
[0148] To help ensure that the plurality of inference outputs 204 are accurate, the deep learning neural network 202 may first undergo training. Aspects of such training are described Figures 7 to 14 in various non-limiting respects.
[0149] Figure 7 FIG. 700 illustrates a block diagram of an exemplary non-limiting system 700 including a training component and a training data set, according to one or more embodiments described herein, which may facilitate deep learning image analysis with increased modularity and reduced footprint. As shown, in some cases, system 700 may include the same components as system 200 and may further include a training component 702 and a training data set 704.
[0150] In various embodiments, the access component 110 may electronically receive, retrieve, obtain, or otherwise access the training data set 704 from any suitable source. In various aspects, the training component 702 may execute the deep learning neural network 202 based on the training data set 704. Aspects of such training are described Figures 8 to 14 in various non-limiting respects.
[0151] Figure 8 FIG. 800 illustrates a block diagram according to one or more embodiments described herein, showing an exemplary non-limiting embodiment of the training data set 704. In various aspects, as shown, the training data set 704 may include a set of training medical imaging inputs 802 and a set of multiple ground truth annotations 804.
[0152] In various cases, the set of training medical imaging inputs 802 can include n inputs, where n is any suitable positive integer: training medical imaging input 802(1) through training medical imaging input 802(n). In various cases, a training medical imaging input can be any suitable electronic data having the same format, size, or dimension as the medical imaging data 104. For example, if the medical imaging data 104 includes CT scan images, MRI scan images, and X-ray scan images of the anatomical structure of a medical imaging patient, then each training medical imaging input can similarly include CT scan images, MRI scan images, and X-ray scan images of the corresponding anatomical structure of the corresponding medical patient.
[0153] In various aspects, the set of multiple ground truth annotations 804 can respectively correspond to (e.g., in a one-to-one manner) the set of training medical imaging inputs 802. Thus, since the set of training medical imaging inputs 802 can have n inputs, the set of multiple ground truth annotations 804 can have n multiple: multiple ground truth annotations 804(1) through multiple ground truth annotations 804(n). In various cases, each multiple ground truth annotation can be any suitable electronic data representing or indicating the inference output that would be achieved if the multiple inference tasks were performed correctly or accurately on the corresponding training medical imaging input.
[0154] For example, multiple ground truth annotations 804(1) may correspond to a training medical imaging input 802(1). Recall that, as described above, it may be desirable for the deep learning neural network 202 to perform a total of (p)(q) inference tasks on the input medical imaging data. Thus, the multiple ground truth annotations 804(1) may be considered to represent or indicate that if each of these (p)(q) inference tasks were performed accurately or correctly on the training medical imaging input 802(1), a total of (p)(q) inference results would be obtained. In other words, the multiple ground truth annotations 804(1) may be considered to represent or indicate the ground truth results that should be output by the multiple task-specific heads 602 in response to the deep learning neural network 202 being performed on the training medical imaging input 802(1). Specifically, the multiple ground truth annotations 804(1) may include an annotation 804(1)(1)(1), where the annotation 804(1)(1)(1) may be considered the ground truth result that should be output by the task-specific head 602(1)(1) in the case of the deep learning neural network 202 being performed on the training medical imaging input 802(1). Similarly, the multiple ground truth annotations 804(1) may include an annotation 804(1)(1)(q), where the annotation 804(1)(1)(q) may be considered the ground truth result that should be output by the task-specific head 602(q)(1) in the case of the deep learning neural network 202 being performed on the training medical imaging input 802(1). Additionally, the multiple ground truth annotations 804(1) may include an annotation 804(1)(p)(1), where the annotation 804(1)(p)(1) may be considered the ground truth result that should be output by the task-specific head 602(p)(1) in the case of the deep learning neural network 202 being performed on the training medical imaging input 802(1). Additionally, the multiple ground truth annotations 804(1) may include an annotation 804(1)(p)(q), where the annotation 804(1)(p)(q) may be considered the ground truth result that should be output by the task-specific head 602(p)(q) in the case of the deep learning neural network 202 being performed on the training medical imaging input 802(1).
[0155] In various cases, the annotations 804(1)(1)(1) through 804(1)(1)(q) may be collectively considered a subset 804(1)(1) of the multiple ground truth annotations 804(1). Similarly, in various cases, the annotations 804(1)(p)(1) through 804(1)(p)(q) may be collectively considered a subset 804(1)(p) of the multiple ground truth annotations 804(1).
[0156] In various aspects, each of the multiple ground truth annotations 804(1) can be manually created by a person skilled in the art based on the training medical imaging input 802(1). In various other aspects, each of the multiple ground truth annotations 804(1) can be generated by a respective pre-trained teacher network. In particular, there can be a total of (p)(q) teacher networks, each of which is pre-trained to perform a respective one of a total of (p)(q) inference tasks on an input medical image. Thus, each of these total (p)(q) teacher networks can independently perform on the training medical imaging input 802(1), resulting in a respective output, and such respective output can be regarded as or considered to be one of the multiple ground truth annotations 804(1).
[0157] As another example, the multiple ground truth annotations 804(n) can correspond to the training medical imaging input 802(n). Similarly, as described above, it may be desirable for the deep learning neural network 202 to perform a total of (p)(q) inference tasks on the input medical imaging data. Thus, the multiple ground truth annotations 804(n) can be regarded as representing or indicating that if each of these (p)(q) inference tasks is accurately or correctly performed on the training medical imaging input 802(n), a total of (p)(q) inference results are obtained. That is, the multiple ground truth annotations 804(n) can be considered to represent or indicate the ground truth results that should be output by the multiple task-specific heads 602 in response to the deep learning neural network 202 being performed on the training medical imaging input 802(n). More specifically, the multiple ground truth annotations 804(n) can include the annotation 804(n)(1)(1), where the annotation 804(n)(1)(1) can be regarded as the ground truth result that should be output by the task-specific head 602(1)(1) in the case of the deep learning neural network 202 being performed on the training medical imaging input 802(n). Similarly, the multiple ground truth annotations 804(n) can include the annotation 804(n)(1)(q), where the annotation 804(n)(1)(q) can be regarded as the ground truth result that should be output by the task-specific head 602(1)(q) in the case of the deep learning neural network 202 being performed on the training medical imaging input 802(n). In addition, the multiple ground truth annotations 804(n) can include the annotation 804(n)(p)(1), where the annotation 804(n)(p)(1) can be regarded as the ground truth result that should be output by the task-specific head 602(p)(1) in the case of the deep learning neural network 202 being performed on the training medical imaging input 802(n). In addition, the multiple ground truth annotations 804(n) can include the annotation 804(n)(p)(q), where the annotation 804(n)(p)(q) can be regarded as the ground truth result that should be output by the task-specific head 602(p)(q) in the case of the deep learning neural network 202 being performed on the training medical imaging input 802(n).
[0158] In various cases, Annotations 804(n)(1)(1) through 804(n)(1)(q) can be collectively regarded as a subset 804(n)(1) of multiple true Annotations 804(n). Similarly, in various cases, Annotations 804(n)(p)(1) through 804(n)(p)(q) can be collectively regarded as a subset 804(n)(p) of multiple true Annotations 804(n).
[0159] In various aspects, each of the multiple true Annotations 804(n) can be manually made by a technician based on the training medical imaging input 802(n). In various other aspects, each of the multiple true Annotations 804(n) can be generated by a corresponding pre-trained teacher network. In particular, as described above, there can be a total of (p)(q) teacher networks, each of which is pre-trained to perform a corresponding one of a total of (p)(q) inference tasks on the input medical image. Thus, each of these total (p)(q) teacher networks can be independently performed on the training medical imaging input 802(n), resulting in a corresponding output, and such a corresponding output can be regarded as or considered to be one of the multiple true Annotations 804(n).
[0160] Figure 9 Exemplary non - limiting block diagram 900 is illustrated showing how a deep learning neural network 202 can be trained in accordance with one or more embodiments described herein.
[0161] In various aspects, the training component 702 can initialize the trainable internal parameters (e.g., convolutional kernels, weight matrices, bias values) of the deep learning neural network 202 (e.g., common backbone network 402, multiple modality - specific backbone networks 404, multiple task - specific backbone networks 406, multiple task - specific heads 602) in any suitable manner (e.g., random initialization) before starting the training.
[0162] In various aspects, the training component 702 can select a training medical imaging input 902 and multiple true Annotations 904 corresponding to the training medical imaging input 902 from the training data set 704. In various cases, as shown, the multiple true Annotations 904 can include a subset 904(1), where the subset 904(1) can have Annotations 904(1)(1) through 904(1)(q). Also as shown, the multiple true Annotations 904 can include a subset 904(p), where the subset 904(p) can have Annotations 904(p)(1) through 904(p)(q).
[0163] In various cases, the training component 702 may execute the deep learning neural network 202 on the training medical imaging input 902, such that the deep learning neural network 202 generates a plurality of outputs 906. More specifically, the training medical imaging input 902 may complete a forward pass through the backbone network portion 302, through the concatenation portion 304, and through the head portion 306, and the plurality of task-specific heads 602 may generate the plurality of outputs 906. More specifically, the task-specific head 602(1)(1) may generate the output 906(1)(1), the task-specific head 602(1)(q) may generate the output 906(1)(q), the task-specific head 602(p)(1) may generate the output 906(p)(1), and the task-specific head 602(p)(q) may generate the output 906(p)(q). In various cases, the outputs 906(1)(1) to 906(1)(q) may be collectively regarded as a subset 906(1) of the plurality of outputs 906. Similarly, the outputs 906(p)(1) to 906(p)(q) may be collectively regarded as a subset 906(p) of the plurality of outputs 906.
[0164] In various aspects, the plurality of outputs 906 may be regarded as prediction or inference results that the deep learning neural network 202 believes should correspond to the training medical imaging input 902 (e.g., predicted / inferred quality-enhanced image, predicted / inferred denoised image, predicted / inferred kernel-transformed image, predicted / inferred segmentation mask, predicted / inferred classification label). In contrast, the plurality of ground truth annotations 904 may be regarded as the correct / accurate results that are known or considered to correspond to the training medical imaging input 902 (e.g., correct / accurate quality-enhanced image, correct / accurate denoised image, correct / accurate kernel-transformed image, correct / accurate segmentation mask, correct / accurate classification label). Note that if the deep learning neural network 202 has not or has hardly been trained so far, the plurality of outputs 906 may be very inaccurate (e.g., the output 906(1)(1) may be very different from the annotation 904(1)(1)); the output 906(1)(q) may be very different from the annotation 904(1)(q); the output 906(p)(1) may be very different from the annotation 904(p)(1); or the output 906(p)(q) may be very different from the annotation 904(p)(q).
[0165] In various aspects, the training component 702 may compute one or more errors or losses (e.g., MAE, MSE, cross-entropy) between the multiple outputs 906 and the multiple ground truth annotations 904. For example, the training component 702 may compute the error or loss between output 906(1)(1) and annotation 904(1)(1), may compute the error or loss between output 906(1)(q) and annotation 904(1)(q), may compute the error or loss between output 906(p)(1) and annotation 904(p)(1), or may compute the error or loss between output 906(p)(q) and annotation 904(p)(q). In some cases, the training component 702 may aggregate such one or more errors or losses (e.g., via averaging).
[0166] In any case, the training component 702 may update the trainable internal parameters of the deep learning neural network 202 (e.g., the common backbone network 402, the multiple modality-specific backbone networks 404, the multiple task-specific backbone networks 406, the multiple task-specific heads 602) via backpropagation (e.g., stochastic gradient descent) driven by such one or more errors or losses.
[0167] In various cases, the training component 702 may repeat this execution and update process for each training medical imaging input in the training dataset 704. This may ultimately cause the trainable internal parameters of the deep learning neural network 202 to be iteratively optimized to accurately perform multiple inference tasks (e.g., a total of (p)(q) inference tasks) on the input medical imaging data. In various cases, the training component 702 may implement any suitable training batch size, any suitable training termination criterion, or any suitable error, loss, or objective function.
[0168] In some aspects, the training component 702 can train the deep learning neural network 202 in a two-stage manner. In the first stage of this two-stage manner, the training component 702 can execute the deep learning neural network 202 on selected training medical imaging inputs and can update the internal parameters of the deep learning neural network 202 via backpropagation as described above. However, in this first stage, the training component 702 can apply any suitable regularization term (e.g., L1 regularization, L2 regularization) to the plurality of task-specific backbone networks 406 or the plurality of task-specific heads 602. Additionally, in such a first stage, the training component 702 can inhibit applying regularization terms to the common backbone network 402 or the plurality of modality-specific backbone networks 404. In various cases, this implementation of regularization can enable the common backbone network 402 or the plurality of modality-specific backbone networks 404 to learn more than the plurality of task-specific backbone networks 406 or the plurality of task-specific heads 602. In the second stage of this two-stage manner, the training component 702 can execute the deep learning neural network 202 on selected training medical imaging inputs and can update the internal parameters of the deep learning neural network 202 via backpropagation as described above. However, in this second stage, the training component 702 can inhibit applying regularization terms to the plurality of task-specific backbone networks 406 or the plurality of task-specific heads 602. Additionally, in this second stage, the training component 702 can freeze the internal parameters of the common backbone network 402 or the plurality of modality-specific backbone networks 404. In various cases, this freezing and lack of regularization can be regarded as fine-tuning the plurality of task-specific backbone networks 406 or the plurality of task-specific heads 602 while keeping the common backbone network 402 or the plurality of modality-specific backbone networks 404 unchanged.
[0169] Note that in various aspects, the internal architecture of the deep learning neural network 202 described herein can be regarded as a modular, white-box architecture, which can reduce waste of computing resources in scenarios where retraining or re-verification is required.
[0170] As a non-limiting example, assume that after training, the deep learning neural network 202 is able to perform all but a few of its (p)(q) inference tasks with sufficient accuracy (e.g., this may occur if the regulatory requirements applicable to those few inference tasks are prospectively changed to require higher accuracy). If the deep learning neural network 202 exhibits a fully connected, black box internal architecture, it will not be known which of the internal parameters of the deep learning neural network 202 are responsible for performing which of the (p)(q) inference tasks. Thus, in such a case, the entire deep learning neural network 202 would have to undergo retraining in order to increase the accuracy of the deep learning neural network 202 with respect to the few inference tasks among the (p)(q) inference tasks. This would be considered a waste of computational resources (e.g., time) on parts of the deep learning neural network 202 that are already able to perform their corresponding inference tasks with sufficient accuracy.
[0171] However, the internal architecture of the deep learning neural network 202 described herein can eliminate this need for full retraining. In particular, it can be known which of the plurality of task-specific backbones 406 and which of the plurality of task-specific heads 602 are responsible for performing the few inference tasks among the (p)(q) inference tasks (e.g., are associated therewith). Thus, in various cases, such responsible task-specific backbones and such responsible task-specific heads can be retrained while the remainder of the deep learning neural network 202 (e.g., the common backbone 402, the plurality of modality-specific backbones 404, the remainder of the task-specific backbones 406, and the remainder of the task-specific heads 602) is frozen or otherwise left unchanged. In this way, computational resources (e.g., time) need not be wasted in retraining parts of the deep learning neural network 202 that are already able to perform their corresponding inference tasks with sufficient accuracy.
[0172] As another non-limiting example, assume that after training, it is desired to teach the deep learning neural network 202 to perform a new inference task. That is, after training the deep learning neural network 202 to perform a total of (p)(q) inference tasks, it may be desired to configure the deep learning neural network 202 to perform a total of (p)(q)+1 inference tasks. If the deep learning neural network 202 exhibits a fully connected, black box internal architecture, this would not be possible without expanding the existing layers of the deep learning neural network 202 and then retraining the entire deep learning neural network 202. As described above, this would be considered a waste of computational resources (e.g., time) to retrain the deep learning neural network 202 to perform the inference tasks (e.g., the original (p)(q) inference tasks) that it is already able to perform with sufficient accuracy.
[0173] However, the internal architecture of the deep learning neural network 202 described herein can eliminate this need for complete retraining. In particular, a new task-specific backbone network can be inserted into the plurality of task-specific backbone networks 406, a new concatenation layer can be inserted into the plurality of concatenation layers 502 such that it is in series with the new task-specific backbone network, and a new task-specific head can be inserted into the plurality of task-specific heads 602 such that it is in series with the new concatenation layer. Additionally, new ground truth annotations corresponding to the new, desired inference task can be obtained (e.g., manually or via a new pre-trained teacher network). Thus, such new ground truth annotations can be used to train the new task-specific backbone network and the new task-specific head, and the remainder of the deep learning neural network 202 (e.g., the common backbone network 402, the plurality of modality-specific backbone networks 404, the remainder of the plurality of task-specific backbone networks 406, the remainder of the plurality of task-specific heads 602) can be frozen or otherwise kept unchanged. In this way, computational resources (e.g., time) need not be wasted in retraining the deep learning neural network 202 to perform tasks that it is already capable of performing with sufficient accuracy.
[0174] As another non-limiting example, assume that after training, it is desired to prevent the deep learning neural network 202 from performing one of its (p)(q) inference tasks. That is, after training the deep learning neural network 202 to perform a total of (p)(q) inference tasks, it may be desired to configure the deep learning neural network 202 to perform a total of (p)(q)-1 inference tasks. If the deep learning neural network 202 exhibits a fully connected, black box internal architecture, this would be impossible without shrinking the existing layers of the deep learning neural network 202 and then retraining the entire deep learning neural network 202. As described above, this would be considered a waste of computational resources.
[0175] However, the internal architecture of the deep learning neural network 202 described herein can eliminate this need for complete retraining. More specifically, a particular one of the (p)(q) inference tasks that is desired to be removed can correspond to a particular one of the plurality of task-specific backbone networks 406, and such a particular task-specific backbone network can be deleted, removed, or otherwise deactivated. Additionally, a particular one of the (p)(q) inference tasks that is desired to be removed can correspond to a particular one of the plurality of stitching layers 502, and such a particular stitching layer can be deleted, removed, or otherwise deactivated. Further still, a particular one of the (p)(q) inference tasks that is desired to be removed can correspond to a particular one of the plurality of task-specific heads 602, and such a particular task-specific head can be deleted, removed, or otherwise deactivated. In various aspects, such deletion, removal, or inactivation can cause the deep learning neural network 202 to no longer perform the particular inference task. However, such deletion, removal, or deactivation may have no effect on other parts of the deep learning neural network 202 (e.g., may have no effect on the common backbone network 402, the plurality of modality-specific backbone networks 404, the remainder of the plurality of task-specific backbone networks 406, the remainder of the plurality of stitching layers 502, or the remainder of the plurality of task-specific heads 602). In this way, an inference task can be selectively removed from the library of the deep learning neural network 202 without affecting how the deep learning neural network 202 performs other inference tasks.
[0176] Figures 10 to 14 Flowcharts of exemplary non-limiting computer-implemented methods 1000, 1100, 1200, 1300, and 1400 for training a deep learning neural network in accordance with one or more embodiments described herein are illustrated. In various cases, an image analysis system 102 can facilitate the computer-implemented methods 1000, 1100, 1200, 1300, and 1400.
[0177] First, consider Figure 10。In various embodiments, operation 1002 may include accessing a deep learning neural network (e.g., 202) via a device (e.g., via 110) operatively coupled to a processor (e.g., 106). In various aspects, the deep learning neural network may include a common backbone network (e.g., 402). In various cases, the deep learning neural network may include a plurality of modality-specific backbone networks (e.g., 404) in parallel with the common backbone network. In various cases, the deep learning neural network may include a plurality of task-specific backbone networks (e.g., 406) in parallel with the common backbone network. In various aspects, the deep learning neural network may include a plurality of splicing layers (e.g., 502) in series with the plurality of task-specific backbone networks, respectively. In various cases, the deep learning neural network may include a plurality of task-specific heads in series with the plurality of splicing layers, respectively. In various cases, a corresponding subset of the plurality of task-specific backbone networks may correspond to a corresponding one of the plurality of modality-specific backbone networks (e.g., 406(1) may correspond to 404(1)). In various aspects, each splicing layer may receive inputs from a corresponding task-specific backbone network, from the modality-specific backbone network corresponding to the corresponding task-specific backbone network, and from the common backbone network (e.g., 502(1)(1) may receive inputs from 406(1)(1), from 404(1), and from 402). In various cases, each task-specific head may receive inputs from a corresponding splicing layer (e.g., 602(1)(1) may receive inputs from 502(1)(1)).
[0178] In various cases, operation 1004 may include accessing a training data set (e.g., 704) via a device (e.g., via 110) and training the deep learning neural network on the training data set. In various aspects, the training data set may include a set of training inputs (e.g., 802) and a corresponding set of a plurality of ground truth annotations (e.g., 804). In various cases, each ground truth annotation may be generated or otherwise obtained from a corresponding teacher network.
[0179] In various cases, operation 1006 may include randomly initializing trainable internal parameters (e.g., convolutional kernels, weight matrices, bias values) of the deep learning neural network by the device (e.g., via 702).
[0180] In various aspects, operation 1008 may include determining, by the device (e.g., via 702), whether any training input in the training dataset has not yet been used to train the deep learning neural network. If so (e.g., if there is at least one training input that has not yet been used to train the deep learning neural network), computer-implemented method 1000 may proceed to operation 1102 of computer-implemented method 1100. Otherwise (e.g., if each training input has already been used to train the deep learning neural network), computer-implemented method 1000 may proceed to operation 1202 of computer-implemented method 1200.
[0181] Now, consider Figure 11 . In various embodiments, operation 1102 may include selecting, by the device (e.g., via 702) and from the training dataset, a training input (e.g., 902) that has not yet been used to train the deep learning neural network and a corresponding plurality of ground truth annotations (e.g., 904).
[0182] In various aspects, operation 1104 may include executing, by the device (e.g., via 702), the deep learning neural network on the training input such that the deep learning neural network produces a plurality of inference outputs (e.g., 906).
[0183] In various cases, operation 1106 may include computing, by the device (e.g., via 702), at least one loss between the plurality of inference outputs and the selected plurality of ground truth annotations. In various cases, the at least one loss may include a regularization term that is applied to the plurality of task-specific backbone networks but not to the common backbone network.
[0184] In various aspects, operation 1108 may include updating, by the device (e.g., via 702) and via backpropagation driven by the at least one loss, the trainable internal parameters of the common backbone network, the plurality of modality-specific backbone networks, the plurality of task-specific backbone networks, and the plurality of task-specific heads.
[0185] In various cases, computer-implemented method 1100 may proceed back to operation 1008.
[0186] Note that operations 1008 - 1108 may be iterated until each training input has been used to train the deep learning neural network. This may be considered the first stage of training.
[0187] Now, consider Figure 12。In various embodiments, operation 1202 may include determining by the device (e.g., via 702) whether any training input in the training dataset has not been used twice to train the deep learning neural network. If so (e.g., if there is at least one training input that has not been used twice to train the deep learning neural network), then computer-implemented method 1200 may proceed to operation 1204. Otherwise (e.g., if each training input has been used twice to train the deep learning neural network), then computer-implemented method 1200 may proceed to operation 1302 of computer-implemented method 1300.
[0188] In various aspects, operation 1204 may include selecting by the device (e.g., via 702) and from the training dataset a training input (e.g., 902) that has not been used twice to train the deep learning neural network and a corresponding plurality of ground truth annotations (e.g., 904).
[0189] In various aspects, operation 1206 may include executing by the device (e.g., via 702) the deep learning neural network on the training input such that the deep learning neural network produces a plurality of inference outputs (e.g., 906).
[0190] In various cases, operation 1208 may include computing by the device (e.g., via 702) at least one loss between the plurality of inference outputs and the selected plurality of ground truth annotations. In various cases, the at least one loss may lack a regularization term applied to the plurality of task-specific backbone networks. This is in contrast to the at least one loss computed at operation 1106.
[0191] In various aspects, operation 1210 may include updating by the device (e.g., via 702) and via backpropagation driven by the at least one loss the trainable internal parameters of the plurality of modality-specific backbone networks, the plurality of task-specific backbone networks, and the plurality of task-specific heads, but not updating the trainable internal parameters of the common backbone network. In other words, the common backbone network may be considered frozen. In some cases, the plurality of modality-specific backbone networks may also be frozen during this operation.
[0192] In various cases, computer-implemented method 1200 may proceed back to operation 1202.
[0193] Note that operations 1202 - 1210 may be iterated until each training input has been used twice to train the deep learning neural network. This may be considered the second phase of training.
[0194] Now, consider Figure 13。In various embodiments, operation 1302 may include determining, by the device (e.g., via 702), whether a new inference task capability is desired. If not (e.g., if it is not desired to teach a deep learning neural network how to perform a new inference task), the computer-implemented method 1300 may end at operation 1304. If so (e.g., if it is desired to teach a deep learning neural network how to perform a new inference task), the computer-implemented method 1300 may proceed to operation 1306.
[0195] In various aspects, operation 1306 may include: inserting, by the device (e.g., via 702), a new task-specific backbone network corresponding to the new inference task capability into the plurality of task-specific backbone networks; inserting, by the device (e.g., via 702), a new concatenation layer in series with the new task-specific backbone network into the plurality of concatenation layers; and inserting, by the device (e.g., via 702), a new task-specific head in series with the new concatenation layer into the plurality of task-specific heads. In various cases, the new task-specific backbone network may correspond to one of the plurality of modality-specific backbone networks. In various cases, the new concatenation layer may receive inputs from the new task-specific backbone network, from any modality-specific backbone network corresponding to the new task-specific backbone network, and from the common backbone network. In various aspects, the new task-specific head may receive an input from the new concatenation layer.
[0196] In various cases, operation 1308 may include inserting, by the device (e.g., 702), new ground truth annotations corresponding to the new inference task capability into each of the plurality of ground truth annotations in the training dataset. In some cases, these new ground truth annotations may be generated by a new teacher network that has been trained to perform the new inference task capability.
[0197] In various aspects, operation 1310 may include randomly initializing, by the device (e.g., via 702), the trainable internal parameters of the new task-specific backbone network and the new task-specific head, while keeping all other trainable parameters of the deep learning network unchanged (e.g., freezing the common backbone network, the remaining portions of the plurality of modality-specific backbone networks, the remaining portions of the plurality of task-specific backbone networks, and the remaining portions of the plurality of task-specific heads). As shown, the computer-implemented method 1300 may then proceed to operation 1402 of the computer-implemented method 1400.
[0198] Now, consider Figure 14。In various embodiments, operation 1402 may include determining, by the device (e.g., via 702), whether any training input in the training dataset has not been used to train a new task-specific backbone network and a new task-specific head (e.g., inserted at operation 1306). If not (e.g., if all training inputs have been used to train the new task-specific backbone network and the new task-specific head inserted at operation 1306), then the computer-implemented method 1400 may proceed back to operation 1302. If so (e.g., if there is at least one training input that has not been used to train the new task-specific backbone network and the new task-specific head inserted at operation 1306), then the computer-implemented method 1400 may proceed to operation 1404.
[0199] In various aspects, operation 1404 may include selecting, by the device (e.g., via 702) and from the training dataset, a training input (e.g., 902) that has not been used to train the new task-specific backbone network and the new task-specific head, and selecting, by the device (e.g., via 702) and from the training dataset, a corresponding plurality of ground truth annotations (e.g., 904). Note that, at this time, the selected plurality of ground truth annotations may have new ground truth annotations corresponding to the new inference task capabilities (e.g., due to operation 1308).
[0200] In various cases, operation 1406 may include performing, by the device (e.g., via 702), a deep learning neural network on the selected training input such that the deep learning neural network produces a plurality of inference outputs (e.g., 906).
[0201] In various cases, operation 1408 may include computing, by the device (e.g., via 702), at least one loss between the plurality of inference outputs and the selected plurality of ground truth annotations. In various aspects, the at least one loss may not have a regularization term applied to the new task-specific backbone network.
[0202] In various cases, operation 1410 may include updating, by the device (e.g., via 702) and via backpropagation driven by the at least one loss, the trainable internal parameters of the new task-specific backbone network and the new task-specific head, but not updating the trainable internal parameters of any other part of the deep learning neural network. In other words, the remaining part of the deep learning neural network may be frozen.
[0203] In various cases, the computer-implemented method 1400 may proceed back to operation 1402.
[0204] The present inventors have experimentally verified the various embodiments described herein. In particular, an exemplary non - limiting embodiment of a deep - learning neural network 202 was created, where such a deep - learning neural network was trained to perform tissue segmentation tasks, combined pneumothorax segmentation and classification tasks, and combined COVID segmentation and classification tasks on input X - ray images depicting a patient's chest cavity. Thus, such a deep - learning neural network has a common backbone network, three task - specific backbone networks (e.g., one per task), three concatenation layers (e.g., one per task), and three task - specific heads (e.g., one per task). Since the non - limiting deep - learning neural network focuses only on the X - ray imaging modality, it does not include modality - specific backbone networks. Such a deep - learning neural network was trained on a training dataset whose ground truth annotations were generated by a pre - trained tissue segmentation teacher network, a pre - trained pneumothorax segmentation and classification teacher network, and a pre - trained COVID segmentation and classification teacher network. Such a deep - learning neural network was trained in the two - stage manner described above, and the performance of such a deep - learning neural network was compared with the performance of the three teacher networks.
[0205] The deep - learning neural network consumes less total computer memory (e.g., has fewer total internal parameters) than the sum of the three teacher networks. In other words, although the deep - learning neural network has a smaller footprint (e.g., although it occupies less space) than the three teacher networks, it can perform the same three tasks as the three teacher networks. Additionally, despite such a smaller footprint, the deep - learning neural network also exhibits performance comparable to that of the three teacher networks. Specifically, with respect to pneumothorax segmentation and classification, the appropriate teacher network exhibits an accuracy score of 0.85, an area under the curve (AUC) of 0.94, and a Dice score of 0.72, while the deep - learning neural network exhibits an accuracy score of 0.86, an AUC of 0.93, and a Dice score of 0.66. Additionally, with respect to COVID classification and segmentation, the appropriate teacher network exhibits AUCs of 0.91, 0.98, and 0.98 for different classification cases, a COVID - specific Dice score of 0.59, and a pneumonia - specific Dice score of 0.54, while the deep - learning neural network exhibits 0.94, 0.98, 0.98, 0.57, and 0.53 respectively. Additionally, with respect to tissue segmentation, the appropriate teacher network exhibits Dice scores of 0.717, 0.617, and 0.767 for different tissue types, while the deep - learning neural network exhibits 0.701, 0.622, and 0.776 respectively. As demonstrated by these experimental results, the deep - learning neural network is capable of achieving accuracy comparable to that of the three teacher networks while occupying less space (e.g., having fewer internal parameters) than the combined three teacher networks.
[0206] Figure 15FIG. 1500 is a flowchart of an exemplary non - limiting computer - implemented method that illustrates one or more embodiments described herein. The method may facilitate deep - learning image analysis with increased modularity and reduced footprint. In various scenarios, an image analysis system 102 may facilitate the computer - implemented method 1500.
[0207] In various embodiments, operation 1502 may include accessing medical imaging data (e.g., 104) via a device (e.g., via 110) operatively coupled to a processor (e.g., 106).
[0208] In various aspects, operation 1504 may include performing a plurality of inference tasks on the medical imaging data via a device (e.g., via 112) and via the execution of a deep - learning neural network (e.g., 202), where the deep - learning neural network may include a common backbone network (e.g., 402) in parallel with a plurality of task - specific backbone networks (e.g., 406), and where the plurality of task - specific backbone networks may respectively correspond to the plurality of inference tasks.
[0209] Although not explicitly shown in Figure 15 the deep - learning neural network may also include a plurality of modality - specific backbone networks (e.g., 404) in parallel with the common backbone network and in parallel with the plurality of task - specific backbones, where a corresponding modality - specific backbone in the plurality of modality - specific backbones may correspond to a corresponding subset of the plurality of task - specific backbone networks (e.g., 404(p) may correspond to 406(p)).
[0210] Although not explicitly shown in Figure 15 the deep - learning neural network may also include a plurality of combination layers (e.g., 502) serially coupled to the plurality of task - specific backbone networks, respectively. In various aspects, the common backbone network may receive the medical imaging data as an input and may produce a first intermediate output (e.g., 408), the plurality of modality - specific backbone networks may receive the medical imaging data as an input and may produce a plurality of second intermediate outputs (e.g., 410), and the plurality of task - specific backbone networks may receive the medical imaging data as an input and may produce a plurality of third intermediate outputs (e.g., 412). In various scenarios, the plurality of combination layers may combine (e.g., in some cases, concatenate) a corresponding third intermediate output among the plurality of third intermediate outputs with the first intermediate output and with a corresponding second intermediate output among the plurality of second intermediate outputs (e.g., 502(p)(1) may combine 412(p)(1) with 410(p) and with 408), thereby producing a plurality of combinations (e.g., 504).
[0211] Although not shown explicitly in Figure 15Although not explicitly shown in, the deep learning neural network may further include a plurality of task - specific heads (e.g., 602) connected in series with a plurality of combination layers respectively. In various cases, the plurality of task - specific heads may receive a plurality of combinations as inputs and may generate a plurality of inference outputs (e.g., 204) respectively corresponding to the plurality of inference tasks.
[0212] Although not Figure 15 explicitly shown in, the computer - implemented method 1500 may further include presenting at least one of the plurality of inference outputs on a graphical user interface by a device (e.g., via 114).
[0213] Although not Figure 15 explicitly shown in, the deep learning neural network may be trained via a first stage and a second stage. The first stage may include training a common backbone and the plurality of task - specific backbone networks based on a regularization term applied to the plurality of task - specific backbone networks, and the second stage may include freezing the common backbone network and removing the regularization term. In various cases, training the deep learning neural network may occur in a supervised manner based on ground - truth annotations (e.g., 804) generated by a plurality of teacher networks. In various cases, a new task - specific backbone network may be added to the plurality of task - specific backbone networks, and the new task - specific backbone network may be trained while freezing the common backbone network and the remaining parts of the plurality of task - specific backbone networks.
[0214] The various embodiments described herein may be regarded as computerized tools for facilitating deep - learning image analysis with increased modularity and reduced footprint. A deep - learning neural network having an internal architecture as described herein is capable of performing a plurality of inference tasks with an accuracy comparable to that of a plurality of teacher networks. However, such a deep - learning neural network may consume fewer computational resources compared to such a plurality of teacher networks. In addition, discrete parts of the deep - learning neural network may be selectively retrained or re - validated without affecting other parts of the deep - learning neural network. This reduced computational resource consumption and this increased modularity / transparency are undoubtedly specific and practical technical advantages. Therefore, the various embodiments described herein undoubtedly constitute useful and practical applications of a computer.
[0215] Although the disclosure herein mainly describes various embodiments applied to medical imaging data, this is only a non - restrictive example. In various aspects, the teachings described herein may be extrapolated to any suitable type of electronic imaging data (e.g., not limited to imaging data in a medical / clinical context only).
[0216] In various examples, a machine learning algorithm or model can be implemented in any suitable manner to facilitate any suitable aspect described herein. To facilitate some of the machine learning aspects of the above-described machine learning aspects of the various implementations, consider the following discussion of artificial intelligence (AI). The various implementations described herein can employ artificial intelligence to facilitate the automation of one or more features or functions. These components can employ various AI-based schemes to perform the various implementations / examples disclosed herein. To provide or assist with the numerous determinations (e.g., determine, detect, infer, estimate, predict, prognose, assess, derive, forecast, detect, calculate) described herein, the components described herein can examine all or a subset of the data to which they are granted access and can provide reasoning about or determine the state of a system or environment from a set of observations captured via events or data. For example, determinations can be used to identify a particular context or action, or a probability distribution of a state can be generated. These determinations can be probabilistic; that is, the calculation of the probability distribution of the state of interest is based on the consideration of data and events. A determination can also refer to techniques for composing higher-level events from a set of events or data.
[0217] Such determinations can result in the construction of new events or actions from a set of observed events or stored event data, regardless of whether the events are temporally close and regardless of whether the events and data are from one or more events and data sources. The components disclosed herein can employ various classification (explicitly trained (e.g., via training data) and implicitly trained (e.g., via observed behavior, preferences, historical information, receiving extrinsic information, etc.)) schemes or systems (e.g., support vector machines, neural networks, expert systems, Bayesian belief networks, fuzzy logic, data fusion engines, etc.) in conjunction with performing automated or determined actions related to the claimed subject matter. Thus, classification schemes or systems can be used to automatically learn and perform a variety of functions, actions, or determinations.
[0218] A classifier can map an input attribute vector z = (z 1 , z 2 , z 3 , z 4 , z n)The confidence that the input maps to a certain class, such as according to f(z)=confidence(class). Such classification can use probability- or statistics-based analysis (e.g., analyzing utility and cost considerations) to determine the action to be automatically performed. Support Vector Machine (SVM) can be an example of a classifier that can be used. SVM operates by finding a hypersurface in the space of possible inputs, where the hypersurface attempts to separate the triggering criteria from non-triggering events. Intuitively, this makes the classification correct for testing data that is close to but different from the training data. Other directed and undirected model classification methods include, for example, Naive Bayes, Bayesian networks, decision trees, neural networks, fuzzy logic models, or any of the probability classification models that can provide different independent models. Classification as used herein also includes statistical regression for developing a priority model.
[0219] The disclosure herein describes non-limiting examples. For ease of description or explanation, when discussing various examples, various parts of the disclosure herein utilize the terms "each", "every", or "all". Such usage of the terms "each", "every", or "all" is non-limiting. In other words, when the disclosure herein provides a description of "each", "every", or "all" of a particular object or component applied to some particular objects or components, it should be understood that this is a non-limiting example, and it should also be understood that in various examples, it may be the case that such description applies to less than "each", "every", or "all" of the particular objects or components.
[0220] To provide additional context to the various embodiments described herein, Figure 16 and the following discussion is intended to provide a brief general description of a suitable computing environment 1600 in which the various embodiments described herein can be implemented. Although the embodiments have been described above in the general context of computer-executable instructions that can run on one or more computers, those skilled in the art will recognize that these embodiments can also be implemented in combination with other program modules or as a combination of hardware and software.
[0221] Generally, program modules include routines, programs, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In addition, those skilled in the art will understand that the methods of the present invention can be practiced with other computer system configurations, including single-processor or multi-processor computer systems, minicomputers, mainframe computers, Internet of Things (IoT) devices, distributed computing systems, and personal computers, handheld computing devices, microprocessor-based or programmable consumer electronics, etc., each of which is operably coupled to one or more associated devices.
[0222] The illustrated embodiments of the present implementation can also be practiced in a distributed computing environment where particular tasks are performed by remote processing devices linked through a communications network. In a distributed computing environment, program modules can be located in both local and remote memory storage devices.
[0223] Computing devices generally include a variety of media, which can include computer-readable storage media, machine-readable storage media, or communication media, where the use of these two terms is different from each other in this article, as described below. Computer-readable storage media or machine-readable storage media can be any available storage media accessible by a computer and include volatile and non-volatile media, removable and non-removable media. By way of example and not limitation, computer-readable storage media or machine-readable storage media can be implemented in conjunction with any method or technology for storing information such as computer-readable or machine-readable instructions, program modules, structured data, or unstructured data.
[0224] Computer-readable storage media can include, but are not limited to, random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), Blu-ray disc (BD) or other optical disc storage devices, magnetic tape cartridges, tapes, magnetic disk storage devices or other magnetic storage devices, solid state drives or other solid state storage devices, or other tangible or non-transitory media that can be used to store the desired information. In this regard, the terms "tangible" or "non-transitory" as applied to storage devices, memory, or computer-readable media should be understood to exclude only propagating transitory signals themselves as a modifier and not to forego the rights to all standard storage devices, memory, or computer-readable media that are not only propagating transitory signals themselves.
[0225] Computer-readable storage media can be accessed by one or more local or remote computing devices, for example, via an access request, query, or other data retrieval protocol, to perform various operations with respect to the information stored by the media.
[0226] Communication media typically contain computer-readable instructions, data structures, program modules, or other structured or unstructured data in a data signal, which can be, for example, a modulated data signal such as a carrier wave or other transmission mechanism, and include any information delivery or transmission medium. The term "modulated data signal" or "signal" refers to a signal that sets or changes one or more of its characteristics to encode information in one or more signals. By way of example and not limitation, communication media include wired media (such as a wired network or a direct wired connection) and wireless media (such as acoustic, RF, infrared, and other wireless media).
[0227] Refer again to Figure 16, An example environment 1600 for implementing various embodiments for the aspects described herein includes a computer 1602, which includes a processing unit 1604, a system memory 1606, and a system bus 1608. The system bus 1608 couples system components including, but not limited to, the system memory 1606 to the processing unit 1604. The processing unit 1604 can be any of a variety of commercially available processors. Dual microprocessors and other multiprocessor architectures can also be used as the processing unit 1604.
[0228] The system bus 1608 can be any of several types of bus structures that can be further interconnected to a memory bus (with or without a memory controller), a peripheral bus, and a local bus using any of a variety of commercially available bus architectures. The system memory 1606 includes a ROM 1610 and a RAM 1612. The basic input / output system (BIOS) can be stored in non-volatile memory such as ROM, erasable programmable read-only memory (EPROM), EEPROM, where the BIOS contains basic routines that assist in transferring information between elements within the computer 1602, such as during startup. The RAM 1612 can also include high-speed RAM, such as static RAM for caching data.
[0229] The computer 1602 also includes an internal hard disk drive (HDD) 1614 (e.g., EIDE, SATA), one or more external storage devices 1616 (e.g., a magnetic floppy disk drive (FDD) 1616, a memory stick or flash drive reader, a memory card reader, etc.), and a drive 1620, such as a solid-state drive, an optical disc drive, which can read from or write to a disk 1622 (such as a CD-ROM disc, a DVD, a BD, etc.). Alternatively, in the case of a solid-state drive, the disk 1622 will not be included unless separated. Although the internal HDD 1614 is shown as being within the computer 1602, the internal HDD 1614 can also be configured for external use in a suitable infrastructure (not shown). Additionally, although not shown in the environment 1600, a solid-state drive (SSD) can be used as a supplement or alternative to the HDD 1614. The HDD 1614, the external storage device 1616, and the drive 1620 can be connected to the system bus 1608 via an HDD interface 1624, an external storage interface 1626, and a drive interface 1628, respectively. The interface 1624 for external drive implementations can include at least one or both of the universal serial bus (USB) and the Institute of Electrical and Electronics Engineers (IEEE) 1394 interface technologies. Other external drive connection technologies are within the contemplation of the embodiments described herein.
[0230] The drive and its associated computer-readable storage medium provide non-volatile storage of data, data structures, computer-executable instructions, and the like. For computer 1602, the drive and storage medium are adapted to store any data in a suitable digital format. Although the above description of computer-readable storage media refers to corresponding types of storage devices, those skilled in the art should understand that other types of storage media that are computer-readable (whether currently existing or developed in the future) can also be used in the exemplary operating environment, and furthermore, any such storage media may contain computer-executable instructions for performing the methods described herein.
[0231] Multiple program modules may be stored in the drive and RAM 1612, including an operating system 1630, one or more application programs 1632, other program modules 1634, and program data 1636. All or part of the operating system, application programs, modules, or data may also be cached in RAM 1612. The systems and methods described herein may be implemented using a variety of commercially available operating systems or combinations of operating systems.
[0232] Computer 1602 may optionally include emulation technology. For example, a hypervisor (not shown) or other intermediary may emulate the hardware environment for operating system 1630, and the emulated hardware may optionally be different from Figure 16 that illustrated in. In such an implementation, operating system 1630 may include one VM among multiple virtual machines (VMs) hosted at computer 1602. Additionally, operating system 1630 may provide a runtime environment to application programs 1632, such as a Java runtime environment or a.NET framework. A runtime environment is a consistent execution environment that allows application programs 1632 to run on any operating system that includes the runtime environment. Similarly, operating system 1630 may support containers, and application programs 1632 may be in the form of containers that are lightweight, independent, executable software packages that include, for example, the code of the application program, the runtime, system tools, system libraries, and settings.
[0233] Furthermore, computer 1602 may be enabled using a security module such as a Trusted Platform Module (TPM). For example, in the case of a TPM, the boot component is hashed in the next boot component and waits for the result to match a security value before loading the next boot component. This process may occur at any layer in the code execution stack of computer 1602, such as applied to the application execution level or the operating system (OS) kernel level, thereby achieving security at any code execution level.
[0234] A user may input commands and information into the computer 1602 through one or more wired / wireless input devices (e.g., keyboard 1638, touch screen 1640, and pointing devices such as mouse 1642). Other input devices (not shown) may include a microphone, an infrared (IR) remote control, a radio frequency (RF) remote control or other remote controls, a joystick, a virtual reality controller or virtual reality headset, a gamepad, a stylus, an image input device (e.g., a camera), a gesture sensor input device, a visual motion sensor input device, an emotion or face detection device, a biometric input device (e.g., a fingerprint or iris scanner), etc. These input devices and other input devices are often connected to the processing unit 1604 through an input device interface 1644, which may be coupled to the system bus 1608, but these input devices and other input devices may be connected through other interfaces (such as a parallel port, an IEEE 1394 serial port, a game port, a USB port, an IR interface, interfaces, etc.).
[0235] A monitor 1646 or other type of display device may also be connected to the system bus 1608 via an interface (such as a video adapter 1648). In addition to the monitor 1646, a computer typically also includes other peripheral output devices (not shown), such as speakers, printers, etc.
[0236] The computer 1602 may operate in a networked environment using a logical connection to one or more remote computers (such as remote computer 1650) via wired or wireless communication. The remote computer 1650 may be a workstation, a server computer, a router, a personal computer, a portable computer, a microprocessor-based entertainment appliance, a peer device, or other common network nodes, and typically includes many or all of the elements described with respect to the computer 1602, but for simplicity, only the memory / storage device 1652 is illustrated. The depicted logical connections include a wired / wireless connection to a local area network (LAN) 1654 or a larger network (e.g., a wide area network (WAN) 1656). Such LAN and WAN networking environments are common in offices and companies and facilitate enterprise-wide computer networks (such as intranets), all of which may be connected to a global communication network (e.g., the Internet).
[0237] When used in a LAN networking environment, the computer 1602 may be connected to the local network 1654 through a wired or wireless communication network interface or adapter 1658. The adapter 1658 may facilitate wired or wireless communication with the LAN 1654, which may also include a wireless access point (AP) disposed thereon for communicating with the adapter 1658 in wireless mode.
[0238] When used in a WAN networking environment, computer 1602 may include a modem 1660 or may be connected to a communication server on WAN 1656 via other devices for establishing communication over the WAN 1656 (such as via the Internet). The modem 1660, which can be internal or external and a wired or wireless device, may be connected to the system bus 1608 via the input device interface 1644. In a networking environment, program modules depicted relative to computer 1602 or portions thereof may be stored in the remote memory / storage device 1652. It should be understood that the network connections shown are examples and other means of establishing a communication link between computers may be used.
[0239] When used in a LAN or WAN networking environment, in addition to or instead of the external storage device 1616 as described above, computer 1602 may access a cloud storage system or other network-based storage systems, such as but not limited to network virtual machines that provide one or more aspects of information storage or processing. Generally, the connection between computer 1602 and the cloud storage system may be established, for example, via adapter 1658 or modem 1660 over LAN 1654 or WAN 1656, respectively. When connecting computer 1602 to an associated cloud storage system, the external storage interface 1626 may manage the storage provided by the cloud storage system, just as with other types of external storage devices, by means of adapter 1658 or modem 1660. For example, the external storage interface 1626 may be configured to provide access to cloud storage sources as if those cloud storage sources were physically connected to computer 1602.
[0240] Computer 1602 may be capable of operating to communicate with any wireless device or entity operatively disposed to communicate wirelessly, such as a printer, scanner, desktop or portable computer, portable data assistant, communication satellite, any equipment or location associated with a wirelessly detectable tag (such as a kiosk, newsstand, store shelf, etc.), and a telephone. This may include Wi-Fi and wireless technologies. Thus, the communication may be of a predefined structure like a conventional network or merely an ad hoc communication between at least two devices.
[0241] Figure 17FIG. 1700 is a schematic block diagram of an example computing environment 1700 with which the disclosed subject matter may interact. The example computing environment 1700 includes one or more clients 1710. The clients 1710 can be hardware or software (e.g., threads, processes, computing devices). The example computing environment 1700 also includes one or more servers 1730. The servers 1730 can also be hardware or software (e.g., threads, processes, computing devices). For example, the server 1730 can house threads to perform transformations by adopting one or more embodiments as described herein. A possible communication between the clients 1710 and the servers 1730 can be in the form of data packets suitable for transfer between two or more computer processes. The example computing environment 1700 includes a communication framework 1750 that can be used to facilitate communications between the clients 1710 and the servers 1730. The clients 1710 are operatively connected to one or more client data repositories 1720 that can be used to store information local to the clients 1710. Similarly, the servers 1730 are operatively connected to one or more server data repositories 1740 that can be used to store information local to the servers 1730.
[0242] The present invention can be a system, a method, an apparatus, or a computer program product at any possible technical detail integration level. The computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to execute aspects of the present invention. The computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer-readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer-readable storage medium can also include the following: a portable computer floppy disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in grooves recorded with instructions, and any suitable combination of the foregoing items. As used herein, a computer-readable storage medium should not be construed to be a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., light pulses passing through an optical fiber cable), or an electrical signal transmitted through a wire.
[0243] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or downloaded to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, or a wireless network). The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the corresponding computing / processing device. The computer-readable program instructions for carrying out operations of the present invention may be assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may establish a connection with an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, an electronic circuit, including, for example, a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), can execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit so as to perform aspects of the present invention.
[0244] Aspects of the present invention are described below with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions. These computer-readable program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more blocks of the flowchart or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, a programmable data processing apparatus, or other devices to function in a particular manner, such that the computer-readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the functions / acts specified in one or more blocks of the flowchart or block diagram. The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other devices to produce a computer-implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other devices implement the functions / acts specified in one or more blocks of the flowchart or block diagram.
[0245] The flowchart and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by a special purpose hardware-based system that performs the specified functions or acts, or combinations of special purpose hardware and computer instructions.
[0246] Although the subject matter has been described above in the general context of computer-executable instructions of a computer program product that runs on one or more computers, those skilled in the art will recognize that the present disclosure may also be implemented in whole or in part in conjunction with other program modules. Generally, program modules include routines, programs, components, data structures, etc. that perform particular tasks or implement particular abstract data types. In addition, those skilled in the art will appreciate that the computer-implemented methods of the present invention may be practiced using other computer system configurations, including single-processor or multi-processor computer systems, minicomputers, mainframe computers, as well as computers, hand-held computing devices (such as PDAs, telephones), microprocessor-based or programmable consumer or industrial electronics, and the like. The illustrated aspects may also be practiced in a distributed computing environment where tasks are performed by remote processing devices linked through a communications network. However, some (if not all) aspects of the present disclosure may be practiced on stand-alone computers. In a distributed computing environment, program modules may be located in local and remote memory storage devices.
[0247] As used in this application, the terms "component", "system", "platform", "interface", etc. may refer to or may include a computer-related entity or an entity related to an operating machine having one or more specific functionalities. Entities disclosed herein may be hardware, a combination of hardware and software, software, or software in execution. For example, a component may be, but is not limited to, a program running on a processor, a processor, an object, an executable file, an execution thread, a program, or a computer. By way of illustration, both an application running on a server and the server can be components. One or more components may reside within a process or execution thread, and a component may be located on one computer or distributed between two or more computers. As another example, a corresponding component may execute according to various computer-readable media storing various data structures. Components may communicate via local or remote processes, such as in accordance with signals having one or more data packets (e.g., data from one component that interacts with another component in a local system, a distributed system, or a network such as the Internet with other systems) via a signal. As another example, a component may be a device having specific functionality provided by mechanical parts operated by an electrical or electronic circuit, which is operated by a software or firmware application executed by a processor. In such cases, the processor may be inside or outside the device and may execute at least a portion of the software or firmware application. As yet another example, a component may be a device that provides specific functionality through electronic components rather than mechanical parts, where the electronic components may include a processor or other components for executing software or firmware that at least partially imparts functionality to the electronic components. In one aspect, a component may be emulated, for example, via a virtual machine within a cloud computing system, as an electronic component.
[0248] In addition, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless otherwise specified or clear from the context, "X employs A or B" is intended to mean any natural inclusive permutation. That is, if X employs A; X employs B; or X employs both A and B, then "X employs A or B" is satisfied in any of the foregoing instances. As used herein, the term "and / or" is intended to have the same meaning as "or". In addition, unless otherwise specified or clear from the context that it is for the singular form, the articles "a" and "an" used in this specification and the drawings shall generally be understood to mean "one or more". As used herein, the terms "example" or "exemplary" are used to mean serving as an example, instance, or illustration. To avoid doubt, the subject matter disclosed herein is not limited by such examples. Additionally, any aspect or design described herein as "example" or "exemplary" should not necessarily be understood as more preferred or advantageous than other aspects or designs, nor does it mean excluding equivalent exemplary structures and techniques known to those of ordinary skill in the art.
[0249] As used in this specification, the term "processor" can generally refer to any computing processing unit or device, including but not limited to a single-core processor; a single processor with software multithreaded execution capabilities; a multi-core processor; a multi-core processor with software multithreaded execution capabilities; a multi-core processor with hardware multithreaded technology; a parallel platform; and a parallel platform with distributed shared memory. Additionally, a processor can refer to an integrated circuit, an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a field-programmable gate array (FPGA), a programmable logic controller (PLC), a complex programmable logic device (CPLD), discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. Further, a processor can utilize nanoscale architectures (such as but not limited to molecule- and quantum-dot-based transistors, switches, and gates) in order to optimize space usage or enhance the performance of user equipment. A processor can also be implemented as a combination of computing processing units. In the present disclosure, terms such as "repository", "storage device", "data repository", "data storage device", "database", and substantially any other information storage component related to the operation and functionality of a component are used to refer to a "memory component", an entity embodied in a "memory", or a component that includes a memory. It should be understood that the memory or memory components described herein can be volatile memory or non-volatile memory, or can include both volatile memory and non-volatile memory. By way of illustration and not limitation, non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable ROM (EEPROM), flash memory, or non-volatile random access memory (RAM) (e.g., ferroelectric RAM (FeRAM)). For example, volatile memory can include RAM that can act as an external cache memory. By way of illustration and not limitation, RAM can be provided in various forms (such as synchronous RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), direct Rambus RAM (DRRAM), direct Rambus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM)). Additionally, the disclosed memory components of the systems or computer-implemented methods herein are intended to include but not limited to include these and any other suitable types of memory.
[0250] The foregoing description includes only examples of systems and computer-implemented methods. Of course, it is not possible to describe every conceivable combination of components or computer-implemented methods for the purposes of describing the present disclosure, but many other combinations and permutations of the present disclosure are possible. Additionally, to the extent that the terms "comprising," "having," "owning," etc. are used in the detailed description, the claims, the appendices, and the drawings, such terms are intended to be inclusive in a manner similar to the term "including" as that term is interpreted when used as a transitional word in a claim.
[0251] The descriptions of the various embodiments have been presented for purposes of illustration, but are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent without departing from the scope and spirit of the described embodiments. The terms used herein were chosen to best explain the principles of the embodiments, the practical application, or technical improvements over technologies found in the marketplace, or to enable others of ordinary skill in the art to understand the embodiments disclosed herein.
[0252] Various non-limiting aspects are described in the following clauses.
[0253] Clause 1: A system, the system comprising: a processor that executes computer-executable components stored in a non-transitory computer-readable memory, the computer-executable components comprising: an access component that accesses medical imaging data; and an inference component that performs a plurality of inference tasks on the medical imaging data via execution of a deep learning neural network, wherein the deep learning neural network includes a common backbone network in parallel with a plurality of task-specific backbone networks, and wherein the plurality of task-specific backbone networks respectively correspond to the plurality of inference tasks.
[0254] Clause 2: The system according to any of the preceding clauses, wherein the deep learning neural network further includes a plurality of modality-specific backbone networks in parallel with the common backbone network and in parallel with the plurality of task-specific backbone networks, wherein a corresponding modality-specific backbone network of the plurality of modality-specific backbone networks corresponds to a corresponding subset of the plurality of task-specific backbone networks.
[0255] Clause 3: The system according to any of the preceding clauses, wherein the deep learning neural network further comprises a plurality of combination layers connected in series with the plurality of task-specific backbone networks respectively, wherein the common backbone network receives the medical imaging data as input and generates a first intermediate output, wherein the plurality of modality-specific backbone networks receive the medical imaging data as input and generate a plurality of second intermediate outputs, wherein the plurality of task-specific backbone networks receive the medical imaging data as input and generate a plurality of third intermediate outputs, and wherein the plurality of combination layers combine a corresponding third intermediate output among the plurality of third intermediate outputs with the first intermediate output and a corresponding second intermediate output among the plurality of second intermediate outputs, thereby generating a plurality of combinations.
[0256] Clause 4: The system according to any of the preceding clauses, wherein the deep learning neural network further comprises a plurality of task-specific heads connected in series with the plurality of combination layers respectively, wherein the plurality of task-specific heads receive the plurality of combinations as input and generate a plurality of inference outputs respectively corresponding to the plurality of inference tasks.
[0257] Clause 5: The system according to any of the preceding clauses, wherein the computer-executable component further comprises: a result component that presents at least one of the plurality of inference outputs on a graphical user interface.
[0258] Clause 6: The system according to any of the preceding clauses, wherein the deep learning neural network is trained via a first stage and a second stage, wherein the first stage includes training the common backbone network and the plurality of task-specific backbone networks based on a regularization term applied to the plurality of task-specific backbone networks, and wherein the second stage includes freezing the common backbone network and removing the regularization term.
[0259] Clause 7: The system according to any of the preceding clauses, wherein the deep learning neural network is trained in a supervised manner based on ground truth annotations generated by a plurality of teacher networks.
[0260] Clause 8: The system according to any of the preceding clauses, wherein a new task-specific backbone network is added to the plurality of task-specific backbone networks, and wherein the new task-specific backbone network is trained while freezing the common backbone network and the remaining portions of the plurality of task-specific backbone networks.
[0261] In various aspects, any one or more combinations of Clauses 1 to 8 can be implemented.
[0262] Clause 9: A computer-implemented method, the method comprising: accessing medical imaging data by a device operatively coupled to a processor; and performing a plurality of inference tasks on the medical imaging data by the device and via execution of a deep learning neural network, wherein the deep learning neural network includes a common backbone network in parallel with a plurality of task-specific backbone networks, and wherein the plurality of task-specific backbone networks respectively correspond to the plurality of inference tasks.
[0263] Clause 10: The computer-implemented method according to any of the preceding clauses, wherein the deep learning neural network further includes a plurality of modality-specific backbone networks in parallel with the common backbone network and in parallel with the plurality of task-specific backbone networks, wherein a respective modality-specific backbone network of the plurality of modality-specific backbone networks corresponds to a respective subset of the plurality of task-specific backbone networks.
[0264] Clause 11: The computer-implemented method according to any of the preceding clauses, wherein the deep learning neural network further includes a plurality of combination layers respectively in series with the plurality of task-specific backbone networks, wherein the common backbone network receives the medical imaging data as an input and produces a first intermediate output, wherein the plurality of modality-specific backbone networks receive the medical imaging data as an input and produce a plurality of second intermediate outputs, wherein the plurality of task-specific backbone networks receive the medical imaging data as an input and produce a plurality of third intermediate outputs, and wherein the plurality of combination layers combine a respective third intermediate output of the plurality of third intermediate outputs with the first intermediate output and a respective second intermediate output of the plurality of second intermediate outputs, thereby producing a plurality of combinations.
[0265] Clause 12: The computer-implemented method according to any of the preceding clauses, wherein the deep learning neural network further includes a plurality of task-specific heads respectively in series with the plurality of combination layers, wherein the plurality of task-specific heads receive the plurality of combinations as inputs, and produce a plurality of inference outputs respectively corresponding to the plurality of inference tasks.
[0266] Clause 13: The computer-implemented method according to any of the preceding clauses, further comprising: presenting, by the device, at least one of the plurality of inference outputs on a graphical user interface.
[0267] Clause 14: The computer-implemented method according to any of the preceding clauses, wherein the deep learning neural network is trained via a first phase and a second phase, wherein the first phase includes training the common backbone network and the plurality of task-specific backbone networks based on a regularization term applied to the plurality of task-specific backbone networks, and wherein the second phase includes freezing the common backbone network and removing the regularization term.
[0268] Clause 15: A computer-implemented method according to any of the preceding clauses, wherein the deep learning neural network is trained in a supervised manner based on ground truth annotations generated by a plurality of teacher networks.
[0269] Clause 16: A computer-implemented method according to any of the preceding clauses, wherein a new task-specific backbone network is added to the plurality of task-specific backbone networks, and wherein the new task-specific backbone network is trained while freezing the common backbone network and the remainder of the plurality of task-specific backbone networks.
[0270] In various aspects, any one or more combinations of Clauses 9 to 16 can be implemented.
[0271] Clause 17: A computer program product for facilitating deep learning image analysis with increased modularity and reduced footprint, the computer program product comprising a non-transitory computer-readable memory having program instructions contained therein, the program instructions being executable by a processor associated with one or more medical imaging devices to cause the processor to perform the following operations: access imaging data generated by the one or more medical imaging devices; generate a plurality of inference outputs based on the imaging data via execution of a deep learning neural network, wherein the deep learning neural network includes a shared backbone network in parallel with a plurality of task-specific backbone networks, and wherein the plurality of task-specific backbone networks respectively correspond to the plurality of inference outputs; and present at least one of the plurality of inference outputs on an electronic display of the one or more medical imaging devices.
[0272] Clause 18: A computer program product according to any of the preceding clauses, wherein the deep learning neural network further includes a plurality of modality-specific backbone networks in parallel with the shared backbone network and in parallel with the plurality of task-specific backbone networks, wherein a respective modality-specific backbone network of the plurality of modality-specific backbone networks corresponds to a respective subset of the plurality of task-specific backbone networks.
[0273] Clause 19: A computer program product according to any of the preceding clauses, wherein the deep learning neural network further includes a plurality of concatenation layers connected in series with the plurality of task-specific backbone networks, wherein the shared backbone network receives the imaging data as input and produces a first intermediate output, wherein the plurality of modality-specific backbone networks receive the imaging data as input and produce a plurality of second intermediate outputs, wherein the plurality of task-specific backbone networks receive the imaging data as input and produce a plurality of third intermediate outputs, and wherein the plurality of concatenation layers concatenate a respective third intermediate output of the plurality of third intermediate outputs with the first intermediate output and a respective second intermediate output of the plurality of second intermediate outputs, thereby producing a plurality of concatenation results.
[0274] Clause 20: A computer program product as described in any of the foregoing clauses, wherein the deep learning neural network further comprises a plurality of task-specific heads connected in series with the plurality of splicing layers respectively, wherein the plurality of task-specific heads receive the plurality of splicing results as inputs and generate the plurality of inference outputs.
[0275] In various aspects, any one or more combinations of Clauses 17 to 20 can be implemented.
[0276] In various aspects, any one or more combinations of Clauses 1 to 20 can be implemented.
Claims
1. A system, the system comprises: a processor that executes computer-executable components stored in a non-transitory computer-readable memory, the computer-executable components comprising: an access component that accesses medical imaging data; and an inference component that performs a plurality of inference tasks on the medical imaging data via execution of a deep learning neural network, wherein the deep learning neural network includes a common backbone network in parallel with a plurality of task-specific backbone networks, and wherein the plurality of task-specific backbone networks respectively correspond to the plurality of inference tasks.
2. The system according to claim 1, wherein the deep learning neural network further includes a plurality of modality-specific backbone networks in parallel with the common backbone network and in parallel with the plurality of task-specific backbone networks, wherein a corresponding modality-specific backbone network among the plurality of modality-specific backbone networks corresponds to a respective subset of the plurality of task-specific backbone networks.
3. The system according to claim 2, wherein the deep learning neural network further includes a plurality of combination layers respectively in series with the plurality of task-specific backbone networks, wherein the common backbone network receives the medical imaging data as an input and generates a first intermediate output, wherein the plurality of modality-specific backbone networks receive the medical imaging data as an input and generate a plurality of second intermediate outputs, wherein the plurality of task-specific backbone networks receive the medical imaging data as an input and generate a plurality of third intermediate outputs, and wherein the plurality of combination layers combine a respective third intermediate output among the plurality of third intermediate outputs with the first intermediate output and a respective second intermediate output among the plurality of second intermediate outputs, thereby generating a plurality of combinations.
4. The system according to claim 3, wherein the deep learning neural network further includes a plurality of task-specific heads respectively in series with the plurality of combination layers, wherein the plurality of task-specific heads receive the plurality of combinations as inputs, and generate a plurality of inference outputs respectively corresponding to the plurality of inference tasks.
5. The system according to claim 4, wherein the computer-executable components further comprise: a result component that presents at least one of the plurality of inference outputs on a graphical user interface.
6. The system according to claim 1, wherein the deep learning neural network is trained via a first phase and a second phase, wherein the first phase includes training the common backbone network and the plurality of task-specific backbone networks based on a regularization term applied to the plurality of task-specific backbone networks, and wherein the second phase includes freezing the common backbone network and removing the regularization term.
7. The system according to claim 6, wherein the deep learning neural network is trained in a supervised manner based on ground truth annotations generated by a plurality of teacher networks.
8. The system according to claim 6, wherein a new task-specific backbone network is added to the plurality of task-specific backbone networks, wherein the new task-specific backbone network is trained while freezing the common backbone network and the remainder of the plurality of task-specific backbone networks.
9. A computer-implemented method, the computer-implemented method comprising: accessing medical imaging data via a device operatively coupled to a processor; and performing a plurality of inference tasks on the medical imaging data via the device and via execution of a deep learning neural network, wherein the deep learning neural network includes a common backbone network in parallel with a plurality of task-specific backbone networks, and wherein the plurality of task-specific backbone networks respectively correspond to the plurality of inference tasks.
10. The computer-implemented method according to claim 9, wherein the deep learning neural network further includes a plurality of modality-specific backbone networks in parallel with the common backbone network and in parallel with the plurality of task-specific backbone networks, wherein a respective modality-specific backbone network of the plurality of modality-specific backbone networks corresponds to a respective subset of the plurality of task-specific backbone networks.
11. The computer-implemented method according to claim 10, wherein the deep learning neural network further includes a plurality of combination layers respectively in series with the plurality of task-specific backbone networks, wherein the common backbone network receives the medical imaging data as an input and generates a first intermediate output, wherein the plurality of modality-specific backbone networks receive the medical imaging data as an input and generate a plurality of second intermediate outputs, wherein the plurality of task-specific backbone networks receive the medical imaging data as an input and generate a plurality of third intermediate outputs, and wherein the plurality of combination layers combine a respective third intermediate output of the plurality of third intermediate outputs with the first intermediate output and a respective second intermediate output of the plurality of second intermediate outputs, thereby generating a plurality of combinations.
12. The computer-implemented method according to claim 11, wherein the deep learning neural network further includes a plurality of task-specific heads respectively in series with the plurality of combination layers, wherein the plurality of task-specific heads receive the plurality of combinations as inputs, and generate a plurality of inference outputs respectively corresponding to the plurality of inference tasks.
13. The computer-implemented method according to claim 12, the computer-implemented method further comprising: presenting at least one of the plurality of inference outputs on a graphical user interface via the device.
14. The computer-implemented method according to claim 9, wherein the deep learning neural network is trained via a first phase and a second phase, wherein the first phase includes training the common backbone network and the plurality of task-specific backbone networks based on a regularization term applied to the plurality of task-specific backbone networks, and wherein the second phase includes freezing the common backbone network and removing the regularization term.
15. The computer-implemented method according to claim 14, wherein the deep learning neural network is trained in a supervised manner based on ground truth annotations generated by a plurality of teacher networks.
16. The computer-implemented method according to claim 14, wherein a new task-specific backbone network is added to the plurality of task-specific backbone networks, and wherein the new task-specific backbone network is trained while freezing the common backbone network and the remainder of the plurality of task-specific backbone networks.
17. A computer program product for facilitating deep learning image analysis with increased modularity and reduced footprint, the computer program product comprising a non-transitory computer-readable memory having program instructions embodied therein, the program instructions being executable by a processor associated with one or more medical imaging devices to cause the processor to perform the following operations: Access imaging data generated by the one or more medical imaging devices; Generate a plurality of inference outputs based on the imaging data via execution of a deep learning neural network, wherein the deep learning neural network includes a shared backbone network in parallel with a plurality of task-specific backbone networks, and wherein the plurality of task-specific backbone networks respectively correspond to the plurality of inference outputs; and Present at least one of the plurality of inference outputs on an electronic display of the one or more medical imaging devices.
18. The computer program product according to claim 17, wherein the deep learning neural network further includes a plurality of modality-specific backbone networks in parallel with the shared backbone network and in parallel with the plurality of task-specific backbone networks, wherein a respective modality-specific backbone network of the plurality of modality-specific backbone networks corresponds to a respective subset of the plurality of task-specific backbone networks.
19. The computer program product according to claim 18, wherein the deep learning neural network further includes a plurality of concatenation layers connected in series with the plurality of task-specific backbone networks, wherein the shared backbone network receives the imaging data as input and produces a first intermediate output, wherein the plurality of modality-specific backbone networks receive the imaging data as input and produce a plurality of second intermediate outputs, wherein the plurality of task-specific backbone networks receive the imaging data as input and produce a plurality of third intermediate outputs, and wherein the plurality of concatenation layers concatenate a respective third intermediate output of the plurality of third intermediate outputs with the first intermediate output and a respective second intermediate output of the plurality of second intermediate outputs, thereby producing a plurality of concatenation results.
20. The computer program product according to claim 19, wherein the deep learning neural network further includes a plurality of task-specific heads connected in series with the plurality of concatenation layers, wherein the plurality of task-specific heads receive the plurality of concatenation results as input and produce the plurality of inference outputs.