Tile-based image augmentation for neural networks
By sharding medical image blocks and applying variable image augmentation technology, neural network analysis is performed after generating augmented image blocks, the performance degradation of neural networks under different imaging systems is solved, and image segmentation and classification accuracy is improved.
Patent Information
- Application Number
- CN202380084479.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-09
- Filing Date
- 2023-12-08
- Publication Date
- 2025-07-18
AI Technical Summary
Existing neural networks have problems with degradation in segmentation or classification accuracy when processing medical images acquired by different imaging systems, especially when the scanner protocol or imaging system changes.
The image blocks are divided into multiple image blocks, each of which has a size, shape and position. Each block is augmented through variable image augmentation technology to generate an augmented image block, and then analyzed using a neural network.
It improves the segmentation and classification accuracy of neural networks when processing images acquired by different imaging systems, and reduces performance fluctuations caused by differences in imaging systems.
Smart Images

Figure CN120344989A_ABST
Abstract
Description
Technical Field
[0001] The various embodiments generally relate to computing technology. More specifically, the various embodiments relate to patch-based image augmentation for neural networks. Background Art
[0002] Today, neural networks are state-of-the-art for a wide range of medical image analysis tasks such as segmentation, detection, or classification. A typical problem with neural networks is their performance, e.g., in terms of segmentation or classification accuracy, when applying them to medical images (e.g., magnetic resonance (MR) or computed tomography (CT) images) that have been acquired with slightly varying scanner protocols, MR systems with different field strengths, or imaging systems from different vendors, compared to the image data used for training, etc. Specifically, a performance degradation can be observed in such cases. Attempts have been made to solve this problem via augmentation techniques where histogram transformation, filtering, GANs, etc. are used to modify the style of the images. However, even with current augmentation techniques, the performance of an image analysis algorithm based on a neural network, such as segmentation, still degrades when applied to images acquired with systems from other vendors. Summary of the Invention
[0003] Therefore, there is a need for improved image augmentation for training and inference applications. The object of the disclosed technology is solved by the subject matter of the appended independent claims, where further embodiments are incorporated in the dependent claims, the drawings, and the following description.
[0004] Disclosed herein are improved computing systems, methods, and computer-readable media for providing patch-based image augmentation for neural networks. According to one or more embodiments, a computer-implemented method includes: dividing an image patch into a plurality of image slices, each image slice having a size, shape, and position relative to the image patch; generating an augmented image patch by applying variable image augmentation to each image slice, wherein at least two augmented image slices have different augmentations; and performing an image analysis task by applying a neural network to the augmented image patch.
[0005] According to one or more embodiments, a computer-implemented system includes a processor and a memory coupled to the processor, the memory including instructions that, when executed by the processor, cause the computing system to perform operations including: dividing an image patch into a plurality of image slices, each image slice having a size, shape, and position relative to the image patch; generating an augmented image patch by applying variable image augmentation to each image slice, wherein at least two augmented image slices have different augmentations; and performing an image analysis task by applying a neural network to the augmented image patch.
[0006] According to one or more embodiments, a computer-implemented method includes: dividing each of a plurality of image patches into a plurality of image tiles, each image tile having a size, a shape, and a position relative to the corresponding image patch; generating a plurality of augmented image patches by applying variable image augmentation to each image tile of an image patch for each image patch, wherein, for at least one augmented image patch, at least two augmented image tiles have different augmentations; and training a neural network using the plurality of augmented image patches as training images. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] The various advantages of the embodiments will become apparent to those skilled in the art from reading the following specification and the appended claims, and by reference to the following drawings, in which:
[0008] Figure 1A A block diagram is provided showing an example of a tile-based image augmentation system in accordance with one or more embodiments;
[0009] Figure 1B A diagram is provided showing an example of an image patch used in one or more embodiments;
[0010] Figure 2 A diagram is provided showing an example process of dividing an image patch into image tiles in accordance with one or more embodiments;
[0011] Figures 3A - 3B A diagram is provided showing an example process of generating augmented image patches in accordance with one or more embodiments;
[0012] Figure 4 A diagram is provided showing an example of an augmentation generator in accordance with one or more embodiments;
[0013] Figures 5A - 5B A flowchart is provided showing an example method of tile-based image augmentation for image analysis in accordance with one or more embodiments;
[0014] Figure 6 A block diagram is provided showing an example of a tile-based image augmentation training system in accordance with one or more embodiments;
[0015] Figure 7 A flowchart is provided showing an example method of tile-based image augmentation training in accordance with one or more embodiments; and
[0016] Figure 8 is a diagram showing an example of a computing system used in a tile-based image augmentation system in accordance with one or more embodiments. DETAILED DESCRIPTION
[0017] This disclosure relates to improved computing systems, methods, and computer-readable media for providing tile-based image augmentation for a neural network. As described herein, in an embodiment, the technique operates to divide an image patch into a plurality of image tiles, each image tile having a size, shape, and position relative to the image patch, generate an augmented image patch by applying variable image augmentation to each image tile, wherein at least two of the augmented image patches have different augmentations, and perform an image analysis task by applying a neural network to the augmented image patch. In an embodiment, the technique operates to divide each of a plurality of image patches into a plurality of image tiles, each image tile having a size, shape, and position relative to the respective image patch, generate a plurality of augmented image patches by applying variable image augmentation to each image tile for each image patch, wherein for at least one of the augmented image patches, at least two of the augmented image tiles have different augmentations, and use the plurality of augmented image patches as training images to train a neural network. The technique helps improve the overall performance of neural network-based image analysis by providing enhanced performance for image analysis tasks such as segmentation, detection, or classification, while reducing variance based on differences in the source of the images used for training and / or the images to be processed by the trained system.
[0018] Referring to the components and features described herein, including but not limited to the drawings and associated description, Figure 1A A block diagram of a tile-based image augmentation system 100 is provided that illustrates an example in accordance with one or more embodiments. As Figure 1A shown, the system 100 receives and processes an image patch 110. As described herein (and as Figure 1B shown), an image patch such as image patch 110 refers to a complete image and / or a sub-image of a complete image. The image can include, for example, an image generated by a diagnostic or other medical imaging system, such as an MR image, a CT image, an X-ray, and / or an image obtained via other imaging techniques (such as ultrasound), or an image generated by an imaging sensor, such as a visible or infrared image, etc. The image (and image patch) can have varying dimensions, including, for example, two-dimensional (2D) or three-dimensional (3D) images. As further shown in FIG. 1, the system 100 includes a tile generator 120, an augmentation generator 130, and a neural network 140. The system 100 produces an image analysis result 160. It should be understood that in some embodiments, the system 100 can include more, alternative, or fewer components than those Figure 1A shown, and in some embodiments, some components can be combined with or incorporated into other components.
[0019] The tile generator 120 obtains an image patch (e.g., image patch 110) and divides the image patch into a plurality of smaller units called tiles, such as each image tile 125. Each image tile has a size, shape, and position relative to the image patch 110. Reference is made herein toFigure 2 Further details are provided regarding the generation of image patches via the patch generator 120.
[0020] The augmentation generator 130 operates on the image patches 125 produced by the patch generator 120 from the image patch 110 to generate the augmented image patch 135. The augmentation generator 130 operates by applying variable image augmentations to each image patch, where at least two of the augmented image patches (for a plurality of image patches of the image patch) have different augmentations. Reference is made herein Figures 3A to 3B and Figure 4 Further details are provided regarding the generation of the augmented image patch via the augmentation generator 130.
[0021] The augmented image patch 135 is provided as an input to the neural network 140, which operates to perform image analysis tasks such as, for example, classification, detection, segmentation, etc. That is, the neural network 140 is used for inference / evaluation of images. The neural network 140 is trained to perform the type of image analysis task (e.g., classification, detection, segmentation, etc.). In some embodiments, the output of the neural network 140 provides the output image analysis result 160 of the system 100. In some embodiments, the output of the neural network undergoes further processing to produce the image analysis result 160.
[0022] In some embodiments, a given image patch 110 is input into the system 100 multiple times to produce multiple augmented image patches 135, where each augmented image patch 135 has a unique combination of augmented image patches relative to the other augmented image patches. That is, each augmented image patch has at least one augmented image patch that is different (in at least one aspect) from the augmented image patches in the other augmented image patches, and each augmented image patch can have two, several, many, or all of the image patches that are different from the augmented image patches in the other augmented image patches. Each corresponding augmented image patch 135 is run through the neural network 140 to produce an output of the neural network 140 (e.g., one output for each augmented image patch). In some embodiments, the original image patch 110 is also input into the neural network 140 to produce an additional output of the neural network 140. For example, the augmentation process as described above is applied multiple times as a processing step to the image patch 110 (image patch I0), and a large number of augmented image patches I1,..., I n are generated. These augmented image patches I1,..., I n are each processed by the neural network 140 to produce outputs L1,..., L n . In some embodiments, the image patch I0 (non-augmented) is also processed by the neural network 140 to produce an output L0. Then the corresponding outputs of the neural network 140 (e.g., L0,..., L n or L1,..., L n)to generate the image analysis result 160.
[0023] In an embodiment, the corresponding outputs are combined according to one or more combination (e.g., fusion) techniques. In an embodiment, such combination techniques can include, for example, averaging, weighted averaging, majority voting, weighted voting, etc. For example, if the image analysis task is classification (e.g., binary classification), each corresponding output of the neural network 140 can be a probability map with values between [0, 1]. The corresponding outputs (binary classification) of the neural network 140 can be combined in several ways. As an example (binary classification), each of the corresponding outputs is combined to produce an average value, and then a threshold can be applied to the average value to produce a binary decision. The binary decision becomes the image analysis result 160 based on the average output. As another example (binary classification), for each corresponding output, a threshold can be applied to produce a binary decision, and then majority voting is performed based on these binary decisions to determine the image analysis result 160. As another example (binary classification), for each corresponding output, a threshold can be applied to produce a binary decision for the corresponding output. Weights are assigned to each binary decision depending on, for example, the distance between the corresponding output and the threshold, and then weighted majority voting is performed (based on the assigned weights) to determine the image analysis result 160. A predetermined result can be defined to be applied in the case of equal votes.
[0024] Other ways of combining the corresponding outputs of the neural network 140 can be used and can depend on, for example, the type of image analysis task involved. For example, if the image analysis task is segmentation, each corresponding output can be a per-pixel (or per-voxel) value (e.g., 0 or 1, 0, 1, 1 or 2, etc.) indicating whether the corresponding pixel (or voxel) in the image represents background, a first object, a second object, etc. As an example, combining the segmentation output results can include performing majority voting or weighted majority voting on a per-pixel (or per-voxel) basis. As another example, if the image analysis task is detection, the corresponding outputs can be combined by performing averaging or weighted averaging.
[0025] In some embodiments, depending on the type of augmentation used and the type of image analysis task performed, it may be necessary to reverse the augmentation before combining the corresponding outputs. For example, if mirroring is used as an augmentation component and the neural network output provides a pixel- or voxel-based result, the portion corresponding to the image patch to which the mirroring was applied will need to be "demirrored" to reverse the mirroring. In some embodiments, if the augmentation is a transformation, the augmentation is reversed by applying the inverse transformation to the portion corresponding to the transformed image patch.
[0026] Some or all of the components in system 100 can be implemented using one or more of a central processing unit (CPU), a graphics processing unit (GPU), an artificial intelligence (AI) accelerator, a field programmable gate array (FPGA) accelerator, an application specific integrated circuit (ASIC), and / or via a processor with software or in combination with a processor with software and an FPGA or ASIC. More specifically, the components of system 100 can be implemented as one or more modules, hardware, or any combination thereof, as a set of program or logic instructions stored in a machine or computer readable storage medium such as random access memory (RAM), read only memory (ROM), programmable ROM (PROM), firmware, flash memory, etc. For example, a hardware implementation can include configurable logic, fixed function logic, or any combination thereof. Examples of configurable logic include appropriately configured programmable logic arrays (PLA), FPGAs, complex programmable logic devices (CPLD), and general purpose microprocessors. Examples of fixed function logic include appropriately configured ASICs, combinational logic circuits, and sequential logic circuits. Configurable or fixed function logic can be implemented using complementary metal oxide semiconductor (CMOS) logic circuits, transistor-transistor logic (TTL) logic circuits, or other circuits.
[0027] For example, the computer program code for operating system 100 can be written in any combination of one or more programming languages, including object oriented programming languages (such as JAVA, SMALLTALK, C++, etc.) and conventional procedural programming languages (such as the "C" programming language or similar programming languages). Additionally, the program or logic instructions may include assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine related instructions, microcode, status setting data, configuration data for integrated circuits, and other structural components that localize the electronic circuit and / or hardware (e.g., host processor, central processing unit / CPU, microcontroller, etc.), status information.
[0028] Referring to the components and features described herein (including but not limited to the drawings and associated descriptions), Figure 1B FIG. 170 is provided showing an example of image blocks used in one or more embodiments. A set of complete images 112 can be grouped together as a set of image blocks 110. In some embodiments, the image blocks 110 can be obtained from a collection of training images. In some embodiments, one or more sub-images 114 are obtained from the set of complete images 112. As an example, a sub-image 114 can be obtained by extracting a portion (e.g., a sample or slice) of the complete image 112. In some examples, multiple sub-images 114 can be extracted from the complete image 112, as Figure 1B shown. The set of sub-images 114 can be grouped together as a set of image blocks 110. Thus, as Figure 1BAs shown, the set of image patches 110 can be a complete image or a sub-image extracted from a complete image. In either case, one or more of the image patches 110 can be used as input to the system 100.
[0029] Referring to the components and features described herein (including but not limited to the drawings and associated descriptions), Figure 2 FIG. 200 is provided that illustrates an example process 200 for dividing an image patch into image tiles according to one or more embodiments. Process 200 operates on one or more image patches 110 (e.g., via tile generator 120). For each image patch, process 200 divides the image patch into a set of tiles, where each tile has a size, shape, and position relative to the image patch. For example, process 200 determines tile size / shape 210 and grid position 220 for each tile. Generally, the tile size / shape 210 and / or the grid position 220 will vary between tiles and / or between patches. The tile size can be determined based on one or more factors such as the number of tiles in the patch, a threshold amount (e.g., percentage) of the image patch represented jointly by the image tiles of the patch, etc., and the tile size can vary between the tiles in the image patch. The tile shape can be selected from one or more shapes such as a square, rectangle, etc., and the tile shape can vary between the tiles in the image patch. The grid position 220 of the image tile provides the position of the tile within the image patch and can be determined based on one or more factors such as the number of tiles in the patch, a threshold amount (e.g., percentage) of the image patch represented jointly by the image tiles of the patch, etc. In some embodiments, the image tiles of a given image patch can overlap.
[0030] In some embodiments, the tile size, tile shape, and / or tile position and / or one or more of the one or more factors used to determine the tile size, tile shape, and / or tile position can be selected based on random variation (e.g., via a random process or a random variable). Thus, as an example, for each image patch, the number of image tiles of the patch is determined based on a random variable (which, in an embodiment, can be bounded by the maximum number of tiles per patch and the minimum number of tiles per patch). As another example, for a given image patch, the shape of each image tile can be randomly selected.
[0031] An example of a patch divided into tiles is shown as patch - tile 230 in Figure 2 as Figure 2As shown, the first block (Block 1) has four patches labeled 1, 2, 3, and 4. Patches 1 - 4 vary in size and shape and have some overlap between adjacent patches. A small portion of Block 1 does not have a corresponding patch, indicating that the block can have fewer corresponding patches than the whole block. In the absence of a corresponding patch, the augmented block will use the data from the original image block without augmentation for such portions. The second block (Block 2) has four patches labeled 1, 2, 3, and 4, each having a similar size and shape, no overlap, and representing approximately 100% of Block 2. The third block (Block 3) has six patches labeled 1, 2, 3, 4, 5, and 6, which have different sizes and shapes. As Figure 2 illustrated by the patches generated by process 200 shown, a substantially unrestricted number of variations in patch size, shape, and position can be generated by patch generator 120.
[0032] Process 200 can generally be implemented in patch generator 120 (FIG. 1, already discussed). More specifically, process 200 and / or the functions performed by patch generator 120 can be implemented as one or more modules as a set of logical instructions stored in a machine or computer-readable storage medium (such as RAM, ROM, PROM, firmware, flash memory, etc.), hardware, or any combination thereof. For example, a hardware implementation can include configurable logic, fixed-function logic, or any combination thereof. Examples of configurable logic include appropriately configured PLAs, FPGAs, CPLDs, and general-purpose microprocessors. Examples of fixed-function logic include appropriately configured ASICs, combinational logic circuits, and sequential logic circuits. Configurable or fixed-function logic can be implemented using CMOS logic circuits, TTL logic circuits, or other circuits.
[0033] For example, the computer program code for performing process 200 and / or the functions associated with patch generator 120 can be written in any combination of one or more programming languages, including object-oriented programming languages (such as JAVA, SMALLTALK, C++, etc.) and conventional procedural programming languages (such as the "C" programming language or similar programming languages). Additionally, the program or logical instructions can include assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, status setting data, configuration data for integrated circuits, and other structural components that personalize electronic circuits and / or hardware (e.g., host processors, central processing units / CPUs, microcontrollers, etc.).
[0034] Referring to the components and features described herein (including but not limited to the drawings and associated descriptions), Figure 3AFIG. shows an example process 300 for generating augmented image patches according to one or more embodiments. Process 300 operates on image patches generated by a patch generator 120 (e.g., via an augmentation generator 130). For each image patch, process 300 obtains the patch of the block and augments each patch using variable image augmentation. For example, the variable image augmentation employed by process 300 can include one or more of the following augmentation selections (e.g., augmentation types): filtering; intensity adjustment; added noise (e.g., via a noise generator); style transfer, etc., and there can be one or more parameters associated with any selection. For a given image patch, the variable image augmentation can include one or more of these selections, and when more than one type of augmentation is selected for an image patch, these augmentations can be blended for the image patch. Generally, the augmentations will vary from patch to patch, although in some embodiments, some patches will have the same augmentation. In any case, at least one augmented image patch will have at least two patches with different augmentations. In some embodiments, one or more augmentation selections and / or parameters associated therewith can be selected based on random variations (e.g., via a random process or a randomized variable). For example, the number of augmentations applied to a given image patch can be selected based on a randomized variable. As another example, the selected augmentations for a given image patch can be randomly selected.
[0035] An example of an augmented block based on applying augmentations to an image block divided into image patches is shown in Figure 3A as augmented image block 320. As Figure 3A shown, the first augmented block 322 (block 1) has four augmented patches labeled 1 - 4 (which correspond to the four patches of block 1 shown in Figure 2 ). The four patches of block 1 have different augmentations (as indicated by the different hashes for each image patch shown in the figure). The second augmented block 324 (block 2) has four augmented patches labeled 1 - 4 (which correspond to the four patches of block 2 shown in Figure 2 ). The four patches of block 2 have different augmentations (as indicated by the different hashes or shading for each image patch shown in the figure). The third augmented block 326 (block 3) has six augmented patches labeled 1 - 6 (which correspond to the six patches of block 3 shown in Figure 2 ). The six patches of block 3 have different augmentations (as indicated by the different hashes or shading for each image patch shown in the figure). Further details regarding variable image augmentation are provided herein with reference to Figure 4 . In some embodiments, a geometric transformation can be applied to the entire image block before or after applying per-patch augmentations to the patches of the image block.
[0036] Process 300 can generally be implemented in the augmentation generator 130 (Figure 1, already discussed). More specifically, process 300 and / or the functions performed by the augmentation generator 130 can be implemented as one or more modules, hardware, or any combination thereof as a set of logical instructions stored in a machine or computer-readable storage medium such as RAM, ROM, PROM, firmware, flash memory, etc. For example, a hardware implementation can include configurable logic, fixed-function logic, or any combination thereof. Examples of configurable logic include appropriately configured PLAs, FPGAs, CPLDs, and general-purpose microprocessors. Examples of fixed-function logic include appropriately configured ASICs, combinational logic circuits, and sequential logic circuits. Configurable or fixed-function logic can be implemented using CMOS logic circuits, TTL logic circuits, or other circuits.
[0037] For example, computer program code for performing process 300 and / or functions associated with the augmentation generator 130 can be written in any combination of one or more programming languages, including object-oriented programming languages such as JAVA, SMALLTALK, C++, etc., and conventional procedural programming languages such as the "C" programming language or similar programming languages. Additionally, program or logic instructions can include assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, status setting data, configuration data for integrated circuits, and other structural components that localize the electronic circuit and / or hardware (e.g., host processor, central processing unit / CPU, microcontroller, etc.) personalized status information.
[0038] Figure 3B A figure is provided showing an example of an augmented block 380 generated from an image block 360. Six augmented blocks 380 are generated using a variable augmentation per-patch process (similar to processes 200 and 300), where the patch position is randomly selected on a grid, the patch size is randomly selected from a plurality of patch sizes, and a random amount of noise and / or a gamma transformation with random parameters (contrast and brightness adjustment) is added as an augmentation to the patch.
[0039] Referring to the components and features described herein (including but not limited to the drawings and associated descriptions), Figure 4 A figure is provided showing an example of an augmentation generator 400 for generating augmented image blocks according to one or more embodiments. In an embodiment, the augmentation generator 400 corresponding to the augmentation generator 130 ( Figure 1A and 3A ) augments the patches generated by the patch generator 120 ( Figure 1A and 2)The one or more generated image patches 410 are operated on to apply variable image augmentation to the image patches 410 (e.g., provide variable image augmentation to the image patches 410). The augmentation generator 400 performs process 300 and includes one or more augmentation modules 420 for selecting various types of image augmentations, such as, for example, a filter module 421, an intensity adjustment module 422, a noise generator module 423, a style transfer module 424, etc. In some embodiments, other augmentation modules ( Figure 4 not shown) can provide different / additional types of image augmentations.
[0040] When selected, the filter module 421 applies a filtering operation to the image patch 410. The filtering operation can include any type of image filtering operation, including low-pass filtering, band-pass filtering, high-pass filtering, edge filtering, Gaussian filtering, gradient filtering, etc. The type and / or characteristics of the applied filtering operation can be based on associated parameters (e.g., selectable). In an embodiment, when the filter module 421 is the only selected augmentation for a given image patch 410, the output of the module becomes the augmented image patch 460.
[0041] When selected, the intensity adjustment module 422 applies a transformation to the image patch 410 to produce an adjustment to the image intensity. Examples of intensity adjustment transformations include one or more of histogram transformation, gamma transformation, linear contrast / brightness transformation, etc. The type and / or characteristics of the applied intensity transformation can be based on associated parameters (e.g., selectable). In an embodiment, when the intensity adjustment module 422 is the only selected augmentation for a given image patch 410, the output of the module becomes the augmented image patch 460.
[0042] When selected, the noise generator module 423 applies a noise component to the image patch 410, which can be performed additively or otherwise. The noise component can include any type of noise, such as Gaussian noise, white noise, etc. The type and / or characteristics of the applied noise component can be based on associated parameters (e.g., selectable). In an embodiment, when the noise generator module 423 is the only selected augmentation for a given image patch 410, the output of the module becomes the augmented image patch 460.
[0043] When selected, the style transfer module 424 applies a transformation to the image patch 410 to produce a style transfer, for example, via a generative adversarial network (GAN) or a similar neural network. The type and / or characteristics of the style transfer can be based on associated parameters (e.g., selectable). In an embodiment, when the style transfer module 424 is the only selected augmentation for a given image patch 410, the output of the module becomes the augmented image patch 460.
[0044] In an embodiment, more than one type of augmentation is selected and applied to the image patch 410. In this case, the output of each augmentation module 420 is provided as input to a mixing module 430 to combine the augmentations using a mixing operation, and the output of the mixing module 430 becomes the augmented image patch 460. The mixing operation can be any type of suitable mixing operation (e.g., on a per-pixel basis), such as, for example, providing the average of the inputs, the weighted average of the inputs, etc.
[0045] As previously mentioned, the augmentations typically vary from one image patch to the next to provide variable image augmentation. That is, typically, the augmentations will vary from patch to patch, although in some embodiments, some patches will have the same augmentation. In any case, for variable image augmentation, at least one augmented image block will have at least two patches with different augmentations. Different augmentations can be applied by various methods. As an example, different augmentations can be obtained by changing the type of augmentation applied (e.g., by selecting different augmentation modules from one image patch to the next). As another example, different augmentations can be obtained by changing the parameters for a particular type of augmentation (e.g., applying different filters from one image patch to the next). As another example, different augmentations can be obtained by changing one or more of the augmentation type and / or the parameters for a particular type of augmentation from one image patch to the next.
[0046] In some embodiments, the selection of the augmentations (e.g., augmentation type and / or associated parameters for the augmentation type) is performed randomly. A random generator 440 (e.g., a random number or random variable generator) can be used to randomly select one or more of the augmentation modules 420 and / or one or more parameters associated with the augmentation modules 420. As an example, the random generator 440 can be used to provide randomized module / parameter selection for each image patch 410. As another example, the random generator 440 can be used to provide a list (e.g., a string) of randomized module / parameter selections, each corresponding selection for a different image patch 410.
[0047] In some embodiments, the selection of augmentations (e.g., augmentation types and / or associated parameters for the augmentation types) is performed based on the augmentation queue 450. For example, a predetermined queue (e.g., a list or table) of varying augmentation types and / or associated parameters can be generated or designed. Each entry in the queue can include a single and / or multiple augmentation types to be applied to a given image patch. For a given image patch, the augmentations (types and / or parameters) provided in the queue entry are applied to the image patch, where the output is provided as the augmented image patch 460, or in the case of multiple augmentation types, the output is provided to the mixing module 430 for mixing and then output from the mixing module 430 as the augmented image patch 460. For the next image patch, the next queue entry provides the augmentation type for that patch, and so on until each entry in the queue has been used.
[0048] In some embodiments, additional augmentations can be performed on the image block (e.g., applied to the entire image block before partitioning into image patches and augmenting the patches) and / or an augmented image block can be generated (e.g., applied to the entire augmented image block after patch-based augmentation). For example, spatial transformations such as, for example, non-rigid transformations (e.g., image warping or shape change), mirroring, rotation, translation, scaling, affine / elastic transformations, etc. can be applied to the image block and / or the augmented image block. The type and / or characteristics of the spatial transformation can be based on associated parameters (e.g., selectable). Other augmentations can be selectively applied to the image block and / or the augmented image block in a similar manner.
[0049] Some or all of the components in the augmentation generator 400 can be implemented using one or more of a CPU, GPU, AI accelerator, FPGA accelerator, ASIC, and / or via a processor with software or in combination with a processor with software and an FPGA or ASIC. More specifically, the augmentation generator 400 and / or the functions performed by the augmentation generator 400 can be implemented as one or more modules as a set of logical instructions stored in a machine or computer-readable storage medium (such as RAM, ROM, PROM, firmware, flash memory, etc.), hardware, or any combination thereof. For example, a hardware implementation can include configurable logic, fixed-function logic, or any combination thereof. Examples of configurable logic include appropriately configured PLAs, FPGAs, CPLDs, and general-purpose microprocessors. Examples of fixed-function logic include appropriately configured ASICs, combinational logic circuits, and sequential logic circuits. Configurable or fixed-function logic can be implemented using CMOS logic circuits, TTL logic circuits, or other circuits.
[0050] For example, the computer program code for performing the functions executed by the augmentation generator 400 can be written in any combination of one or more programming languages, including object-oriented programming languages (such as JAVA, SMALLTALK, C++, etc.) and conventional procedural programming languages (such as the "C" programming language or similar programming languages). Additionally, the program or logical instructions can include assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, status setting data, configuration data for integrated circuits, status information that personalizes other structural components local to the electronic circuit and / or hardware (e.g., host processor, central processing unit / CPU, microcontroller, etc.).
[0051] Referring to the components and features described herein (including but not limited to the drawings and associated descriptions), Figures 5A - 5B FIG. 500 (including process blocks 500A and 500B) provides a flow diagram illustrating an example method 500 for tile-based image augmentation for image analysis in accordance with one or more embodiments. The method 500 can generally be implemented in and / or via the components of the system 100 (FIG. 1, already discussed). More specifically, the method 500 can be implemented as one or more modules as a set of logical instructions stored in a machine or computer-readable storage medium (such as RAM, ROM, PROM, firmware, flash memory, etc.), hardware, or any combination thereof. For example, a hardware implementation can include configurable logic, fixed-function logic, or any combination thereof. Examples of configurable logic include appropriately configured PLAs, FPGAs, CPLDs, and general-purpose microprocessors. Examples of fixed-function logic include appropriately configured ASICs, combinational logic circuits, and sequential logic circuits. The configurable or fixed-function logic can be implemented with CMOS logic circuits, TTL logic circuits, or other circuits.
[0052] For example, the computer program code for performing the operations shown in method 500 and / or the functions associated therewith can be written in any combination of one or more programming languages, including object-oriented programming languages (such as JAVA, SMALLTALK, C++, etc.) and conventional procedural programming languages (such as the "C" programming language or similar programming languages). Additionally, the program or logical instructions can include assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, status setting data, configuration data for integrated circuits, status information that personalizes other structural components local to the electronic circuit and / or hardware (e.g., host processor, central processing unit / CPU, microcontroller, etc.).
[0053] Go to Figure 5A, Method 500A begins at the illustrated processing block 510a by dividing an image patch into multiple image slices, where at the illustrated processing block 510b, each image slice has a size, shape, and position relative to the image patch. The illustrated processing block 520a provides for generating an augmented image patch by applying a variable image augmentation to each image slice, where at the illustrated processing block 520b, at least two augmented image slices have different augmentations. The illustrated processing block 530 provides for performing an image analysis task by applying a neural network to the augmented image patch. In an embodiment, the image analysis task is one of classification, detection, or segmentation.
[0054] In some embodiments, for each of the multiple image slices, one or more of the size, shape, or position of the corresponding image slice are randomly selected. In some embodiments, the image slices together represent at least a threshold amount of the image patch. In some embodiments, at least two image slices overlap within the image patch, and the overlapping regions of at least two image slices are blended together after augmentation. In some embodiments, a neural network is trained using multiple augmented training image patches generated from multiple training image patches, where for at least one augmented training image patch, at least two augmented image slices have different augmentations.
[0055] In an embodiment, the variable image augmentation includes selecting one or more augmentation components, where an augmentation component can have one or more parameters to further define (or refine) the augmentation component or its aspects or features to be applied to the corresponding image slice. In some embodiments, the variable image augmentation includes at least one augmentation component selected from a plurality of augmentation components. In some embodiments, the plurality of augmentation components includes one or more of a filter, intensity adjustment, noise generator, or style transfer. In some embodiments, for each of the multiple image slices, at least one augmentation component is randomly selected. In some embodiments, for each of the multiple image slices, a parameter associated with at least one augmentation component is randomly selected. In some embodiments, for each of the multiple image slices, at least one augmentation component is selected based on a predetermined queue. In some embodiments, for each image slice having multiple selected augmentation components, the outputs of each of the selected augmentation components are blended to provide the augmented image slice.
[0056] Now turning to Figure 5B , Method 500B provides for generating multiple augmented image patches by repeatedly dividing and generating operations on an image patch at the illustrated processing block 550a, where at the illustrated processing block 550b, each augmented image patch has a unique combination of augmented image slices relative to the other augmented image patches. Performing an image analysis task ( Figure 5A, the processing block 530) includes, at the illustrated processing block 560, applying a neural network to each of the plurality of augmented image patches to obtain a corresponding output of the neural network for each augmented image patch, and at the illustrated processing block 570, determining a result of the image processing task based on a combination of the corresponding outputs of the neural network. In some embodiments, the original image patches are also input into the neural network to generate additional outputs of the neural network, which are then combined with the other outputs in a similar manner. In some embodiments, combining the corresponding outputs of the neural network includes using one or more combination (e.g., fusion) techniques. In an embodiment, such combination techniques can include, for example, averaging, weighted averaging, majority voting, weighted voting, etc.
[0057] Referring to the components and features described herein (including but not limited to the drawings and associated descriptions), Figure 6 FIG. 6 is a block diagram showing a tile-based image augmentation training system 600 illustrating examples in accordance with one or more embodiments. Certain components, aspects, and operations of system 600 correspond to or are similar to those of system 100 (FIG. 1, already discussed), and thus, some details will not be repeated except as necessary or appropriate to understand system 600. As Figure 6 shown, system 600 receives and processes a plurality of image patches 110. As described herein (and as Figure 1B shown), an image patch such as image patch 110 refers to a complete image and / or a sub-image of a complete image. Images can include, for example, images generated by diagnostic or other medical imaging systems, such as, for example, MR images, CT images, X-rays, and / or images obtained via other imaging techniques (such as, for example, ultrasound), or images generated by imaging sensors, such as, for example, visible or infrared images, etc. Images (and image patches) can have different dimensions, including, for example, two-dimensional (2D) or three-dimensional (3D) images. Image patches 110 can be selected to be used as potential training images (e.g., once augmented as described herein) to train neural network 640.
[0058] As Figure 6 further shown, system 600 includes a tile generator 620, an augmentation generator 630, and a neural network 640, where neural network 640 is untrained. System 600 uses a neural network training algorithm 650 to train neural network 640 and produce a trained neural network 660. It should be understood that in some embodiments, system 100 can include more, alternative, or fewer components than those Figure 1A shown, and in some embodiments, some components can be combined with or incorporated into other components.
[0059] Patch generator 620 operates on image patches 110 and generates per-patch image slices 625 for each input image patch 110. In an embodiment, patch generator 620 corresponds to patch generator 120 ( Figure 1A and 2 , which has been discussed). In an embodiment, per-patch image slice 625 corresponds to per-patch image slice 125 ( Figure 1A , which has been discussed) and / or block-slice 230 ( Figure 2 , which has been discussed).
[0060] Augmentation generator 630 applies variable augmentations to per-patch image slices 625 for each image patch 110 input to system 600 and generates multiple augmented image patches 635, thereby generating, for example, augmented image patches 635 for each input image patch 110. In an embodiment, augmentation generator 630 corresponds to augmentation generator 130 ( Figure 1A and 3A , which has been discussed) and / or augmentation generator 400 ( Figure 4 , which has been discussed). In an embodiment, each augmented image patch 635 corresponds to augmented image patch 135 ( Figure 1A , which has been discussed). In an embodiment, augmented image patch 635 corresponds to augmented image patch 320 ( Figure 3A , which has been discussed).
[0061] Augmented image patches 635 are used as training images and are input to untrained neural network 640. Neural network training algorithm 650 is used to train neural network 640 based on augmented image patches 635. Neural network training algorithm 650 can include any neural network training algorithm suitable for training a neural network to perform a given image analysis task such as classification, detection, segmentation, etc. Once the training process is executed, the result is a trained neural network 660. The trained neural network 660 can be used to perform an appropriate image analysis task (e.g., classification, detection, segmentation, etc.) for which the neural network has been trained on other image patches ( Figure 6 not shown). For example, the trained neural network 660 can replace neural network 140 ( Figure 1A , which has been discussed).
[0062] In some embodiments, in cases where neural network training involves patches of a specific size, patch generator 120 can make the patch size selectable to avoid using the same patch size as the specific patch size used for training for augmentation. In some embodiments, a random patch size is used by patch generator 120 as an alternative to avoid using a specific patch size for training.
[0063] Some or all of the components in system 600 can be implemented using one or more of a CPU, GPU, AI accelerator, FPGA accelerator, ASIC, and / or via a processor with software or in combination with a processor with software and an FPGA or ASIC. More specifically, the components of system 600 can be implemented as one or more modules, hardware, or any combination thereof of a set of programs or logical instructions stored in a machine or computer-readable storage medium (such as RAM, ROM, PROM, firmware, flash memory, etc.). For example, a hardware implementation can include configurable logic, fixed-function logic, or any combination thereof. Examples of configurable logic include a properly configured PLA, FPGA, CPLD, and general-purpose microprocessor. Examples of fixed-function logic include a properly configured ASIC, combinational logic circuits, and sequential logic circuits. Configurable or fixed-function logic can be implemented using CMOS logic circuits, TTL logic circuits, or other circuits.
[0064] For example, the computer program code for operating system 600 can be written in any combination of one or more programming languages, including object-oriented programming languages (such as JAVA, SMALLTALK, C++, etc.) and conventional procedural programming languages (such as the "C" programming language or similar programming languages). Additionally, the program or logical instructions can include assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, status setting data, configuration data for an integrated circuit, and other structural components that localize the electronic circuit and / or hardware (e.g., host processor, central processing unit / CPU, microcontroller, etc.), such as status information.
[0065] Referring to the components and features described herein (including but not limited to the drawings and associated descriptions), Figure 7 A flowchart is provided showing an example method 700 of tile-based image augmentation training used in training a neural network according to one or more embodiments. Method 700 can generally be implemented in system 600 ( Figure 6 , already discussed) and / or via its components. More specifically, method 700 can be implemented as one or more modules, hardware, or any combination thereof of a set of logical instructions stored in a machine or computer-readable storage medium (such as RAM, ROM, PROM, firmware, flash memory, etc.). For example, a hardware implementation can include configurable logic, fixed-function logic, or any combination thereof. Examples of configurable logic include a properly configured PLA, FPGA, CPLD, and general-purpose microprocessor. Examples of fixed-function logic include a properly configured ASIC, combinational logic circuits, and sequential logic circuits. Configurable or fixed-function logic can be implemented using CMOS logic circuits, TTL logic circuits, or other circuits.
[0066] For example, the computer program code for performing the operations shown in method 700 and / or the functions associated therewith can be written in any combination of one or more programming languages, including object-oriented programming languages (such as JAVA, SMALLTALK, C++, etc.) and conventional procedural programming languages (such as the "C" programming language or similar programming languages). Additionally, the program or logical instructions can include assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, status setting data, configuration data of an integrated circuit, and status information that personalizes other structural components local to the electronic circuit and / or hardware (e.g., host processor, central processing unit / CPU, microcontroller, etc.).
[0067] The illustrated processing block 710a provides for dividing each of a plurality of image patches into a plurality of image slices, where at the illustrated processing block 710b, each image slice has a size, shape, and position relative to the corresponding image patch. The illustrated processing block 720a provides for generating a plurality of augmented image patches by applying a variable image augmentation to each image slice of an image patch for each image patch, where at the illustrated processing block 720b, for at least one augmented image patch, at least two augmented image slices have different augmentations. The illustrated processing block 730 provides for training a neural network using the plurality of augmented image patches as training images.
[0068] In some embodiments, for each of the plurality of image slices, one or more of the size, shape, or position of the corresponding image slice is randomly selected. In some embodiments, for each image patch, the image slices of the image patch together represent at least a threshold amount of the image patch. In some embodiments, at least two image patches overlap within the image patch, and the overlapping regions of at least two image slices are blended together after augmentation.
[0069] In an embodiment, the variable image augmentation includes selecting one or more augmentation components, where an augmentation component can have one or more parameters to further define (or refine) the augmentation component or its aspects or features to be applied to the corresponding image slice. In some embodiments, the variable image augmentation includes at least one augmentation component selected from a plurality of augmentation components. In some embodiments, the plurality of augmentation components includes one or more of a filter, intensity adjustment, noise generator, or style transformation. In some embodiments, for each of the plurality of image slices, at least one augmentation component is randomly selected. In some embodiments, for each of the plurality of image slices, a parameter associated with at least one augmentation component is randomly selected. In some embodiments, for each of the plurality of image slices, at least one augmentation component is selected based on a predetermined queue. In some embodiments, for each image slice having a plurality of selected augmentation components, the outputs of each of the selected augmentation components are blended to provide the augmented image slice.
[0070] It should be understood that method 700 can be executed in several ways. As an example, the operations of method 700 can be executed serially, where the first image block is divided into multiple image patches, a first augmented image block is generated by applying variable image augmentation to each image patch of the first image block, and the first augmented image block is used to train a neural network. Then this process can be repeated for the second image block, the third image block, etc. in serial manner. For example, these tasks can be executed in real time in serial manner (e.g., the augmentation applied "on-the-fly" to the image patches in the image block during training). As another example, the operations of method 700 can be executed in batch manner, where each of multiple image blocks is divided into multiple image patches, and a set of augmented image blocks is generated by applying variable image augmentation to each image patch for the multiple image blocks, and then this set of augmented image blocks is stored and used to train a neural network. As another instance, the operations of method 700 can be executed in a hybrid serial-batch manner, where some aspects are executed serially and other operations are executed in batch manner. For example, augmented image blocks can be generated in serial manner (each augmented image block based on a corresponding input image block), and then the augmented image blocks are stored; once all the augmented image blocks are generated, then this set of augmented image blocks is used to train a neural network.
[0071] Referring to the components and features described herein (including but not limited to the drawings and associated descriptions), Figure 8 is a diagram showing a computing system 800 used in system 100 and / or system 600 according to one or more embodiments. Although Figure 8 certain components are shown, computing system 800 can include additional or multiple components connected in various ways. It should be understood that not all examples must include Figure 8 every component shown in Figure 8 As shown, computing system 800 includes one or more processors 802, an I / O subsystem 804, a network interface 806, a memory 808, a data storage 810, an artificial intelligence (AI) accelerator 812, a user interface 816, and / or a display 820. These components are coupled, connected, or otherwise communicate data via an interconnect 814. In some embodiments, computing system 800 interacts with a separate display. Computing system 800 can implement one or more components or features of system 100, system 600, and / or any of the components, features, or methods described herein with reference to Figure 1A 、 1B 、2, 3A, 3B, 4, 5A, 5B, 6, and / or 7.
[0072] Processor 802 includes one or more processing devices, such as a microprocessor, a central processing unit (CPU), a fixed application specific integrated circuit (ASIC) processor, a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a field programmable gate array (FPGA), a digital signal processor (DSP), etc., as well as associated circuits, logic, and / or interfaces. Processor 802 can optionally include or be connected to a memory (e.g., memory 808) that stores executable instructions and / or data as needed or appropriately. Processor 802 can execute such instructions to implement, control, operate, or interface with system 100, system 600, and / or any component, feature, or method described herein with reference to Figure 1A , 1B , 2, 3A, 3B, 4, 5A, 5B, 6, and / or 7, or any component or feature of any component, feature, or method described herein. Processor 802 can transmit, send, or receive messages, requests, notifications, data, etc. to / from other devices. Processor 802 can be implemented as any type of processor capable of performing the functions described herein. For example, processor 802 can be implemented as a single-core or multi-core processor, a digital signal processor, a microcontroller, or other processor or processing / control circuit. The processor can include embedded instructions (e.g., processor code).
[0073] The I / O subsystem 804 includes circuits and / or components adapted to facilitate input / output operations with processor 802, memory 808, and other components of computing system 800.
[0074] The network interface 806 includes suitable logic, circuits, and / or interfaces for sending and receiving data over one or more communication networks using one or more communication network protocols. The network interface 806 can operate under the control of processor 802 and can send / receive various requests and messages to / from one or more other devices. The network interface 806 can include wired or wireless data communication capabilities; these capabilities can support data communication with a wired or wireless communication network (such as network 807), and also include the Internet, a wide area network (WAN), a local area network (LAN), a wireless personal area network, a wide area network, a cellular network, a telephone network, any other wired or wireless network for sending and receiving data signals, or any combination thereof (including, for example, a Wi-Fi network or a corporate LAN). The network interface 806 can support communication via a short-range wireless communication field (such as Bluetooth, NFC, or RFID). Examples of the network interface 806 include, but are not limited to, one or more of an antenna, a radio frequency transceiver, a wireless transceiver, a Bluetooth transceiver, an Ethernet port, a universal serial bus (USB) port, or any other device configured to send and receive data.
[0075] Memory 808 includes suitable logic, circuitry, and / or interfaces to store executable instructions and / or data, when necessary or appropriate, to implement, control, operate, or interface with the system 100, system 600, and / or any components, features, or methods described herein with reference to Figure 1A , 1B , 2, 3A, 3B, 4, 5A, 5B, 6, and / or 7. Memory 808 can be implemented as any type of volatile or non-volatile memory or data storage device capable of performing the functions described herein, and can include random access memory (RAM), read-only memory (ROM), write-once read-many memory (e.g., EEPROM), removable storage drives, hard disk drives (HDDs), flash memory, solid-state memory, etc., and any combination thereof. In operation, memory 808 can store various data and software used during the operation of computing system 800, such as operating systems, applications, programs, libraries, and drivers. Thus, memory 808 can include at least one non-transitory computer-readable medium that includes instructions that, when executed by computing system 800, cause computing system 800 to perform operations to execute one or more functions or features of system 100, system 600, and / or any components, features, or methods described herein with reference to Figure 1A , 1B , 2, 3A, 3B, 4, 5A, 5B, 6, and / or 7. Memory 808 can be communicatively coupled to processor 802 directly or via I / O subsystem 804.
[0076] Data storage 810 can include any type of one or more devices configured for short-term or long-term storage of data, such as memory devices and circuitry, memory cards, hard disk drives, solid-state drives, non-volatile flash memory, or other data storage devices. Data storage 810 can include or be configured as a database, such as a relational database or a non-relational database, or a combination of more than one database. In some examples, the database or other data storage can be physically separated from and / or remote from computing system 800, and / or can be located in another computing device, a database server, a cloud-based platform, or any storage device that communicates data with computing system 800.
[0077] Artificial intelligence (AI) accelerator 812 includes suitable logic, circuitry, and / or interfaces to accelerate artificial intelligence applications, such as, for example, artificial neural networks, machine vision, and machine learning applications, including through parallel processing techniques. In one or more examples, AI accelerator 812 can include a graphics processing unit (GPU). AI accelerator 812 can implement system 100, system 600, and / or any components, features, or methods described herein with reference to Figure 1A , 1Bone or more components or features of the components, features, or methods described by 2, 3A, 3B, 4, 5A, 5B, 6, and / or 7, including neural network 140( Figure 1A ) and / or neural network 640 and / or neural network 660( Figure 6 ), one or more of which. In some examples, computing system 800 includes a second AI accelerator (not shown).
[0078] Interconnect 814 includes any one or more individual physical buses, point-to-point connections, or both connected by appropriate bridges, adapters, or controllers. Interconnect 814 can include (for example) a system bus, a Peripheral Component Interconnect (PCI) bus, a HyperTransport or Industry Standard Architecture bus, a Small Computer System Interface (SCSI) bus, a Universal Serial Bus (USB), an IIC (I2C) bus, or an Institute of Electrical and Electronics Engineers (IEEE) Standard 694 bus (e.g., "FireWire"), or any other interconnect suitable for coupling or connecting the components of computing system 800.
[0079] User interface 816 includes code for presenting information or screens to a user on a display and receiving input (including commands) from the user via an input device. Display 820 can be any type of device for presenting visual information, such as a computer monitor, a flat panel display, or a mobile device screen, and can include a liquid crystal display (LCD), a light-emitting diode (LED) display, a plasma panel, or a cathode ray tube display, etc. Display 820 can include a display interface for communicating with the display. In some examples, display 820 can include a display interface for communicating with a display external to computing system 800.
[0080] In some examples, one or more of the illustrative components of computing system 800 can be incorporated (in whole or in part) within another component or otherwise form a part of another component. For example, memory 808 or portions thereof can be incorporated within processor 802. As another example, user interface 816 can be incorporated within code in processor 802 and / or memory 808. In some examples, computing system 800 can be implemented as, but not limited to, a mobile computing device, a smart phone, a wearable computing device, an Internet of Things device, a laptop computer, a tablet computer, a notebook computer, a computer, a workstation, a server, a multiprocessor system, and / or a consumer electronic device. In some examples, computing system 800 or a portion thereof is implemented as a set of logic instructions in one or more modules stored in at least one non-transitory machine or computer-readable storage medium, such as random access memory (RAM), read-only memory (ROM), programmable ROM (PROM), firmware, flash memory, etc., in configurable logic, such as, for example, a programmable logic array (PLA), a field programmable gate array (FPGA), a complex programmable logic device (CPLD), in fixed-function logic hardware using circuit technology, such as, for example, an application specific integrated circuit (ASIC), complementary metal oxide semiconductor (CMOS), or transistor-transistor logic (TTL) technology, or in any combination thereof.
[0081] Embodiments of each of the above-described systems, devices, components, features, and / or methods (including system 100, process 200, process 300, augmentation generator 400, method 500, system 600, method 700, and / or any other system component) can be implemented in hardware, software, or any suitable combination thereof. For example, a hardware implementation can include configurable logic, fixed-function logic, or any combination thereof. Examples of configurable logic include appropriately configured PLAs, FPGAs, CPLDs, and general-purpose microprocessors. Examples of fixed-function logic include appropriately configured ASICs, combinational logic circuits, and sequential logic circuits. Configurable or fixed-function logic can be implemented using CMOS logic circuits, TTL logic circuits, or other circuits.
[0082] Alternatively or additionally, all or part of the foregoing systems, devices, components, features, and / or methods can be implemented in one or more modules as a set of programs or logical instructions stored in a machine or computer-readable storage medium (such as RAM, ROM, PROM, firmware, flash memory, etc.) to be executed by a processor or computing device. For example, computer program code for performing the operations of a component can be written in any combination of one or more operating system (OS)-applicable / appropriate programming languages, including object-oriented programming languages (such as PYTHON, PERL, JAVA, SMALLTALK, C++, C#, etc.) and conventional procedural programming languages (such as the "C" programming language or similar programming languages).
[0083] Additional notes and examples:
[0084] Example M A 1 includes a computer-implemented method that includes: dividing an image patch into a plurality of image tiles, each image tile having a size, shape, and position relative to the image patch; generating an augmented image patch by applying a variable image augmentation to each image tile, wherein at least two augmented image tiles have different augmentations; and performing an image analysis task by applying a neural network to the augmented image patch.
[0085] Example M A 2 includes Example M A The method of 1, further comprising generating a plurality of augmented image patches by repeatedly dividing and generating operations on the image patch, wherein each augmented image patch has a unique combination of augmented image tiles relative to other augmented image patches, wherein performing the image analysis task includes applying a neural network to each augmented image patch among the plurality of augmented image patches to obtain a corresponding output of the neural network for each augmented image patch, and determining a result of the image processing task based on combining the corresponding outputs of the neural network.
[0086] Example M A 3 includes Example M A 1 or M A The method of 2, wherein for each image tile among the plurality of image tiles, one or more of the size, shape, or position of the corresponding image tile are randomly selected.
[0087] Example M A 4 includes Example M A 1, M A 2 or M A The method of 3, wherein the image tiles together represent at least a threshold amount of the image patch.
[0088] Example M A 5 includes Example M A 1 - M AThe method according to any one of 4, wherein the variable image augmentation includes at least one augmentation component selected from a plurality of augmentation components.
[0089] Example M A 6 includes Example M A 1 - M A The method according to any one of 5, wherein the plurality of augmentation components includes one or more of a filter, intensity adjustment, noise generator, or style transfer.
[0090] Example M A 7 includes Example M A 1 - M A The method according to any one of 6, wherein for each image patch among a plurality of image patches, at least one augmentation component is randomly selected.
[0091] Example M A 8 includes Example M A 1 - M A The method according to any one of 7, wherein for each image patch among a plurality of image patches, parameters associated with at least one augmentation component are randomly selected.
[0092] Example M A 9 includes Example M A 1 - M A The method according to any one of 8, wherein for each image patch among a plurality of image patches, at least one augmentation component is selected based on a predetermined queue.
[0093] Example M A 10 includes Example M A 1 - M A The method according to any one of 9, wherein for each image patch having a plurality of selected augmentation components, the outputs for each selected augmentation component are mixed to provide an augmented image patch.
[0094] Example M A The method of 11 includes Example M A 1 - M A The method according to any one of 10, wherein at least two image patches overlap within an image block, and wherein the overlapping regions of at least two image patches are mixed together after augmentation.
[0095] Example M A 12 includes Example M A 1 - M A The method according to any one of 11, wherein a neural network is trained using a plurality of augmented training image blocks generated from a plurality of training image blocks, wherein for at least one augmented training image block, at least two augmented image patches have different augmentations.
[0096] Example S A1 includes a computing system that includes a processor and a memory coupled to the processor, the memory including instructions that, when executed by the processor, cause the computing system to perform operations that include: dividing an image block into a plurality of image patches, each image patch having a size, a shape, and a position relative to the image block; generating an augmented image block by applying a variable image augmentation to each image patch, wherein at least two of the augmented image patches have different augmentations; and performing an image analysis task by applying a neural network to the augmented image block.
[0097] Example S A 2 includes Example S A The computing system of 1, wherein the instructions, when executed, cause the computing system to perform further operations including generating a plurality of augmented image blocks by repeatedly dividing and generating operations on the image block, wherein each augmented image block has a unique combination of augmented image patches relative to other augmented image blocks, wherein performing the image analysis task includes applying the neural network to each of the plurality of augmented image blocks to obtain a corresponding output of the neural network for each augmented image block, and determining a result of the image processing task based on combining the corresponding outputs of the neural network.
[0098] Example S A 3 includes Example S A 1 or S A The computing system of 2, wherein for each of the plurality of image patches, one or more of the size, shape, or position of the corresponding image patch are randomly selected.
[0099] Example S A 4 includes Example S A 1, S A 2 or S A The computing system of 3, wherein the image patches together represent at least a threshold amount of the image block.
[0100] Example S A 5 includes Example S A 1 - S A The computing system of any one of 4, wherein the variable image augmentation includes at least one augmentation component selected from a plurality of augmentation components.
[0101] Example S A 6 includes Example S A 1 - S A The computing system of any one of 5, wherein the plurality of augmentation components includes one or more of a filter, an intensity adjustment, a noise generator, or a style transfer.
[0102] Example S A 7 includes Example S A 1 - SA A computing system according to any one of 6, wherein, for each of the plurality of image patches, the at least one augmentation component is randomly selected.
[0103] Example S A 8 includes Example S A 1 - S A A computing system according to any one of 7, wherein, for each of the plurality of image patches, a parameter associated with at least one augmentation component is randomly selected.
[0104] Example S A 9 includes Example S A 1 - S A A computing system according to any one of 8, wherein, for each of the plurality of image patches, the at least one augmentation component is selected based on a predetermined queue.
[0105] Example S A 10 includes Example S A 1 - S A A computing system according to any one of 9, wherein, for each image patch having a plurality of selected augmentation components, the outputs for each of the selected augmentation components are mixed to provide an augmented image patch.
[0106] Example S A 11 includes Example S A 1 - S A A computing system according to any one of 10, wherein at least two image patches overlap within an image block, and wherein the overlapping regions of at least two image blocks are mixed together after augmentation.
[0107] Example S A 12 includes Example S A 1 - S A A computing system according to any one of 11, wherein a neural network is trained using a plurality of augmented training image blocks generated from a plurality of training image blocks, and wherein, for at least one augmented training image block, at least two augmented image patches have different augmentations.
[0108] Example C A 1 includes at least one non - transitory computer - readable storage medium including instructions that, when executed by a computing system, cause the computing system to perform operations including: dividing an image block into a plurality of image patches, each image patch having a size, shape, and position relative to the image block; generating an augmented image block by applying variable image augmentation to each image patch, wherein at least two augmented image patches have different augmentations; and performing an image analysis task by applying a neural network to the augmented image block.
[0109] Example CA 2 includes Example C A At least one non-transitory computer-readable storage medium of 1, wherein the instructions, when executed, cause the computing system to perform further operations, the further operations including generating a plurality of augmented image patches by repeatedly dividing and generating operations on the image patches, wherein each augmented image patch has a unique combination of augmented image slices relative to other augmented image patches, wherein. Performing the image analysis task includes applying a neural network to each of the plurality of augmented image patches to obtain a corresponding output of the neural network for each augmented image patch, and determining a result of the image processing task based on combining the corresponding outputs of the neural network.
[0110] Example C A 3 includes Example C A 1 or C A At least one non-transitory computer-readable storage medium of 2, wherein for each of the plurality of image slices, one or more of the size, shape, or position of the corresponding image slice are randomly selected.
[0111] Example C A 4 includes Example C A 1, C A 2 or C A At least one non-transitory computer-readable storage medium of 3, wherein the image slices jointly represent at least a threshold amount of image patches.
[0112] Example C A 5 includes Example C A 1-C A At least one non-transitory computer-readable storage medium of any one of 4, wherein the variable image augmentation includes at least one augmentation component selected from a plurality of augmentation components.
[0113] Example C A 6 includes Example C A 1-C A At least one non-transitory computer-readable storage medium of any one of 5, wherein the plurality of augmentation components includes one or more of a filter, intensity adjustment, noise generator, or style transfer.
[0114] Example C A 7 includes Example C A 1-C A At least one non-transitory computer-readable storage medium of any one of 6, wherein for each of the plurality of image slices, at least one augmentation component is randomly selected.
[0115] Example C A 8 includes Example C A 1-C AAt least one non-transitory computer-readable storage medium according to any one of 7, wherein, for each of the plurality of image patches, parameters associated with at least one augmentation component are randomly selected.
[0116] Example C A 9 includes Example C A 1-C A At least one non-transitory computer-readable storage medium according to any one of 8, wherein, for each of the plurality of image patches, the at least one augmentation component is selected based on a predetermined queue.
[0117] Example C A 10 includes Example C A 1-C A At least one non-transitory computer-readable storage medium according to any one of 9, wherein, for each image patch having a plurality of selected augmentation components, the outputs for each selected augmentation component are mixed to provide an augmented image patch.
[0118] Example C A 11 includes Example C A 1-C A At least one non-transitory computer-readable storage medium according to any one of 10, wherein at least two image patches overlap within an image block, and wherein the overlapping regions of the at least two image patches are mixed together after augmentation.
[0119] Example C A 12 includes Example C A 1-C A At least one non-transitory computer-readable storage medium according to any one of 11, wherein a neural network is trained using a plurality of augmented training image blocks generated from a plurality of training image blocks, and wherein, for at least one augmented training image block, at least two augmented image patches have different augmentations.
[0120] Example M B 1 includes a computer-implemented method, the method comprising: dividing each of a plurality of image blocks into a plurality of image patches, each image patch having a size, a shape, and a position relative to the corresponding image block; generating a plurality of augmented image blocks by applying variable image augmentation to each image patch of each image block, wherein, for at least one augmented image block, at least two augmented image patches have different augmentations; and training a neural network using the plurality of augmented image blocks as training images.
[0121] Example M B 2 includes Example M B The method of 1, wherein, for each of the plurality of image patches, one or more of the size, shape, or position of the corresponding image patch are randomly selected.
[0122] Example M B 3 includes Example M B 1 or M B 2's method, wherein for each image patch, the image slices of the image patch jointly represent at least a threshold amount of the image patch.
[0123] Example M B 4 includes Example M B 1, M B 2 or M B 3's method, wherein the variable image augmentation includes at least one augmentation component selected from a plurality of augmentation components.
[0124] Example M B 5 includes Example M B 1 - M B 4's method, wherein the plurality of augmentation components includes one or more of a filter, intensity adjustment, noise generator, or style transformation.
[0125] Example M B 6 includes Example M B 1 - M B 5's method, wherein for each of the plurality of image slices, at least one augmentation component is randomly selected.
[0126] Example M B 7 includes Example M B 1 - M B 6's method, wherein for each of the plurality of image slices, a parameter associated with at least one augmentation component is randomly selected.
[0127] Example M B 8 includes Example M B 1 - M B 7's method, wherein for each of the plurality of image slices, at least one augmentation component is selected based on a predetermined queue.
[0128] Example M B 9 includes Example M B 1 - M B 8's method, wherein for each image slice having a plurality of selected augmentation components, the outputs for each selected augmentation component are mixed to provide an augmented image slice.
[0129] Example M B 10 includes Example M B 1 - M B 9's method, wherein at least two image slices overlap within the image patch, and wherein the overlapping regions of at least two image slices are mixed together after augmentation.
[0130] Example S B 1 includes a computing system, the computing system includes a processor and a memory coupled to the processor, the memory includes instructions, the instructions when executed by the processor cause the computing system to perform operations, the operations include: dividing each of a plurality of image blocks into a plurality of image patches, each image patch having a size, a shape, and a position relative to the corresponding image block; generating a plurality of augmented image blocks by applying variable image augmentation to each image patch of each image block, wherein, for at least one augmented image block, at least two augmented image patches have different augmentations; and using the plurality of augmented image blocks as training images to train a neural network.
[0131] Example S B 2 includes Example S B The computing system of 1, wherein, for each of the plurality of image patches, one or more of the size, shape, or position of the corresponding image patch are randomly selected.
[0132] Example S B 3 includes Example S B 1 or S B The computing system of 1 or 2, wherein, for each image block, the image patches of the image block together represent at least a threshold amount of the image block.
[0133] Example S B 4 includes Example S B 1, S B 2 or S B The computing system of 1, 2, or 3, wherein the variable image augmentation includes at least one augmentation component selected from a plurality of augmentation components.
[0134] Example S B 5 includes Example S B 1 - S B The computing system of any one of 1 - 4, wherein the plurality of augmentation components includes one or more of a filter, intensity adjustment, noise generator, or style transfer.
[0135] Example S B 6 includes Example S B 1 - S B The computing system of any one of 1 - 5, wherein, for each of the plurality of image patches, at least one augmentation component is randomly selected.
[0136] Example S B 7 includes Example S B 1 - S B The computing system of any one of 1 - 6, wherein, for each of the plurality of image patches, a parameter associated with at least one augmentation component is randomly selected.
[0137] Example S B 8 includes Example S B 1 - S B A computing system according to any one of 7, wherein for each image patch among a plurality of image patches, at least one augmentation component is selected based on a predetermined queue.
[0138] Example S B 9 includes Example S B 1 - S B A computing system according to any one of 8, wherein for each image patch having a plurality of selected augmentation components, the outputs of the augmentation components for each selection are mixed to provide an augmented image patch.
[0139] Example S B 10 includes Example S B 1 - S B A computing system according to any one of 9, wherein at least two image patches overlap within an image block, and wherein the overlapping regions of at least two image patches are mixed together after augmentation.
[0140] Example C B 1 includes at least one non - transitory computer - readable storage medium comprising instructions that, when executed by a computing system, cause the computing system to perform operations including: dividing each of a plurality of image blocks into a plurality of image patches, each image patch having a size, shape, and position relative to the corresponding image block; generating a plurality of augmented image blocks by applying a variable image augmentation to each image patch of each image block, wherein for at least one augmented image block, at least two augmented image patches have different augmentations; and using the plurality of augmented image blocks as training images to train a neural network.
[0141] Example C B 2 includes Example C B The at least one non - transitory computer - readable storage medium of 1, wherein for each image patch among a plurality of image patches, one or more of the size, shape, or position of the corresponding image patch are randomly selected.
[0142] Example C B 3 includes Example C B 1 or C B The at least one non - transitory computer - readable storage medium of 2, wherein for each image block, the image patches of the image block together represent at least a threshold amount of the image block.
[0143] Example C B 4 includes Example C B 1, C B 2 or C BAt least one non-transitory computer-readable storage medium of 3, wherein the variable image augmentation includes at least one augmentation component selected from a plurality of augmentation components.
[0144] Example C B 5 includes Example C B 1-C B At least one non-transitory computer-readable storage medium of any one of 4, wherein the plurality of augmentation components includes one or more of a filter, intensity adjustment, noise generator, or style transformation.
[0145] Example C B .6 includes Example C B 1-C B At least one non-transitory computer-readable storage medium of any one of 5, wherein for each image patch among a plurality of image patches, at least one augmentation component is randomly selected.
[0146] Example C B 7 includes Example C B 1-C B At least one non-transitory computer-readable storage medium of any one of 6, wherein for each image patch among a plurality of image patches, a parameter associated with at least one augmentation component is randomly selected.
[0147] Example C B 8 includes Example C B 1-C B At least one non-transitory computer-readable storage medium of any one of 7, wherein for each image patch among a plurality of image patches, at least one augmentation component is selected based on a predetermined queue.
[0148] Example C B 9 includes Example C B 1-C B At least one non-transitory computer-readable storage medium of any one of 8, wherein for each image patch having a plurality of selected augmentation components, the outputs for each selected augmentation component are mixed to provide an augmented image patch.
[0149] Example C B 10 includes Example C B 1-C B At least one non-transitory computer-readable storage medium of any one of 9, wherein at least two image patches overlap within an image block, and wherein the overlapping regions of the at least two image patches are mixed together after augmentation.
[0150] The embodiments are applicable to all types of semiconductor integrated circuit ("IC") chips. Examples of such IC chips include, but are not limited to, processors, controllers, chipset components, programmable logic arrays (PLAs), memory chips, network chips, system-on-a-chip (SoC), SSD / NAND controller ASICs, etc. Additionally, in some of the figures, signal conductors are represented by lines. Some may be different to indicate more constituent signal paths, some have numerical labels to indicate the number of constituent signal paths, and / or some have arrows at one or more ends to indicate the primary information flow direction. However, this should not be construed in a limiting manner. Instead, such additional details can be used in conjunction with one or more exemplary embodiments to facilitate easier understanding of the circuitry. Regardless of whether there is additional information, any represented signal line may actually include one or more signals that can travel in multiple directions and can be implemented with any suitable type of signal scheme, e.g., digital or analog lines implemented with differential pairs, fiber optic lines, and / or single-ended lines.
[0151] Example sizes / models / values / ranges may be given, but the embodiments are not limited thereto. As manufacturing technology (e.g., lithography) matures over time, it is expected that devices of smaller sizes can be manufactured. Additionally, to simplify the description and discussion, power / ground connections to the IC chips and other components that are known may or may not be shown in the figures so as not to obscure certain aspects of the embodiments. Further, arrangements may be shown in block diagram form to avoid obscuring the embodiments, and also in view of the fact that details regarding the implementation of such block diagram arrangements highly depend on the platform in which the embodiments are to be implemented, i.e., such details should be within the purview of those skilled in the art. In cases where specific details (e.g., circuitry) are set forth to describe example embodiments, it should be apparent to those skilled in the art that the embodiments can be practiced without these specific details or with variations of these specific details. Accordingly, the description is to be regarded as illustrative rather than restrictive.
[0152] The term "coupled" as used herein can be used to refer to any type of direct or indirect relationship between the components being discussed and can apply to electrical, mechanical, fluid, optical, electromagnetic, electromechanical, or other connections, including logical connections via intermediate components (e.g., device A can be coupled to device C via device B). Additionally, the terms "first", "second", etc. as used herein can be used only to facilitate discussion and do not have a particular temporal or chronological significance unless otherwise stated.
[0153] As used in this application and the claims, a list of items joined by the term "one or more of" can mean any combination of the listed terms. For example, the phrase "one or more of A, B, or C" can mean A, B, C; A and B; A and C; B and C; or A, B, and C.
[0154] Those skilled in the art will understand from the foregoing description that the broad techniques of the embodiments can be implemented in various forms. Accordingly, while the embodiments have been described in connection with their specific examples, the true scope of the embodiments should not be so limited, as other modifications will become apparent to those skilled in the art upon study of the drawings, the specification, and the appended claims.
Claims
1. A computer-implemented method, comprising: Dividing an image patch into a plurality of image tiles, each image tile having a size, a shape, and a position relative to the image patch; Generating an augmented image patch by applying variable image augmentation to each image tile, wherein at least two augmented image tiles have different augmentations; And Performing an image analysis task by applying a neural network to the augmented image patch.
2. The method according to claim 1, further comprising: Generating a plurality of augmented image patches by repeating the division and generation operations on the image patch, wherein each augmented image patch has a unique combination of augmented image tiles relative to other augmented image patches; Wherein performing the image analysis task comprises: Applying the neural network to each of the plurality of augmented image patches to obtain a corresponding output of the neural network for each augmented image patch; and Determining the result of the image processing task based on a combination of the corresponding outputs of the neural network.
3. The method according to claim 1, wherein For each of the plurality of image tiles, randomly selecting one or more of the size, the shape, or the position of the corresponding image tile.
4. The method according to claim 1, wherein, The image tiles together represent at least a threshold amount of the image patch.
5. The method according to claim 1, wherein, The variable image augmentation includes at least one augmentation component selected from a plurality of augmentation components.
6. The method according to claim 5, wherein, The plurality of augmentation components include one or more of a filter, intensity adjustment, a noise generator, or style transfer.
7. The method according to claim 5, wherein, For each of the plurality of image tiles, randomly selecting the at least one augmentation component.
8. The method according to claim 5, wherein For each of the plurality of image tiles, randomly selecting a parameter associated with the at least one augmentation component.
9. The method according to claim 5, wherein, For each of the plurality of image tiles, selecting the at least one augmentation component based on a predetermined queue.
10. The method according to claim 5, wherein, For each image tile having a plurality of selected augmentation components, mixing the outputs for each of the selected augmentation components to provide an augmented image tile.
11. The method according to claim 1, wherein, At least two image tiles overlap within the image patch, and wherein the overlapping regions of the at least two image tiles are mixed together after augmentation.
12. The method according to claim 1, wherein, Training the neural network using a plurality of augmented training image patches generated from a plurality of training image patches, wherein for at least one augmented training image patch, at least two augmented image tiles have different augmentations.
13. A computing system, comprising: A processor; And A memory coupled to the processor, the memory including instructions that, when executed by the processor, cause the computing system to perform the following operations, the operations including: Dividing an image patch into a plurality of image tiles, each image tile having a size, a shape, and a position relative to the image patch; Generating an augmented image patch by applying variable image augmentation to each image tile, wherein at least two augmented image tiles have different augmentations; and Performing an image analysis task by applying a neural network to the augmented image patch.
14. The computing system according to claim 13, wherein, The instructions, when executed, cause the computing system to perform further operations, the further operations including: Generating a plurality of augmented image patches by repeating the division and generation operations on the image patches, wherein each augmented image patch has a unique combination of augmented image slices relative to other augmented image patches; wherein performing the image analysis task includes: applying the neural network to each of the plurality of augmented image patches to obtain a corresponding output of the neural network for each augmented image patch; and determining a result of the image processing task based on combining the corresponding outputs of the neural network.
15. The computing system according to claim 13, wherein, For each of the plurality of image slices, randomly selecting one or more of the size, shape, or position of the corresponding image slice, and wherein the image slices together represent at least a threshold amount of the image patch.
16. The computing system according to claim 13, wherein, The variable image augmentation includes at least one augmentation component selected from a plurality of augmentation components, and wherein the plurality of augmentation components includes one or more of a filter, intensity adjustment, noise generator, or style transfer.
17. A computer-implemented method, comprising: dividing each of a plurality of image patches into a plurality of image slices, each image slice having a size, shape, and position relative to the corresponding image patch; generating a plurality of augmented image patches by applying variable image augmentation to each image slice of each image patch for each image patch, wherein for at least one augmented image patch, at least two augmented image slices have different augmentations; and using the plurality of augmented image patches as training images to train a neural network.
18. The method according to claim 17, wherein, For each of the plurality of image slices, randomly selecting one or more of the size, shape, or position of the corresponding image slice, and wherein for each image patch, the image slices of the image patch together represent at least a threshold amount of the image patch.
19. The method according to claim 17, wherein, The variable image augmentation includes at least one augmentation component selected from a plurality of augmentation components, and wherein the plurality of augmentation components includes one or more of a filter, intensity adjustment, noise generator, or style transfer.
20. The method according to claim 19, wherein, For each image slice having a plurality of selected augmentation components, mixing the outputs for each of the selected augmentation components to provide an augmented image slice.