System and method for training a machine model with augmented data

By preserving camera characteristics in augmented images, the method addresses the challenge of model overfitting due to sensor variations, improving model robustness and accuracy across devices with similar sensor configurations.

JP2026012763APending Publication Date: 2026-01-27TESLA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025173584
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-10-11
Filing Date
2025-10-15
Publication Date
2026-01-27

AI Technical Summary

Technical Problem

Existing machine learning models struggle to effectively learn from images captured with varying sensor characteristics, leading to overfitting and reduced robustness when deployed across different devices with diverse sensor configurations.

Method used

The method involves generating augmented images that preserve camera characteristics by using image manipulation functions that maintain intrinsic and extrinsic camera properties, such as scale, orientation, and reflections, while training predictive computer models.

Benefits of technology

This approach enhances model generalization and robustness, allowing the trained model to perform accurately across devices with consistent sensor configurations, particularly in applications like object detection and autonomous vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026012763000001_ABST
    Figure 2026012763000001_ABST
Patent Text Reader

Abstract

To provide a system and method for training a machine model with augmented data.SOLUTION: The method includes identifying a set of images captured by a set of cameras while fixed to one or more image collection systems. For each image in the set of images, a training output for the image is identified. An extended image of a set of extended images is generated for one or more images in the set of images. Generating the augmented image includes modifying the image with an image manipulation function that maintains camera characteristics of the image. The augmented training image is associated with the training output of the image. A set of parameters of the prediction computer model is trained to predict a training output based on an image training set including the image and the set of augmented images.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims priority to U.S. Provisional Patent Application No. 62 / 744,534, filed October 11, 2018, entitled "Training Machine Models with Data Augmentation that Retains Sensor Characteristics," which is incorporated herein by reference in its entirety.

[0002] Embodiments of the present invention relate generally to systems and methods for training data in a machine learning environment, and more particularly to augmenting training data by including additional data, such as sensor characteristics, in the training dataset. [Background technology]

[0003] In typical machine learning applications, data can be augmented in various ways to avoid overfitting a model to the characteristics of the capture equipment used to acquire the training data. For example, in a typical image set used to train a computer model, the images can represent objects captured in many different capture environments with different sensor characteristics relative to the objects being captured. For example, such images may be captured with different sensor characteristics, such as different scales (e.g., significantly different distances within the image), different focal lengths, different lens types, different pre- or post-processing, different software environments, sensor array hardware, etc. These sensors may also differ with respect to various extrinsic parameters, such as the position and orientation of the imaging sensor relative to the environment when the image is captured. All of these different types of sensor characteristics can cause the captured images to present differently and variedly across multiple different images in the image set, making it more difficult to properly train the computer model.

[0004] Many applications of neural networks learn from data captured in a variety of conditions and are deployed with a variety of different sensor configurations (e.g., in apps running on multiple types of mobile phones). To account for differences in the sensors used to capture images, developers can augment the image training data with modifications such as flipping, rotating, or cropping the images, which generalizes the developed model with respect to camera characteristics such as focal length, axis skew, position, and rotation.

[0005] To account for these variations and to deploy the trained network on a variety of sources, the training data can be augmented or manipulated to increase the robustness of the trained model. However, these approaches typically apply transformations that change the camera properties in the augmented images, preventing the model from effectively learning about any particular camera configuration. Summary of the Invention [Means for solving the problem]

[0006] One embodiment is a method for training a set of parameters of a predictive computer model that can include identifying a set of images captured by a set of cameras while fixed to one or more image acquisition systems, identifying, for each image in the set of images, a training output for the image, generating, for one or more images in the set of images, an augmented image of the set of augmented images by modifying the image with an image manipulation function that preserves camera characteristics of the image, and generating an augmented image of the set of augmented images by associating the augmented training image with the training output for the image, and training the set of parameters of the predictive computer model to predict the training output based on the image training set that includes the image and the set of augmented images.

[0007] An additional embodiment may include a system having one or more processors and a non-transitory computer storage medium storing instructions that, when executed by the one or more processors, cause the processors to perform operations including identifying a set of images captured by a set of cameras while fixed to one or more image acquisition systems; for each image in the set of images, identifying a training output for the image; for one or more images in the set of images, generating an augmented image of the set of augmented images by modifying the image with an image manipulation function that preserves camera characteristics of the image, and generating an augmented image of the set of augmented images by associating the augmented training image with the training output for the image; and training a set of parameters of a predictive computer model to predict the training output based on the image training set including the image and the set of augmented images.

[0008] Another embodiment includes a non-transitory computer-readable medium having instructions for execution by a processor, the instructions, when executed by the processor, causing the processor to: identify a set of images captured by a set of cameras while fixed to one or more image acquisition systems; for each image in the set of images, identify a training output for the image; for one or more images in the set of images, generate an augmented image of the set of augmented images by modifying the image with an image manipulation function that preserves camera characteristics of the image, and generate an augmented image of the set of augmented images by associating the augmented training image with the training output for the image; and train a computer model to predict the training output based on the image training set including the image and the set of augmented images. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a block diagram of an environment for training and deploying computer models according to one embodiment.

[0010] [Figure 2] 1A and 1B show exemplary images captured with the same camera characteristics.

[0011] [Figure 3] FIG. 1 is a block diagram of components of a model training system, according to one embodiment.

[0012] [Figure 4] FIG. 1 is a data flow diagram illustrating an example of generating augmented images based on labeled training images, according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0013] The drawings depict various embodiments of the present invention for purposes of illustration only. Those skilled in the art will readily recognize from the following discussion that alternative embodiments of the structures and methods shown herein may be utilized without departing from the principles of the present invention as described herein.

[0014] One embodiment is a system that trains a computer model using augmented images to preserve camera characteristics of the originally captured image. These camera characteristics can include intrinsic or extrinsic characteristics of the camera. Such intrinsic characteristics can include characteristics of the sensor itself, such as dynamic range, field of view, focal length, and lens distortion. Extrinsic characteristics can represent the configuration of the camera relative to the captured environment, such as the camera's angle, scale, or pose.

[0015] These intrinsic and extrinsic characteristics can affect the camera's view of objects and other aspects captured in an image, as well as artifacts and other effects, such as stationary objects that appear in the camera's view due to their positioning on the device or system. For example, a camera mounted on a vehicle may include as part of its view the car's hood, which appears across many images and for all cameras of that configuration mounted in the same way on the same model of car. As another example, these camera characteristics may also include reflections resulting from objects in the camera's view. Reflections may be one type of consistent characteristic that becomes included in many of the images captured by the camera.

[0016] By maintaining, saving, storing, or using the camera characteristics of the images to train the data model while further adding to the training data with the augmented images, the resulting model can be useful across many different devices with the same camera characteristics. Furthermore, augmentation can provide generalization and greater robustness to model predictions, especially when the image is cloudy, occluded, or otherwise does not provide a clear view of detectable objects. These techniques can be particularly useful for object detection and autonomous vehicles. This technique can also be beneficial in other situations where the same camera configuration may be deployed on many devices. Because these devices can have a consistent set of sensors in a consistent orientation, training data can be collected with a given configuration, the model can be trained with augmented data from the collected training data, and the trained model can be deployed to devices with the same configuration. Thus, these techniques avoid augmentation, which results in unnecessary generalization in this context and allows for generalization of other variables with some data augmentation.

[0017] To preserve camera characteristics, image manipulation functions used to generate augmented images are functions that preserve camera characteristics. For example, these manipulations may avoid affecting the angle, scale, or attitude of the camera relative to the captured environment. In embodiments, images augmented by image manipulation functions that affect camera characteristics are not used for training. For example, image manipulation functions that may be used to preserve camera characteristics include cropping, hue / saturation / value jitter, salt and pepper, and domain transitions (e.g., changing from day to night). Functions that may alter camera characteristics and therefore are not used in some embodiments include cropping, padding, flipping (horizontal or vertical), or affine transformations (such as shear, rotate, translate, and skew).

[0018] As a further example, an image can be augmented with a "crop" function that removes portions of the original image. The removed portions of the image can then be replaced with other image content, such as a specified color, blur, noise, or content from another image. The number, size, area, and replacement content of the crops can be varied and can be based on image labels (e.g., regions of interest within the image, or bounding boxes of objects).

[0019] Thus, a computer model can be trained with the images and augmented images and deployed to a device with the camera characteristics of the captured images so that the model can be used for sensor analysis. In particular, this data augmentation and model training can be used for models trained to detect objects or object bounding boxes in images.

[0020] 1 illustrates an environment for training and deploying computer models according to one embodiment. One or more image acquisition systems 140 capture images that can be used by a model training system in training computer models that can be deployed and used by a model application system. These systems are connected via a network 120, such as the Internet, which represents the various wireless or wired communication links through which these devices communicate.

[0021] The model training system 130 trains a computer model having a set of trainable parameters to predict an output given a set of inputs. The model training system 130 in this example typically trains a model based on image inputs to generate output prediction information about the image. For example, in various embodiments, these outputs may identify objects in the image (either by bounding box or segmentation), identify the state of the image (e.g., time of day, weather), or other tags or descriptors of the image.

[0022] For convenience, images are used herein as an exemplary type of sensor data, however, the augmentation and model development described herein can be applied to various types of sensors to augment the training data captured from these sensors while maintaining sensor configuration characteristics.

[0023] Image collection system 140 has a set of sensors that capture information from the environment of image collection system 140. While a single image collection system 140 is shown, many image collection systems 140 can capture images for model training system 130. The sensors for image collection systems 140 have sensor characteristics that can be the same or substantially the same across image collection systems 140. In one embodiment, image collection system 140 is a vehicle or other system that moves through the environment and captures images of the environment with a camera. Image collection system 140 may be operated manually or by a partially or fully automated vehicle. Thus, as image collection system 140 moves through the environment, image collection system 140 can capture and transmit images of the environment to model training system 130.

[0024] The model application system 110 is a system having a set of sensors with the same or substantially the same sensor characteristics as the image collection system. In some examples, the model application system 110 also functions as the image collection system 130, providing captured sensor data (e.g., images) to the model training system 130 for use as further training data. The model application system 110 receives the trained model from the model training system 130 and uses the model with the data sensed by its sensors. Because the images captured from the image collection system 140 and the model application system 110 have the same camera configuration, the model application system 110 can capture its environment in the same manner and from the same perspective (or substantially similar) as the image collection system. After applying the model, the model application system 110 can use the model's output for various purposes. For example, if the model application system 110 is a vehicle, the model can predict the presence of objects in the images, which can be used by the model application system 110 as part of a safety system or as part of an autonomous (or semi-autonomous) control system.

[0025] FIG. 2 illustrates exemplary images captured with the same camera characteristics. In this example, image 200A is captured by a camera on image acquisition system 130. Another image 200B may also be captured by image acquisition system 130, which may be the same image acquisition system or a different image acquisition system 130. While capturing different environments and different objects within the environments, these images maintain the camera characteristics relative to the image capturing the environment. Camera characteristics refer to the configuration and orientation characteristics of the camera that affect how the environment appears within the camera. For example, these camera characteristics may include the angle, scale, and attitude (e.g., viewing position) of the camera relative to the environment. Changing the angle, scale, or position of the camera relative to the same environment from which the images are captured will change the image of the environment. For example, a camera positioned higher will view an object from a different height and show a different portion of the object's underside than a camera positioned lower down. Similarly, these images contain consistent artifacts and effects within the image that are due to the camera configuration that are not part of the environment being analyzed. For example, both images 200A and 200B contain glare and other effects from the windshield, objects on the lower right side of the image occlude the environment, and the windshield occludes the bottom of the image. Thus, images captured from the same camera characteristics typically exhibit the same artifacts, distortions, and capture the environment in the same way.

[0026] FIG. 3 illustrates components of a model training system 130, according to one embodiment. The model training system includes various modules and data stores for training a computer model. The model training system 130 trains the model used by the model application system 110 by augmenting images from the image acquisition system 140 to improve model generalization. The augmented images are generated using image manipulation functions that do not affect (e.g., preserve) the camera configuration of the images. This enables more effective modeling, allowing model parameters to more closely learn weights related to consistent camera characteristics while allowing generalization of model parameters that are more selective and avoid overfitting for aspects of the images that may vary between images.

[0027] The model training system includes a data input module 310 that receives images from the image collection system 140. The data input module 310 can store these images in an image data store 350. The data input module 310 may receive the images as they are generated or provided by the data collection system 140, or may request the images from the image collection system 140.

[0028] The labeling module 320 can identify or apply labels to images in the image data 350. In some examples, the images may already have identified characteristics. Labels can also represent data predicted or output by a trained model. For example, labels can designate specific objects in the environment shown in the image or can include descriptors or "tags" associated with the image. Depending on the application of the model, labels can represent this information in various ways. For example, objects may be associated with bounding boxes in the image, or objects may be segmented from other parts of the image. Thus, labeled images can represent ground truth against which a model is trained. Images may be labeled by any suitable means, typically by a supervised labeling process (e.g., by a user reviewing images and assigning labels to them). These labels can then be associated with images in the image data store 350.

[0029] The image augmentation module 330 can generate additional images based on images captured by the image collection system 140. These images may be generated as part of the training pipeline of the model training module 340, or these augmented images may be generated before training begins in the model training module 340. The augmented images can be generated based on images captured by the image collection system 140.

[0030] 4 illustrates an example of generating augmented images based on labeled training images 400, according to one embodiment. The labeled training images may be images captured by image acquisition system 140. The training images 410 may include unaugmented training images 410A with associated training outputs 420A that correspond to the labeled data in labeled training images 400.

[0031] The image augmentation module 330 generates augmented images by applying image manipulation functions to the labeled training images 400. The image manipulation functions generate modified versions of the labeled training images 400 to change image characteristics for training a model. The image manipulation functions used to generate training images preserve the camera characteristics of the labeled training images 400. Thus, the manipulation functions can preserve the scale, perspective, orientation, and other characteristics of the view of the environment that may be affected by the camera's physical capture characteristics or the camera's position when capturing the environment, which may be consistent across various devices. Thus, the image manipulation functions can affect how visible or distinctly visible an object or other feature of the environment is in the scene, but may not affect the position or size of the object in the image. Exemplary image manipulation functions that preserve camera characteristics and can be applied include cropping, jittering (e.g., of hue, saturation, or color value), salt and pepper (introducing black and white dots), blurring, and domain transition. Multiple of these image manipulation functions can be applied in combination to generate the augmented images. Cropping refers to an image manipulation function that removes part of an image and replaces it with other image content. Domain transition refers to an image manipulation function that modifies an image to accommodate different environmental conditions within the image. For example, a daytime image can be modified to approximate how the image would look at night, or an image taken in sunny conditions can be modified to add a rain or snow effect.

[0032] These augmented images can be associated with the same training output as labeled training image 400. In the example shown in Figure 4, augmented image 410B is generated by applying crops to labeled training image 400, and augmented image 410B can be associated with training output 420B. Similarly, to generate training image 410C, multiple crops are applied to modify multiple portions of the image. In this example, the crops applied to generate training image 410C fill the cropped regions of the image with different patterns.

[0033] In various embodiments, crops may be applied using various parameters and configurations that can vary based on the training image and the position of the training output within the image. Thus, the number, size, location, and replacement image content of crops may vary based on the position of the training output in different embodiments. By way of example, the crop function may apply multiple crops of similar size, or may apply several crops of different semi-randomized sizes within a range. By using multiple crops and varying their sizes, the crops can more closely simulate the effect that real-world obstacles (of various sizes) have on viewing an object, preventing the trained model from learning to compensate for any one particular size crop.

[0034] The range of crop sizes may be based in part on the size of the objects or other labels in the image. For example, the crop may be no more than 40% of the size of the bounding box of the object in the image, or smaller than the bounding box of the smallest object. This ensures that the crop does not completely obscure the target object, and therefore the image continues to contain image data of the object from which the model can learn. The number of crops may also be randomized and selected from a distribution such as a uniform, Gaussian, or exponential distribution.

[0035] Additionally, the location of the crop may be selected based on the location of the object within the image. This may result in some, but not excessive, overlap with the bounding box. The intersection between the object and the crop region may be measured by the portion of the object displaced by the crop, or by the intersection over union (IoU), which may be measured by dividing the intersection of the object with the crop region by the union of the object's area and the crop region. For example, the crop region may be positioned to have an intersection over union value in the range of 20% to 50%. Thus, by including some, but not too much, of the object in the crop, the crop can create more "challenging" instances of partially obscuring the object without removing too much relevant image data. Similarly, the crop may be selected for a specific portion of the image based on the camera's predicted view within the image. For example, the crop may be located primarily in the lower half of the image or in the center of the image, because the center of the image may be the area of ​​most interest (e.g., in the case of a vehicle, this is often in the direction of the vehicle's travel), while the bottom may typically contain artifacts that are always present.

[0036] The replacement image data for the cutout region may be a solid color (e.g., a constant) or another pattern, such as Gaussian noise. As another example, to represent an occlusion or other obstacle, the cutout may be replaced with a patch of image data from another image having the same image type or label. Finally, the cutout may be composited with a region near the cutout, for example, by Poisson composition. By using various composition techniques, such as background patching or composition, they can ensure that the replacement data within the cutout is more difficult to distinguish from the environment, thus providing an example that more closely resembles a real-world obstacle.

[0037] 4 as rectangular regions, the crops applied in generating the augmented images may vary in shape in other embodiments. After generating the augmented images 410B, 410C and associating them with the associated training outputs 420B, 420C, the image augmentation module 330 may add these images to the image data store 350.

[0038] The model training module 340 trains a computer model based on images captured by the image acquisition system 140 and the augmented images generated by the image augmentation module 330. These images can be used as an image training set for model training. In one embodiment, the machine learning model is a neural network model, such as a feedforward network, a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), or a self-organizing map (SOM), trained by the model training module 340 based on training data. After training, the computer model can be stored in the trained computer model store 370. The model receives sensor data (e.g., images) as input and outputs output predictions according to the model's training. When training a model, the model learns (or "trains") a set of parameters that predict outputs based on input images, as evaluated by a loss function of the training data. That is, during training, the training data is evaluated according to the current parameter set to generate predictions. Its predictions for training inputs can be compared to specified outputs (e.g., labels) to evaluate loss (e.g., using a loss function), and parameters can be modified via an optimization algorithm to optimize a set of parameters so as to reduce the loss function. Although called "optimization," these algorithms can reduce loss with respect to a set of parameters, but may not be guaranteed to find the "optimal" values ​​of the parameters given a set of inputs. For example, a gradient descent optimization algorithm may seek a local minimum rather than a global minimum.

[0039] By training a computer model on augmented training data, the computer model can perform with improved accuracy when applied to sensor data from a physical sensor operating in an environment that has the sensor characteristics of the data being captured. These sensor characteristics (e.g., camera characteristics) are represented in the images used to train the data, so that the augmentation preserves these characteristics. In one embodiment, the training data does not include augmented images generated by image manipulation functions that alter the camera characteristics of the images, such as operations that crop, pad, flip (vertically or horizontally), or apply an affine transformation (e.g., shear, rotate, translate, skew) to the image.

[0040] After training, the model distribution module 380 can distribute the trained model to systems for applying the trained model. In particular, the model distribution module 380 can send the trained model (or its parameters) to the model application system 110 for use in detecting characteristics of images based on sensors of the model application system 110. Thus, predictions from the model can be used in the operation of the model application system 110, for example, in object detection and control of the model application system 110.

[0041] The foregoing description of embodiments of the invention has been presented for purposes of illustration and is not intended to be exhaustive or to limit the invention to the precise form disclosed. Those skilled in the art will recognize that many modifications and variations are possible in light of the above disclosure.

[0042] Some portions of this specification describe embodiments of the invention in terms of algorithms and symbolic representations of operations on information. These algorithmic descriptions and representations are commonly used by those skilled in the data processing arts to effectively convey the substance of their work to others skilled in the art. While these operations are described functionally, computationally, or logically, they should be understood to be implemented by computer programs or equivalent electrical circuits, microcode, or the like. Further, it has proven convenient at times to refer to arrangements of these operations as modules, without loss of generality. The described operations and their associated modules may be embodied in software, firmware, hardware, or any combination thereof.

[0043] Any of the steps, operations, or processes described herein may be performed or implemented by one or more hardware or software modules, alone or in combination with other devices. In one embodiment, the software modules are implemented by a computer program product comprising a computer-readable medium containing computer program code that can be executed by a computer processor to perform any or all of the described steps, operations, or processes.

[0044] Embodiments of the present invention may also relate to apparatus (e.g., systems) for performing the operations herein. This apparatus may be specially configured for the required purposes and / or may comprise a general-purpose computing device selectively activated or reconfigured by a computer program stored in a computer. The computing device may be a system or device of one or more processors and / or computer systems. Such computer programs may be stored on a non-transitory tangible computer-readable storage medium or any type of medium suitable for storing electronic instructions that may be coupled to a computer system bus. Furthermore, any computing system referred to herein may include a single processor or may be an architecture that utilizes a multiple processor design to increase computing power.

[0045] Embodiments of the present invention may also relate to products produced by the computational processes described herein. Such products may include information resulting from the computational processes, where the information is stored on a non-transitory tangible computer-readable storage medium and may include any embodiment of the computer program product or other data combination described herein.

[0046] Finally, the language used herein has been selected primarily for ease of reading and descriptive purposes, and may not have been selected to delineate or limit the subject matter of the present invention. Accordingly, it is intended that the scope of the invention be limited not by this detailed description, but by any claims that issue on an application based thereon. Accordingly, the disclosure of embodiments of the present invention is intended to be illustrative, but not limiting, of the scope of the invention, which is set forth in the appended claims.

Claims

1. 1. A method for training a set of parameters of a predictive computer model, comprising: identifying a set of images captured by a set of cameras while fixed to one or more image acquisition systems; for each image in the set of images, identifying a training output for that image; generating an augmented image of a set of augmented images for one or more images in the set of images, generating an augmented image of the set of augmented images by modifying the image with an image manipulation function that preserves camera characteristics of the image; and generating an augmented image of a set of augmented images by associating the augmented training images with the training output of the images; training a set of parameters of the predictive computer model to predict the training output based on an image training set including the image and the set of augmented images.

2. The method of claim 1 , wherein the training output is an object in the image.

3. The method of claim 1 , wherein the training set of images does not include images generated by image manipulation functions that change camera properties of the images.

4. The method of claim 3 , wherein the image manipulation functions that change camera characteristics include cropping, padding, horizontal or vertical flipping, or affine transformation.

5. The method of claim 1 , wherein the image manipulation function is crop, hue, saturation, value jitter, salt and pepper, domain transition, or any combination thereof.

6. The method of claim 1 , wherein the image manipulation function is a crop applied to the image based on the position of the training output within the image.

7. The method of claim 1 , wherein the image manipulation function is a crop applied to a portion of the image that is smaller than a bounding box of the training output.

8. The method of claim 1 , wherein the image manipulation function is a crop applied to a portion of the image that overlaps the location of the training output within the image.

9. 1. A system having one or more processors and a non-transitory computer storage medium storing instructions that, when executed by the one or more processors, cause the processors to: generating an augmented image of the set of augmented images that identifies a set of images captured by the set of cameras while fixed to one or more image acquisition systems; generating, for each image in the set of images, an augmented image of a set of augmented images that identify a training output for that image; generating an augmented image of a set of augmented images for one or more images in the set of images, generating an augmented image of the set of augmented images by modifying the image with an image manipulation function that preserves camera characteristics of the image; and generating an augmented image of a set of augmented images by associating the augmented training images with the training output of the images; training the set of parameters of the predictive computer model to predict the training output based on an image training set including the image and the set of augmented images.

10. The system of claim 9 , wherein the training set of images does not include images generated by image manipulation functions that alter camera properties of the images.

11. The system of claim 10 , wherein the image manipulation function that changes the camera characteristics includes cropping, padding, horizontal or vertical flipping, or affine transformation.

12. The system of claim 9 , wherein the image manipulation function is crop, hue, saturation, value jitter, salt and pepper, domain transition, or any combination thereof.

13. The system of claim 9 , wherein the image manipulation function is a crop applied to a portion of the image that overlaps the location of the training output within the image.

14. A non-transitory computer-readable medium having instructions for execution by a processor, the instructions, when executed by the processor, causing the processor to: identifying a set of images captured by a set of cameras while fixed to one or more image acquisition systems; for each image in the set of images, identifying a training output for that image; generating an augmented image of a set of augmented images for one or more images in the set of images, generating an augmented image of the set of augmented images by modifying the image with an image manipulation function that preserves camera characteristics of the image; and generating an augmented image of a set of augmented images by associating the augmented training images with the training output of the images; training the computer model to learn to predict the training output based on an image training set that includes the image and the augmented set of images.

15. The non-transitory computer-readable medium of claim 14 , wherein the training set of images does not include images generated by an image manipulation function that alters camera properties of the images.

16. 16. The non-transitory computer-readable medium of claim 15, wherein the image manipulation function that changes a camera characteristic comprises cropping, padding, horizontal or vertical flipping, or affine transformation.

17. 15. The non-transitory computer-readable medium of claim 14, wherein the image manipulation function is crop, hue, saturation, value jitter, salt and pepper, domain transition, or any combination thereof.

18. 15. The non-transitory computer-readable medium of claim 14, wherein the image manipulation function is a crop applied to a portion of the image that is smaller than a bounding box of the training output.

19. 15. The non-transitory computer-readable medium of claim 14, wherein the image manipulation function is a crop applied to a portion of the image that is smaller than a bounding box of the training output.

20. 15. The non-transitory computer-readable medium of claim 14, wherein the image manipulation function is a crop applied to a region in the image that overlaps with the location of the training output.