Method for processing image data to apply machine learning model
By converting image data into line-of-sight ray display and preprocessing, the machine learning model is trained using full convolution network and Dropout technology, and the cross-domain applicability problem under distortion of different camera devices is solved, and the general applicability and robustness of the model on different camera devices is achieved.
Patent Information
- Application Number
- CN202380084550.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-09
- Filing Date
- 2023-10-20
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, machine learning models have the problem of insufficient cross-domain generalization capabilities when processing distorted images of different imaging devices, resulting in poor applicability of models between different imaging devices.
By converting image data into ray ray display and preprocessing on ray ray display, including grid display and normalization, the machine learning model is trained using a full convolutional network and Dropout technology to make it distorted to the camera device unchanged.
The general applicability of machine learning models on different camera devices is realized, the need for large-scale retraining is reduced, and the robustness and adaptability of the model is improved.
Smart Images

Figure CN120345008A_ABST
Abstract
Description
Field of the Invention
[0001] The present invention relates to a method for processing image data for applying a machine learning model. Furthermore, the present invention also relates to a computer program and a device for this purpose. Background Art
[0002] It is known from the prior art that computer vision algorithms can be used to improve images captured by imaging devices, especially in vehicles. Here, adapting computer vision algorithms to different imaging device distortions is generally referred to as a domain adaptation problem. Herein, domain invariance is ensured by a specific training method in which the same CNN is trained on images with different imaging device distortions so that the resulting weights can be generalized across domains.
[0003] In "X. Peng, Y. L. Murphey, S. Stent, Y. Li, and Z. Zhao, 'Spatial focal loss for pedestrian detection in fisheye imagery,' 2019 IEEE Winter Conference on Applications of Computer Vision (WACV), pp. 561–569, 2019", specific types of costs are described that reweight samples with different imaging device distortions in order to facilitate the learning of domain-invariant features. Summary of the Invention
[0004] The subject matter of the present invention is a method having the features of claim 1, a computer program having the features of claim 9, and a device having the features of claim 10. Further features and details of the present invention result from the corresponding dependent claims, the description, and the drawings. Herein, the features and details described in connection with the method according to the invention of course also apply to the computer program according to the invention and the device according to the invention, and vice versa respectively, such that the disclosure regarding the individual aspects of the invention is always cross-referenced or can be cross-referenced to each other.
[0005] The subject matter of the present invention is in particular a method for processing, in particular preprocessing, image data for applying a machine learning model, the method comprising the following steps, which are preferably carried out successively and / or repeatedly:
[0006] - obtaining image data, wherein the image data is obtained from an image detection using an imaging device, preferably an imaging device of a vehicle,
[0007] - converting, preferably transforming (Konvertieren), the image data into a line-of-sight ray representation,
[0008] - Provide input for a machine learning model based on the image data in the line-of-sight ray display, preferably in order to apply the machine learning model through this input, preferably in order to analyze and evaluate the image data, for example for classification and / or segmentation.
[0009] The input is based on the line-of-sight ray display and thus in particular on a camera-independent representation of the image data. This has the advantage that a machine learning model, preferably in the form of an artificial neural network, can be applied independently of the specific camera device type. The vehicle can for example be configured as a motor vehicle, preferably a passenger car. The vehicle in particular has driving functions for autonomous driving and / or for assisting the vehicle driver. The driving functions can be implemented to control the vehicle based on the image data and / or the output of the applied machine learning model, for example steering and / or acceleration.
[0010] Furthermore, it can be envisaged that the image data in the line-of-sight ray display is represented by image points, where providing the input can include the following steps for further preprocessing:
[0011] - Provide a grid display in which one grid has a plurality of grid cells, where the image points are assigned to these grid cells.
[0012] Here, the grid density and the exact assignment can be hyperparameters of the method according to the invention. In the context of machine learning, a hyperparameter is in particular a parameter whose value is used to control the learning process. The geometry of the grid can be predefined if necessary and for example corresponds to the projection plane of the line-of-sight rays of the line-of-sight ray display, preferably corresponds to the image plane of the line-of-sight ray display.
[0013] The image points can be points on the grid into which the line-of-sight rays enter. In other words, the image points can each represent the intersection points of the line-of-sight rays of the line-of-sight ray display. Accordingly, the image points can also be referred to as ray intersection points.
[0014] Here, the expression "Sight Ray", in English, is used in the professional convention to show, for example, as described in Z. Zhang, Camera Parameters (Intrinsic, Extrinsic), pp. 81–85. Boston, MA: Springer US, 2014 or Ikeuchi, K. (eds) Computer Vision. Springer, Boston, MA. Essentially, this show describes such points at which the light rays of the points or pixels generated in the image data illuminate a hypothetical plane (also called the projection plane within the scope of the present invention) on the optical path. Therefore, this show is also called "Sight Ray show". The hypothetical plane is often also called the "image plane". It should be noted that: although it is called a "plane", any virtual surface, such as a spherical surface, can also be constructed as the image plane, and the sight ray show of the image, that is, the image data, relative to this surface can be calculated.
[0015] First, the image data can exist in the form of an image show, that is, for example, in the form of a pixel matrix show. Here, this can be an image data show as usually provided by a camera device. Subsequently, the image data can be converted into a sight ray show. In the next step, the sight ray show can be transformed into a show related to a grid, preferably a two-dimensional (2D) grid. This can be called the provision of a grid show. This means: a grid, preferably a 2D grid, is defined on the image plane (or the projection plane onto which the sight rays are projected), and each sight ray is assigned to a grid cell in this grid. Therefore, the image points can be such points on the grid where the sight rays enter the image plane or the projection plane.
[0016] In addition, optionally, it is set that different numbers of image points are assigned to the grid cells, where the provision of the input includes the following steps for further preprocessing:
[0017] - Perform normalization based on the grid show, preferably by normalizing the distance between the cell center point of the corresponding grid cell and the image points respectively assigned to this grid cell in the grid show.
[0018] Here, the grid show can be transformed into a convolutional show and / or normalized to a 2D grid composed of vectors of a fixed length. This can have the following advantages: the grid show can be processed by machine learning algorithms, especially standard CNN schemes.
[0019] In addition, the following advantages can be achieved: By means of the method according to the invention, image data can be transformed into a distortion-free display, enabling the machine learning model to be used as a general algorithm, for example for recognizing objects in image data or for semantic segmentation of image data. This enables the method according to the invention to be used, for example, in machines such as vehicles or robots in order to control the machine based on object recognition.
[0020] In addition, within the scope of the invention, it can be provided that the feature map is calculated by normalization, where the feature map includes, for each grid cell, a feature vector that is calculated from the image points of the corresponding grid cell. This enables a series of tasks, such as semantic segmentation or object recognition, to be performed based on the feature map with the aid of a machine learning model. For example, it is also possible to map labels onto the grid display for this purpose in order to enable the machine learning model to be trained.
[0021] In addition, the normalization can also be based on applying a neural network to the image points of the corresponding grid cell. For example, a feature map in a 2D grid plane can be obtained by applying a neural network in the form of a CNN (Convolutional Neural Network or convolutional neural network). Thereby, the distortion of the imaging device can be directly modeled. Therefore, it can also be optionally provided that the neural network includes at least one convolutional layer.
[0022] In addition, it can be envisaged that the feature map includes, for each grid cell, a unique feature vector, preferably a feature vector having at least one channel. If applied to all grid cells, the result of the process can be a feature map having the same size as the above-mentioned 2D grid.
[0023] Preferably, within the scope of the invention, it is provided that a machine learning model is applied by means of the provided input, where, preferably, the vehicle is controlled based on the application of the machine learning model, and the machine learning model is preferably trained with Dropout, preferably a specific form of Dropout. Vehicle functions such as driver assistance functions and / or autonomous driving functions can be provided, and these vehicle functions analyze and evaluate the output of the applied machine learning model in order to control the vehicle based on this. In addition, Dropout can be used in the first layer during the training of the machine learning model in order to make the machine learning model robust to changes in the optical parameters.
[0024] A machine learning model, hereinafter also simply referred to as the model, can be pre-trained with image data from a camera device of a certain type. Thanks to the advantageous conversion into a line-of-sight ray representation, the machine learning model trained in this way is still suitable for use with other types of camera devices. Here, the type of camera device and / or imaging errors, such as camera distortion, during the application of the machine learning model can be different from the camera device used for training the machine learning model. Therefore, it is preferably possible to train the model in such a way that the model remains invariant to differences in the optical characteristics between the camera devices. Therefore, it is preferably possible to apply a model, such as a CNN, to camera device distortions that were not present during training by means of the method according to the invention.
[0025] Furthermore, directly modeling camera device distortion also offers several advantages. First, by explicitly modeling prior knowledge about distortion, the learning problem can be made less difficult. The neural network does not have to implicitly learn the transformations between different distortions because these transformations are explicitly modeled in the CNN structure. This frees up capacity in the model that can be used to learn task-relevant features for the corresponding task. In addition, explicitly training for invariance with respect to camera device distortion enables the use of different machine learning models, such as machine learning models trained for completely different camera device distortions in the form of deep learning models, especially when the sensors of two camera devices are compatible. Therefore, a machine learning model can also be used, for example, with a fisheye camera device even if the model was not trained for such a camera device distortion. This also means that for the case of applying a trained machine learning model to a dataset with other camera device distortions (e.g., images with a skewed viewing angle), large-scale labeling and retraining are not required.
[0026] Likewise, the subject matter of the present invention is a computer program, in particular a computer program product, which includes instructions that, when the computer program is executed by a computer, cause the computer to execute the method according to the invention. Therefore, the computer program according to the invention offers the same advantages as those detailed with respect to the method according to the invention.
[0027] Likewise, the subject matter of the present invention is a device for data processing that is configured to execute the method according to the invention. For example, a computer executing the computer program according to the invention can be configured as the device. The computer can have at least one processor for executing the computer program. A non-volatile data memory can also be provided in which the computer program can be stored and read out by the processor for execution.
[0028] Likewise, the subject matter of the present invention is a computer-readable storage medium comprising a computer program according to the present invention. The storage medium is configured, for example, as a data memory such as a hard disk and / or a non-volatile memory and / or a memory card. The storage medium can be integrated, for example, into the computer.
[0029] In addition, the method according to the present invention can also be implemented as a computer-implemented method.
[0030] Further advantages, features and details of the present invention can be derived from the following description, in which embodiments of the present invention are described in detail with reference to the accompanying drawings. Here, the features mentioned in the claims and in the description can be of essential significance for the invention individually or in any combination. Description of the Drawings
[0031] The drawings show:
[0032] Figure 1 : Exemplary visualization of different optical parameters of a camera device.
[0033] Figure 2 : Definition of a 2D grid on an image plane for transforming a line of sight ray into a relative display.
[0034] Figure 3 : Application of a CNN to a relative line of sight ray display.
[0035] Figure 4 : According to Figure 3 An exemplary implementation of an embodiment, where the CNN is a Transformer network.
[0036] Figure 5 : Visualization of an embodiment of the present invention in the form of a box.
[0037] Figure 6 : Schematic visualization of a method, device and computer program according to an embodiment of the present invention.
[0038] In the following drawings, the same reference numerals are used for the same technical features of different embodiments. Detailed Description of the Embodiments
[0039] Since a camera device typically uses a method of arranging individual sensors in a raster form to record visual information, the output of such a camera device is a matrix consisting of pixel values. Such pixels are in Figure 1Exemplarily shown in the following. First, it is often meaningful to transform this pixel matrix into other representations as follows. Using the parameters of the imaging device itself, coordinates can be calculated for each pixel: the coordinates at which light impinges on the image plane (see, for example, Z. Zhang, Camera Parameters (Intrinsic, Extrinsic), pp. 81–85. Boston, MA: Springer US, 2014). Essentially, this representation describes the points at which the light rays that give rise to the pixel impinge on a hypothetical plane in the optical path. Therefore, this representation is also referred to as the “line-of-sight ray representation”. It should be noted that this representation can always be related to the plane that is commonly referred to as the “image plane”. Of course, any virtual plane can be constructed and the line-of-sight ray representation of the image with respect to this plane can be calculated. For training a model that is independent of the imaging device, it may be meaningful to select a virtual plane that can be easily generalized between different imaging devices, such as a sphere surrounding the imaging device.
[0040] In Figure 1 the different optical parameters of the imaging device are illustrated. In the first image ( Figure 1 a), a barrel distortion is applied, which results in more pixels being mapped at the edges of the image than in the middle. In contrast, in the second image ( Figure 1 b), such a distortion is applied, in which case more pixels are mapped in the middle region and fewer pixels are mapped at the edges.
[0041] Computer vision algorithms using deep learning have been extremely successful in many applications. These algorithms are based on a specific representation of the information detected by the imaging device: a series of pixels arranged in a matrix structure with specific adjacency relationships: an image is typically an H×W×C matrix with multiple values, where H and W are the height and width of the image, and C is the number of channels, for example 3 for an RGB image. However, this representation is highly related to the imaging device, because the optical parameters of the imaging device determine the pixel representation, such as the pixel representation of an object in the real world (see Figure 1)。Therefore, traditional machine learning models trained on such images can only learn the camera-device-specific representation and will have problems when the representation changes due to the use of other camera devices—even if the object being captured is the same (see, e.g., X. Peng, Y. L. Murphey, S. Stent, Y. Li, and Z. Zhao, “Spatial focal loss for pedestrian detection in fisheye imagery,” 2019 IEEE Winter Conference on Applications of Computer Vision (WACV), pp. 561–569, 2019).
[0042] Embodiments of the present invention can overcome these drawbacks by the following steps:
[0043] First, the transformation from an image presentation to a line-of-sight ray presentation can be carried out with the aid of known methods. Second, it can be set that points are subsequently converted into relative coordinates along a 2D grid. Third, a fully convolutional network can be applied to the presentation for each grid cell, and fourth, a regular CNN can subsequently be applied to the resulting feature map in order to solve a specific task (e.g., semantic segmentation). Additionally, fifth, Dropout can be used in the first layer during training in order to make the model robust to changes in the optical parameters. The steps of the method are described in more detail subsequently.
[0044] First, the image presentation of the image, preferably the pixel matrix presentation of the image, can be transformed into a line-of-sight ray presentation. This can be achieved with traditional methods from the prior art (see, e.g., Z. Zhang, Camera Parameters (Intrinsic, Extrinsic), pp. 81–85. Boston, MA: Springer US, 2014. In: Ikeuchi, K. (eds) Computer Vision. Springer, Boston, MA). Essentially, the presentation consists of a set of points pi ∈ P, where pi = (x i , y i , rgb i ), and x i , y i are the coordinates in the image plane, and rgb iis the color information of the pixel (note that the same approach can be used for each color space, not only RGB, and can be simply extended to other modalities such as RGBD).
[0045] Then, in the next step, the view ray representation can be transformed into a representation related to a two-dimensional (2D) grid. This means that a 2D grid is defined on the image plane (or the plane onto which the view rays are projected), and each view ray is assigned to a grid cell in this grid. This serves as a preprocessing step and only needs to be performed once, as Figure 2 shown. In this regard, the grid representation is also mentioned. Thus, the relative view ray representation consists of points l i ∈L, where l i =(Δx i , Δy i , rgb i ), where Δx i , Δy i are the coordinates relative to the nearest grid point. Essentially, through these steps, the view ray representation can be mapped onto a 2D grid. However, this representation may not be directly usable by a machine learning model because there are different numbers of mapped view rays in each grid cell.
[0046] In Figure 2 the definition of the 2D grid on the image plane is visualized for transforming the view rays into a relative representation. The grid density and the exact assignment method can be hyperparameters of the method according to an embodiment of the present invention.
[0047] In a further step, a transformation of the relative view ray representation, preferably the grid representation, into a convolutional representation can be performed. In other words, the relative view ray representation can be normalized onto a 2D grid consisting of vectors of a fixed length. This can have the advantage that the relative view ray representation can be processed with a standard CNN scheme - as Figure 3 shown. For this purpose, an artificial neural network such as a CNN can be applied to one or more points l iabove in order to compute some features for that grid cell. Here, the artificial neural network can for example involve one or more one-dimensional convolutional layers and a global average pooling. What may be crucial is that the resulting feature map must have a unique (with one or more channels) feature vector for each grid cell. If applied to all grid cells, the result of this process can be a feature map having the same size as the above-mentioned 2D grid. The CNN can be defined in various ways, for example as a series consisting of convolutions and poolings with variable input size, as in Figure 3 or defined as a Transformer network, as used in “A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, An image is worth 16x16 words: Transformers for image recognition at scale, ‘ArXiv, vol. abs / 2010.11929, 2020” and described in Figure 4
[0048] In Figure 3 the application of the CNN on the relative line-of-sight ray display is visualized. The CNN preferably works with a variable input size N_i×D, where N i is the number of line-of-sight rays assigned to grid cell i, and D is the input dimension (5 for 2D RGB). The output of this block can then be produced in a fully convolutional manner and has the form 1×F, where F is the selected feature dimension (this is a hyperparameter of the model, as in all other convolutional layers). In Figure 4 an alternative implementation according to Figure 3 is shown, where the CNN is a Transformer network.
[0049] Finally, a standard CNN can be applied to the feature map output in the previous step in order to perform a series of tasks such as semantic segmentation or object recognition. Here, it is also preferably possible to map the labels to the above-mentioned grid display in order to enable training.
[0050] Special advantages can be achieved in the following way: The resulting model is trained such that it is robust to changes in the line-of-sight ray representation. This can be achieved in particular by dynamically changing the assignment from P to L. For example, a form of Dropout can be incorporated, in which some l i points are randomly masked from each grid cell, which effectively results in less information being available for the grid cell. Since the CNN applied to these L i points is fully convolutional and is thus jointly utilized by all grid cells, this results in an information processing step that is independent of distortion and that acts on images of imaging devices with different optical properties.
[0051] In Figure 5 the embodiments of the invention are visualized in the form of a box: A (relative) line-of-sight ray representation (A) is established, a CNN is applied to obtain a feature map (B) on a 2D grid plane, and a standard CNN is applied to the resulting features (C).
[0052] In Figure 6 a method, a device, and a computer program according to an embodiment of the invention are schematically shown. Here, the method 100 for processing image data 205 to apply a machine learning model 200 can include, in a first method step 101, acquiring 101 the image data 205, where the image data 205 is obtained from an image capture using an imaging device 40. Then, a conversion of the image data 205 to a line-of-sight ray representation can be set according to a second method step 102. Subsequently, according to a third method step 103, an input for the machine learning model 200 can be provided based on the image data 205 in the line-of-sight ray representation. The providing can include, for example, a data transfer in which the image data - if necessary after further preprocessing - is delivered as an input to the machine learning model.
[0053] In Figure 2 it is shown that the image data 205 in the line-of-sight ray representation can be represented by image points 215. In addition, a grid representation 210 can be provided by further preprocessing, in which the grid has a plurality of grid cells, and the image points 215 are assigned to these grid cells. Here, the image points 215 can be assigned to the grid cells in different numbers. Subsequently, a normalization can be set, in which the distance between the corresponding cell center point and the corresponding image point 215 in the grid representation is normalized, for example by a neural network 230, whereupon a feature map 220 is obtained (see Figure 3 ).
[0054] The feature map 220 can advantageously include a unique feature vector for each grid cell, preferably a feature vector having at least one channel. As shown in Figure 4 , a Transformer 230 can also be used as a neural network 230 that pre-generates "Embeddings" 235 in order to generate the feature map 220.
[0055] The above explanations of the embodiments only describe the present invention within the scope of examples. Of course, individual features of the embodiments, as long as they are technically meaningful, can be freely combined with each other without departing from the scope of the present invention.
Claims
1. A method (100) for processing image data (205) to apply a machine learning model (200), comprising: - obtaining (101) the image data (205), wherein the image data (205) is obtained from image detection using an imaging device (40), - converting (102) the image data (205) into a line-of-sight ray display, - providing (103) an input for the machine learning model (200) based on the image data (205) in the line-of-sight ray display.
2. The method according to claim 1, wherein, the image data (205) in the line-of-sight ray display is represented by image points (215), and wherein the providing (103) of the input includes the following steps for further preprocessing: - providing a grid display (210), in which a grid has a plurality of grid cells, and wherein the image points (215) are assigned to the grid cells.
3. The method according to claim 2, wherein, the image points (215) are assigned to the grid cells in different numbers, and wherein the providing (103) of the input includes the following steps for further preprocessing: - performing normalization based on the grid display (210), preferably by normalizing the distance between the cell center point of the corresponding grid cell and the image points (215) respectively assigned to the grid cell in the grid display.
4. The method according to claim 3, wherein, a feature map (220) is calculated by the normalization, and wherein the feature map (220) includes feature vectors for the grid cells respectively, and the feature vectors are calculated from the image points (215) of the corresponding grid cells.
5. The method according to claim 3 or 4, wherein, the normalization is based on applying a neural network (230) to the image points (215) of the corresponding grid cells.
6. The method according to claim 5, wherein, the neural network (230) includes at least one convolutional layer.
7. The method according to any one of the above claims, wherein, the feature map (220) includes a unique feature vector for each grid cell, preferably a feature vector having at least one channel.
8. The method according to any one of the above claims, wherein, the machine learning model (200) is applied through the provided input, and wherein a vehicle (1) is controlled based on applying the machine learning model (200), and wherein the machine learning model is preferably trained with Dropout.
9. A computer program (20), which includes instructions that, when the computer program (20) is implemented by a computer (10), cause the computer to implement the method (100) according to any one of the above claims.
10. A device (10) for data processing, which is arranged to implement the method (100) according to any one of claims 1 to 8.