Training machine learnable model to estimate relative object scale
Patent Information
- Application Number
- JP2022080909
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-05-18
- Filing Date
- 2022-05-17
- Publication Date
- 2025-05-20
AI Technical Summary
Existing machine-learnable models for estimating object scale in images require extensive supervised training with labeled data, leading to high computational complexity and effort, and there is a need for more efficient and less complex methods to obtain relative scale estimates.
A scale estimator is integrated with a feature extractor to provide relative scale estimates, trained using unsupervised approaches by generating training targets through spatial scaling of image patches, allowing the model to learn relative scales without explicit ground truth, using a shallow neural network for aggregation and estimation.
The method allows for efficient training of scale estimators with reduced computational complexity, enabling accurate estimation of relative object scales for scene geometry analysis, which can be applied in autonomous vehicles and robotic systems.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a system and a computer-implemented method for training a machine learning-enabled model for estimating the relative scale of objects in an image. The present invention further relates to a system and a computer-implemented method for estimating the relative scale of objects in an image, for example, to determine the scene geometry (geometric arrangement of the scene) of the image. Furthermore, the present invention relates to a computer-readable medium containing data representing instructions for a processor system for carrying out any computer-implemented method. [Background technology]
[0002] Background technology The field of computer vision relates to enabling machines to "see" and understand their environment. The core task of computer vision is to identify and classify objects in images, such as pedestrians or vehicles in camera images acquired by autonomous vehicles, or parts being processed by manufacturing robots.
[0003] Scale is an inherent attribute of objects shown in an image, a fundamental property like location or appearance. Here, the term "scale" may refer to the apparent size of an object in an image, which can depend on the distance from the camera to the object, the camera's focus, etc. Computer vision tasks typically need to consider the (varying) scale of objects in an image. For example, image classification preferably requires scale invariance to achieve accurate classification results. In image segmentation, scale equivariance is important because the output map needs to scale proportionally to the input. In object detection or object tracking, scale invariance and scale equivariance are important.
[0004] Scale invariance or equivariance is typically addressed in computer vision by providing a sufficient variety of examples in the training data, for example, objects at different scales. It is also possible to provide scale invariance or equivariance by fitting machine learning models and / or their training. For example, the publication "Scale-Equivariant Steerable Networks" 2019, https: / / arxiv.org / abs / 1910.11093v1, describes incorporating a mechanism for scale equivariance within a CNN to improve its performance, where performance can be understood as the CNN's ability to properly classify images. This scale equivariance mechanism is based on constructing the filters in the convolutional layer of the neural network to be a weighted sum of basis filters (also called basis functions), where the weights can be trained during the CNN's training.
[0005] While it is known that machine learning models can be fitted and / or trained for computer vision tasks that require scale invariance or equivariance, or at least to a sufficient degree, it is sometimes desirable to obtain an explicit representation of the relative scale of objects in an image. This may, for example, allow for the estimation of the scene geometry depicted in the image. For instance, if a scene is densely packed with similar objects, such as flowers in a flower field, the relative scale of the objects in the image of this scene may indicate a scene geometry in which the surface of the objects is a plane containing objects whose surfaces are sloped toward the horizon. This may be evident from the fact that objects placed near the bottom of the image have a larger apparent size and, consequently, a decreasing apparent size toward the center of the image. In other words, the relative scale of objects allows conclusions to be drawn based on the scene geometry, which can be useful in many real-world applications, such as autonomous driving where camera images indicate the possibility of traffic congestion by a dense field of vehicles. Another example is an environment with many pedestrians, where the scene geometry may be used to identify pedestrians close to the autonomous vehicle and filter them based on importance, allowing the autonomous vehicle to ignore pedestrians that are not important because they are far away. In general, gaining an understanding between objects at different scales can make it possible to identify relationships between objects, such that objects of the same or similar scales may be related. This can be used to generate relationship graphs of objects. [Prior art documents] [Non-patent literature]
[0006] [Non-Patent Document 1] “Scale-Equivariant Steerable Networks,” 2019, https: / / arxiv.org / abs / 1910.11093v1 [Overview of the project] [Problems that the invention aims to solve]
[0007] It is known that using supervised learning, a machine learning model can be trained to provide an explicit representation of the scale of objects in an image, for example, by providing training data annotated with the objects' scales, either pixel-wise or relative to the image resolution. However, this has the drawback of requiring a large amount of training data showing objects at various scales, and a great deal of manual work to carefully construct a machine learning model that can explicitly indicate scale, for example, by providing labels.
[0008] For example, the ability to obtain scale estimators that are easier to train and have limited computational complexity using unsupervised methods would be an advantage. [Means for solving the problem]
[0009] Summary of the Invention According to a first aspect of the present invention, a computer-implemented method and a corresponding system are provided for training a machine learning-enabled model for estimating the relative scale of objects in an image, as defined by claims 1 and 14, respectively. According to a further aspect of the present invention, a computer-implemented method and a corresponding system for estimating the relative scale of objects in an image are provided, as defined by claims 7 and 15, respectively. According to a further aspect of the present invention, a computer-readable medium is provided that includes instructions for causing a processor system to perform any of the computer-implemented methods, as defined by claim 13.
[0010] A scale estimator is provided which may be trained to provide relative scale estimates, which may include machine learning-enabled model parts such as neural networks, according to the means described above. This scale estimator may be provided as an “add-on” to a feature extractor. Such a feature extractor may be an existing “known” feature extractor configured to take, for example, a 64x64 pixel image patch or an image patch having any other suitable spatial dimension as input and extract multiple features from the image patch that relate to one or more objects in the image. Examples of these features include different types of edges, textures, corners, etc. These features may be defined manually, or they may be machine-learned, for example, by a feature extractor previously trained on training data containing examples of objects. The feature extractor may provide multiple feature maps as outputs, as is known by itself. For example, if the feature extractor includes a convolutional neural network (CNN), such feature maps may consist of or be represented by the output channels of the CNN. Furthermore, the feature extractor may be fitted, and in the case of a machine learning-based feature extractor, it may be trained to be at least appropriately scale-equivalent. This manifests as a feature map having a scale dimension. Thus, the feature extractor can provide filtered responses along each of the scale dimensions. Such feature extractors themselves are known, for example, as far as scale-equivalent CNNs (SE-CNNs) described herein, from concurrently pending European Patent Application No. 20195059, incorporated herein by reference, whose input and convolutional layers can constitute an example of a feature extractor like the one described herein.
[0011] The scale estimator may be configured to aggregate each feature map obtained from the output of the feature extractor. For example, such aggregation may include aggregating filter responses along various dimensions of the feature map, such as its spatial dimension. In particular, the maximum filter response may be identified along the scale dimension, and as a result, the aggregation of feature maps may provide a feature-level scale estimate. In a particular example, if the CNN has 512 output channels, the aggregation may result in 512 feature-level scale estimates, each derived from the maximum filter response along the respective scale dimension of the feature map. The scale estimator may further include a machine learning-capable model portion, such as the aforementioned neural network, which may be configured to receive feature-level scale estimates as input and output patch-level scale estimates representing the overall scale estimate of the image patch provided as input.
[0012] For this purpose, the machine learning-capable model portion may be trained on training data. However, instead of relying on supervised training where patch-level scale estimates are provided manually or at least externally as ground truth, a suitable target for training may be generated during training. In detail, for example, it may be sufficient for a scale estimator to learn the relative scale of objects in order to learn that one object is closer to the camera than another. Such a relative scale may not represent an absolute measure of scale, and thus it may not be possible to draw conclusions about the absolute size of an object, such as, for example, that an object is 2m wide. Nevertheless, such a relative scale may be sufficient for various purposes, including the estimation of scene geometry as described above. The training target for estimating relative scale may be generated by the means described in the claim by spatial scaling of image data of image patches of the training image according to at least two known scale factors. For example, the image data in an image patch may be reduced by a factor of, for example, 0.75, or enlarged by a factor of, for example, 1.5. It will be understood that such enlargement may include cropping and reduction may include padding so that image patches of equal dimensions are obtained. In other examples, the image data of an image patch may be scaled, for example, using a unit scale factor, with a factor of 1.0 and a factor of 1.5. Various other examples of such scale factors are conceivable. Such scale factors may be referred to as “known” or “actual” scale factors.
[0013] Next, the feature extractor and scale estimator may be applied to image patches containing scaled image data, resulting in at least two patch-level scale estimates. While the absolute sizes of objects in either image patch may be unknown, the relative sizes of objects between the two image patches may be known, represented by a relationship between two known scale coefficients. This relationship may be expressed, for example, as a difference (e.g., 2.0 - 0.5 = 1.5) or a ratio (e.g., 2.0 / 0.5 = 4.0), and may be referred to as the actual relative scale. Here, "actual" refers to the fact that the image data has actually been scaled according to each scale coefficient, and "relative" refers to a numerical value representing the relationship between the scale coefficients (e.g., difference or ratio). A similar relationship may be calculated for the patch-level scale estimates, resulting in the estimation of relative scales. The loss function may be formulated to fit the parameters of machine learning-enabled model parts, such as the weights of a neural network, so that training learns to estimate the actual relative scale. In detail, the loss function can represent the mismatch between the actual relative scale and the estimated relative scale, and training can seek to minimize this mismatch. Thus, the scale estimator can learn to estimate the relative scale better. This does not require manually provided ground truth, because the actual relative scale may be generated internally. Thus, training the scale estimator may be easier. Furthermore, the scale estimator may be simply provided as an add-on to an (existing) feature extractor. This contributes to separation of concerns, as there is no need to be overly attached to the feature extractor itself, but rather the scale estimator itself is trained to adapt to the feature map provided by a particular feature extractor.Furthermore, such scale estimators have proven to be relatively simple architecturally, as they "simply" require aggregating feature maps into feature-level scale estimates and combining these feature-level scale estimates into patch-level scale estimates. Such combinations may be performed by a relatively simple machine learning-capable model portion compared to the feature extractor itself. For example, in many applications, a shallow neural network with only one hidden layer may suffice. Therefore, if a feature extractor is already available for purposes such as object detection or classification, a scale estimator may be added at a relatively low cost in terms of computational complexity and / or training effort.
[0014] Optionally, each feature map includes at least two spatial dimensions and a scale dimension, where the scale estimator is configured to aggregate each feature map across at least two spatial dimensions by averaging, weighting, or majority voting. The spatial dimensions of the feature maps may not be particularly relevant to scale estimation. Therefore, the spatial dimensions may be reduced to, for example, 1×1 by aggregating the feature maps across the spatial dimensions by averaging, weighting, majority voting, or similar techniques. In a particular example, a global mean pooling layer may be used to reduce an H×W×S feature map to a 1×1×S feature map (where "S" represents the scale dimension).
[0015] Optionally, the scale estimator is configured to identify the spatial scale at which the filter response is maximized and to aggregate each feature map across the scale dimension by using the spatial scale identifier as a feature-level scale estimate or part thereof. The feature map may be aggregated from, for example, 1×1×S to 1×1×1 by identifying the spatial scale at which the filter response is maximized and using the spatial scale identifier as a feature-level scale estimator or part thereof. For example, a predetermined set of scales.
Number
[0016] Optionally, the machine-learnable model part of the scale estimator includes a neural network. For example, this neural network may be a shallow neural network having one hidden layer. Such a neural network or generally a shallow multi-layer perceptron (MLP) has been found to be sufficient to learn to combine the feature-level scale estimate into the patch-level scale estimate. Such an MLP may require fewer resources for implementation and may be easily trained considering its relatively few parameters.
[0017] Optionally, the error term defines the mean squared error or mean squared deviation between the actual relative scale and the estimated relative scale. These mean squared error (MSE) or mean squared deviation (MSD) are both suitable as error functions, but require few resources for evaluation during implementation and execution.
[0018] Optionally, a scene geometry map indicating the scene geometry of an image is generated by the following steps, namely, - Applying a feature extractor and a scale estimator to a plurality of image patches of the image to obtain a plurality of patch-level scale estimates; - Generating a scene geometry map for the image as a representation of the plurality of patch-level scale estimates associated with the plurality of image patches. and is generated by.
[0019] As described elsewhere, the patch-level scale estimates may be combined into the scene geometry map, for example, by constructing an array representing the image patches using each position within an array containing the respective patch-level scale estimates. Such an array can be made similar to a map for an image and may represent a scene geometry as described elsewhere in this specification.
[0020] Optionally, the feature extractor and the scale estimator are applied to overlay image patches of an image. Since the scale estimator can generate one patch-level scale estimate per image patch, if the scale estimator is applied to non-overlapping image patches, the resulting scene geometry map may be relatively low resolution compared to the input image. To increase the resolution of the scene geometry map, the feature extractor and the scale estimator may be applied to overlapping image patches. This can provide a more detailed and accurate scene geometry map for the image. [[ID=⑤]] [[ID=⑥]]
[0021] [[ID=⑦]] [[ID=⑧]]Optionally, the scene geometry map is generated by the following steps, namely, [[ID=⑨]] [[ID=⑩]]- subtracting the minimum value of a plurality of patch-level scale estimates from the plurality of patch-level scale estimates in the scene geometry map, and / or [[ID=⑪]] [[ID=⑫]]- spatially expanding the scene geometry map to the spatial resolution of the image [[ID=⑬]] [[ID=⑭]]and is generated thereby. [[ID=⑮]] [[ID=⑯]]
[0022] [[ID=⑰]] Optionally, the image may be acquired from a sensor configured to sense the environment of a computer-controlled entity, and the scene geometry map of the image is analyzed. Based on the analysis results, control data is generated for the computer-controlled entity to adapt and control the computer-controlled entity to the environment. The computer-controlled entity, such as a robotic system or an autonomous vehicle, may be controlled based on the analysis results of the scene geometry map. For example, the image from which the scene geometry map is generated may be acquired by an on-board camera, and this scene geometry map is used to show the scene geometry acquired by the on-board camera. For example, this scene geometry map may indicate that there is traffic congestion ahead of an autonomous vehicle, in which case it may be desirable to control the vehicle differently, for example, by slowing down.
[0023] Those skilled in the art will understand that two or more of the above-described embodiments, implementations, and / or any other embodiments of the present invention may be combined in any way deemed useful.
[0024] Any modification and variation of any system, any computer-implemented method, or any computer-readable medium, corresponding to another described modification and variation of the entity, can be carried out in accordance with this description by a person skilled in the art.
[0025] These and other embodiments of the present invention will become apparent from the embodiments described by example in the following description and from the accompanying drawings, and will be further clarified by reference to these embodiments and accompanying drawings. [Brief explanation of the drawing]
[0026] [Figure 1] This diagram illustrates a system for training a machine learning-enabled model that estimates the relative scale of objects within an image. [Figure 2]This figure illustrates the steps of a computer-implemented method for training a machine learning-enabled model to estimate the relative scale of objects within an image. [Figure 3] This figure shows a feature extractor and scale estimator applied to image data of an image patch to generate a patch-level scale estimate for the image patch. [Figure 4] This figure illustrates the calculation of the error term for training a scale estimator, which involves scaling image data in an image patch according to two different scale factors by estimating the scale factors using the scale estimator, where the error function represents the inconsistency between the relationships of the respective scale factors. [Figure 5A] This figure shows the input images for the feature extractor and scale estimator. [Figure 5B] This figure shows an input image divided into multiple non-superimposed image patches, where each image patch is used as input to a feature extractor and a scale estimator. [Figure 5C] This figure shows a scene geometry map generated by a scale estimator, which includes patch-level scale estimates for each image patch, and is spatially scaled to the image resolution. [Figure 6] This diagram illustrates a system for estimating the relative scale of objects within an image. [Figure 7] This figure shows a (semi)autonomous vehicle that includes the system described in Figure 5, which generates and analyzes a scene geometry map and controls the vehicle based on it. [Figure 8] This diagram illustrates the steps of a computer-implemented method for estimating the relative scale of objects within an image. [Figure 9] This is a diagram showing a computer-readable medium containing data.
[0027] It should be noted that these drawings are purely illustrative and not drawn to scale. In these drawings, elements corresponding to elements already described may have the same reference numeral.
[0028] The following list of reference numerals is provided to facilitate the interpretation of the drawings and should not be construed as limiting the scope of the claims.
[0029] Explanation of the symbols A system for training a machine learning-enabled model to estimate the relative scale of objects in 100 images. 120 Processor Subsystems 140 Data Storage Interfaces 150 Data Storage 152 training data 154 Feature Extractor Data Representation 156 Data representation of scale estimators A method for training a machine learning-enabled model to estimate the relative scale of objects in 200 images. 210 Provides a feature extractor. Provides a 220-scale estimator. Access 230 training data. 240 training sessions 245 Repeat for the next image patch. 250 Spatially scale the image patch to obtain scaled image patches. 260 Apply feature extractors and scale estimators. Optimize the machine learning-enabled model portion of the 270-scale estimator. 300 image patch 310 Image patch with reduced image data 320 Image patches with enlarged image data 360 Feature Extractor 380 Scale Estimator 400 images showing a flower field 410 Images divided into multiple image patches 412 Image Patch 420 Scene Geometry Maps 430 Enlarged Scene geometry map enlarged to 440 image resolution A system for estimating the relative scale of objects in 500 images. 520 Processor Subsystems 540 Data Storage Interfaces 550 Data Storage 552 Image Data 554 Feature Extractor Data Representation 556 Data representation of scale estimators 560 Sensor Data Interface 562 Sensor Data 570 Control Interface 572 Control Data 600 Environment 610 (Semi-)Autonomous Vehicle 620 sensors 622 Camera 630 Actuator 632 Electric motor How to estimate the relative scale of objects in 700 images 710 Provides a feature extractor. Provides a 720-scale estimator. 730 Apply the feature extractor and scale estimator to the image patch. 740 Repeat for the next image patch. Outputs a data representation of the 750 patch level scale estimate. 800 Computer-readable media 810 Non-temporary data [Modes for carrying out the invention]
[0030] Detailed description of the embodiment The following describes a system and a computer-implemented method for training a machine learning-readable model for estimating the relative scale of objects in an image, with reference to Figures 1 and 2; a feature extractor and a scale estimator being applied to image data of an image patch to generate patch-level scale estimates for the image patch, with reference to Figure 3; training the machine learning-readable model portion of the scale estimator, with reference to Figure 4; the scale estimator being applied to an input image to generate a scene geometry map, with reference to Figures 5A to 5C; a system and a computer-implemented method for estimating the relative scale of objects in an image, with reference to Figures 6 and 8; and an autonomous vehicle incorporating the system described in Figure 6, with reference to Figure 7. Figure 9 shows a computer-readable medium used in embodiments of the invention described in the claims.
[0031] Figure 1 shows a system 100 for training a machine learning-enabled model that estimates the relative scale of objects in an image. This system 100 may include an input interface subsystem for accessing training data 152 for training. For example, as shown in Figure 1, the input interface subsystem may include or be comprised of a data storage interface 140 that can provide access to the training data 152 on data storage 150. For example, the data storage interface 140 may be a memory interface or a persistent storage interface such as a hard disk or SSD interface, or it may be a personal, local or wide area network interface such as Bluetooth, Zigbee, Wi-Fi interface, Ethernet or fiber optic interface. The data storage 150 may be internal data storage of the system 100, such as memory, a hard drive or SSD, or it may be external data storage, such as network-accessible data storage.
[0032] In some embodiments, the data storage 150 may further include a data representation 154 for the feature extractor and a data representation 156 for the scale estimator, both of which are discussed in detail below and may be accessed by the system 100 from the data storage 150. However, it will be understood that the training data 152, the data representation 154 for the feature extractor, and the data representation 156 for the scale estimator may each be accessed from different data storages, for example, through different data storage interfaces. Each data storage interface may be of the type described above for the data storage interface 140. In other embodiments, the data representations 154, 156 for the feature extractor and / or scale estimator may be generated internally by the system 100, for example, based on design parameters or design specifications, and therefore do not necessarily have to be explicitly stored on the data storage 150.
[0033] System 100 may further include a processor subsystem 120, which may be configured to train a scale estimator 156, in particular a machine learning-enabled model portion of the scale estimator 156, on training data 152 in a manner described elsewhere herein during the operation of System 100. For example, training by the processor subsystem 120 may include the step of running an algorithm that optimizes the parameters of the scale estimator 156 using a training goal, such as a loss function. In some embodiments, a feature extractor 154 may include or consist of a machine learning-enabled model, and the processor subsystem 120 may be configured to train the feature extractor 154 on training data 152 or on different or additional training data.
[0034] System 100 may further include an output interface for outputting a data representation of a trained scale estimator, which is also referred to as a machine-trained scale estimator, and the data is also referred to as trained scale estimator data. "Trained" will be understood, both herein and elsewhere, to mean that at least the machine-learnable model portion of the scale estimator has been trained. For example, as shown in Figure 1, the output interface may consist of a data storage interface 140, which in these embodiments is an input / output ("IO") interface through which the trained scale estimator is stored in a data storage device 150. For example, a data representation 156 defining an "untrained" scale estimator may be replaced, at least partially, by a trained scale estimator data representation during or after training, such that the parameters of the scale estimator 156, in particular the parameters of the machine-learnable model portion of the scale estimator 156, are adapted to reflect training based on the training data 152. In other embodiments, the data representation of the trained scale estimator may be stored separately from the data representation 156 of the "untrained" scale estimator. In some embodiments, the output interface may be separate from the data storage interface 140, but generally the data storage interface 140 may be of the type described above.
[0035] Figure 2 shows a computer-implemented method 200 for training a machine learning-enabled model, particularly a scale estimator, for estimating the relative scale of objects in an image. This method 200 may be adapted to the operation of system 100 in Figure 1, but it is not necessarily required to be adapted to the operation of other types of systems, apparatus, devices, or entities, or to steps in a computer program.
[0036] The Method 200 is shown to include a step 210 in which a feature extractor is provided as described elsewhere in this Spec., a step 220 in which a scale estimator is provided as described elsewhere in this Spec., a step 230 in which training data is accessed, including a set of training images, in a step titled "Accessing Training Data". The present method 200 is shown to further include a step 240 in which a machine learning-ready model portion of a scale estimator is trained on training data, in a step titled “Train,” which includes a step 250 in which the image data of the image patches of the training images is spatially scaled by at least two known scale coefficients to obtain at least two further image patches, in a step titled “Spatially scale image patches to obtain scaled image patches,” which includes a step 260 in which the feature extractor and scale estimator are applied to at least two further image patches to obtain at least two patch-level scale estimates, and a step 270 in which the parameters of the machine learning-ready model portion of the scale estimator are optimized by minimizing the error term of the loss function, in a step titled “Optimize machine learning-ready model portion of scale estimator.” This training step 240 may include multiple iterative loops that are repeated across different image patches of the training images and across different training images (not shown), for example, as indicated by arrow 245 in Figure 2.
[0037] Continuing with the estimation of the relative scale of objects in an image, the means described herein utilize a feature extractor and a scale estimator. The feature extractor may, but does not necessarily, be a machine learning-capable feature extractor, and may be trained separately from the scale estimator, for example, by different types of systems, and / or at different moments in time, based on different types of training data. For example, a system and method for training a scale estimator may use a pre-trained feature extractor that has been previously trained by a method known in itself. In an unrestricted example, the feature extractor may be a scale-equivalent convolutional neural network (SE-CNN) trained to extract features from image patches. Such feature extraction may result in the output of a feature map for each feature, which in the CNN example is also referred to as the CNN's "channels".
[0038] For example, consider a function F: x → y. Here, x and y are the input and output tensors, and F represents a feature extractor, such as an SE-CNN. The input tensor may have the shape batch_size × 3 × height × width, while the output tensor may have the shape batch_size × num_channels × num_scales × height' × width'. Here, "batch_size" can represent the number of image patches used as input, "3" can represent the three color components of the image data (e.g., RGB or YUV), "height" and "width" can be the height and width of each image patch (e.g., 64 × 64 pixels), "num_channels" can represent the number of feature maps generated as output, "num_scales" can represent the number of scales on which features were detected, which in turn can correspond to the scale dimension of the feature map, and "height'" and "width'" can represent the height and width of the feature map, thereby representing the spatial dimension of the feature map.
[0039] As is known, a feature extractor may be configured to detect multiple features in an image patch to obtain multiple feature maps as output, where the multiple features relate to one or more objects in the image, and each feature map is generated by applying a filter to the image data of the image patch, and each feature map includes a filter response across a set of different spatial scales along the scale dimension. Thus, the feature extractor may be configured to take scale information into account by providing a filter response across different scales. As is known, the feature extractor may be configured by the scales to be used, for example, with respect to the number of scales and scale coefficients. For example, the scale may be defined as a hyperparameter of the feature extractor. The number of scales and the step size may be selected depending on the specific application. For example, if the image is a camera image acquired by an on-board camera of a vehicle likely to be showing congestion, the relative size of other vehicles in the congestion is expected to vary slightly from vehicle to vehicle. Thus, a relatively small step size, such as 1.4, can be used between scales. Also, a very distant car may be up to 1 / 8 the size of a nearby car, hence 9 scales.
number
[0040] As an example of a feature extractor, an ImageNet-pretrained CNN, such as the SE-CNN described in European Patent Application No. 20195059, may be used. Furthermore, it can be assumed that the feature map represents the features of only one object. Thus, each feature map may be spatially aggregated using, for example, a global spatial mean pooling layer P. This layer may be provided as the last layer of the feature extractor, or as a separate layer following the feature extractor. In a particular example, the feature extractor may have 512 output channels. For example, after aggregation using the global spatial mean pooling layer, the output tensor may have a shape of batch_size × 512 × 9, where "9" refers to the number of scales. From each output, the scale may be extracted where the maximum filtered response is obtained. This may be done, for example, by maximizing pooling across the scale dimensions using the argmax operator. As a result, 512 predictions of the scale may be obtained for each image patch. These predictions are referred to elsewhere as feature-level scale estimates. Next, a shallow multilayer perceptron G may be used to regress these 512 feature-level scale estimates onto a single patch-level scale estimate, where G may represent an example of what is referred to elsewhere as the machine learning-enabled model portion of the scale estimator. The shallow multilayer perceptron may be, for example, a neural network with one hidden layer, a deep neural network, a linear regressor, or any differentiable model that can map vectors (feature-level scale estimates) to scalars (patch-level scale estimates). In this regard, it should be noted that while the scale estimator may include a shallow machine learning-enabled model portion, it is not a requirement, as the scale estimator may also include a deep machine learning-enabled model portion, for example, with multiple hidden layers.
[0041] Figure 3 illustrates an example of the above, where an image patch 300 of size H×V×3, where H×V is, for example, 64×64 pixels, is input to the feature extractor F360, resulting in the generation of multiple feature maps of size H'×V'×9, where "9" is the scale number in this example. A global spatial average pooling layer P may be used to spatially aggregate the feature maps into a 1×1×9 feature map, which may be followed by an argmax operator capable of generating a single 1×1×1 feature-level scale estimate for each image patch. The machine learning-capable model portion of the scale estimator 380, for example a shallow multilayer perceptron G, may then be used to combine all feature-level scale estimates, for example, number 512, into a single patch-level scale estimate.
[0042] An example of a scale estimator with one hidden layer can be described by PyTorch-like pseudocode, as shown in the following code extract. [Table 1]
[0043] Continuing to refer to Figure 3, note that the global spatial average pooling layer P is shown to be separated from the feature extractor F. In some embodiments, the global spatial average pooling layer P or a similar type of spatial aggregator function may be part of the same integrated network that also includes the feature extractor F. In other words, a network may be provided that includes both the feature extractor F and a global spatial average pooling layer P or a similar function that follows the feature extractor. Various other types of decomposition of the elements of the feature extractor and scale estimator are also possible.
[0044] In some examples, the feature extractor F may be part of an object detector or classifier. Such an object detector or classifier may include an additional network layer that processes the feature map of the feature extractor F to obtain object segmentation or classification. In such examples, the feature map may represent internal data of the object detector or classifier, which may be accessed by a scale estimator that estimates the scale. Thus, in such examples, the feature map may be used for both object detection or classification by the object detector or classifier and for scale estimation.
[0045] The scale estimator, in particular its machine learning-capable model portion, such as a multilayer perceptron G, may be trained using an appropriate dataset. For example, datasets of various classes of natural images, such as ImageNet or STL-10, may be used. Training may include defining the training objective, for example, by defining a loss function. Figure 4 illustrates the calculation of this loss function. Here, image patches 300 from the training data may be scaled according to two different scale factors, for example by interpolation, where L is the scale factor with γ1 and γ2 respectively. γ1 ,L γ2 It has been shown that such a scale factor may be identified as such. For example, such a scale factor may be randomly sampled from a range of, for example, 0.5 to 2.0. This results in two image patches 310,320 containing scaled image data, which are connected to a network N including a feature extractor and a scale estimator. θ It may be supplied to the network N. The scale estimator here has a machine learning-enabled model portion with parameter θ. θ As explained with reference to Figure 3, the scaling factor is calculated by generating a patch-level scale estimate for each of the scaled image patches 310,320.
number
number
number
[0046] This loss function is expressed by a term γ1-γ2 representing the difference between two known scale factors γ1 and γ2, and by an estimated patch level scale.
number
number
[0047] Continuing to refer to the training, note that any other suitable loss function may be used, for example, one that uses a different error term than the MSE, such as the non-squared error. Additionally, instead of using the difference of the scaling factors, a ratio or other type of expression of the relationship of the scaling factors may be used. For example, L scale The following equation
number
[0048] In certain cases, the relationship between the scale factors may be expressed as the logarithm of the difference between the scale factors or the logarithm of their ratios.
[0049] In some examples, the loss function may be defined by taking into account more than two image patches, for example, by using three or more scaling factors. In some examples, at least one of the scaling factors is, for example, 1.0, representing a unitary scaling factor.
[0050] A combination of a feature extractor and a scale estimator may be used to estimate the relative scale of objects in an image, provided that a machine learning-capable model portion of the scale estimator, particularly one with parameter θ, has been trained.
[0051] Figures 5A to 5C show the inferences by the scale estimator. Here, Figure 5A shows an example of the input image 400 to the feature extractor and the scale estimator, and in this specific example, a scene in the form of a flower field is shown. Figure 5B shows the input image 410 divided into non-overlapping image patches 412, where each image patch is used as an input to the feature extractor and the scale estimator. Figure 5C shows the scene geometry map 420 generated by the scale estimator, and this scene geometry map 420 may be generated as an image-like representation of the patch-level scale estimation values for each image patch. Here, different greyscales indicate different scale estimation values. It can be seen that in this scene geometry map, this scene includes objects near the camera at the bottom of the image, and further includes objects near the center and top of the image separated from the camera. As shown by the arrow 430, the scene geometry map 420 may be spatially enlarged to a higher resolution, such as the resolution of the input image 410, using, for example, bicubic interpolation, to yield an enlarged scene geometry map 440. In some examples, the minimum estimated patch-level scale estimation value γ min may be subtracted from each individual patch-level scale estimation value, for example, before or after generating the scene geometry maps 420, 440. This inference procedure may be described by the following PyTorch-like pseudo-code.
Table 3
[0052] Various uses for estimates of the relative scale of objects are conceivable, and the generation of scene geometry maps presented here is merely one example. Nevertheless, the ability to easily generate scene geometry maps by estimating patch-level scale estimates for image patches in an input image is likely to be an advantage in many applications. Such scene geometry maps are likely to be particularly accurate for images of scenes where identical or similar types of objects, such as cars, people, and flowers, appear in dense clusters, such as in images of traffic jams, crowded spaces, sports stadiums, concerts, and fields.
[0053] A specific example is footage from an in-car camera. When a vehicle encounters traffic congestion, the road itself and road markings (e.g., lane markings) may become invisible or only partially visible. The scene may also be very dense in that there is a high density of other vehicles in front of the vehicle. In this case, the road geometry, such as its curvature, may be estimated from a scene geometry map, which may be obtained by estimating the relative position and scale of the cars within the scene.
[0054] Another example is a traffic surveillance camera located somewhere above a wide road, which typically monitors pedestrians or vehicles crossing the road. Since the traffic surveillance camera can perform automatic object detection, it would be good to train it to actually detect people crossing the road. A scene geometry map may be used for health checks in an example where an open-roof double-decker bus is passing by, which might indicate that the person detected by the camera appears to be on the ground surface and therefore unlikely to be actually crossing the road. Thus, in such cases and similar cases, a scene geometry map may be used as an additional input to decision logic following image-based object detection.
[0055] Figure 6 shows a system 500 for estimating the relative scale of objects in an image using a feature extractor and a scale estimator, as described elsewhere. This system 500 may include an input interface subsystem for accessing the data representations of the feature extractor and the scale estimator. For example, as also shown in Figure 6, the input interface subsystem may include a data storage interface 540, through which the data representations 554 of the feature extractor and 556 of the scale estimator are accessed. Generally, the data storage interface 540 and the data storage 550 may be of the same type as those described with reference to Figure 1 for the data storage interface 140 and the data storage 150. Figure 6 further shows that the data storage 550 includes image data 552 of an image to which the feature extractor and the scale estimator can be applied to estimate the relative scale of objects in the image. For example, the image data 552 may be acquired by a camera's image sensor, or it may be sensor data acquired by another type of spatial sensor, such as a LiDAR or radar. This sensor data may be represented as an image. In some embodiments, such sensor data may be received directly from the sensor 620, for example, via a sensor data interface 560 or another type of interface, instead of from the accessed data storage 550. In such embodiments, the sensor data 562 may be received by the system 500 "live," for example, in real time or pseudo-real time.
[0056] System 500 may further include a processor subsystem 520, which may be configured to apply feature extractors and scale estimators to image data 552 and / or sensor data 562 as image data to generate at least one patch-level scale estimate, or in some cases multiple patch-level scale estimates, for each image patch of the image data, for example, in the form of a scene geometry map, during the operation of System 500. Generally, the processor subsystem 520 may be configured to perform any functions as previously described with reference to Figures 3 to 5C and elsewhere. It will be further understood that the same considerations and optional additional implementations as those for the processor subsystem 120 in Figure 1 apply to the processor subsystem 520 in Figure 6. It will be further understood that similar considerations and optional implementations are generally applicable to System 500 in Figure 6, as to System 100 in Figure 1, unless otherwise specified.
[0057] Figure 6 further illustrates various optional components of system 500. For example, in some embodiments, system 500 may include a sensor data interface 560 for direct access to sensor data 562 acquired by a sensor 620 in the environment 600. Sensor 620 may, but does not necessarily, be part of system 500. Sensor 620 can take any suitable form, such as an image sensor or other types of spatial sensors. The sensor data interface 560 may have any suitable form corresponding to a type of sensor, including, but not limited to, the types of data storage interfaces described above, with respect to a low-level communication interface, an electronic bus, or a data storage interface 540.
[0058] In some embodiments, the system 500 may include an output interface, such as a control interface 570, for providing control data 572 to an actuator 630 in, for example, an environment 600. Such control data 572 may be generated by a processor subsystem 520 to control the actuator 630 based on an analysis of the output of a scale estimator. For example, the actuator 630 may be an electrical, hydraulic, pneumatic, thermal, magnetic, and / or mechanical actuator. Specific but non-limiting examples include electric motors, electroactive polymers, hydraulic cylinders, piezoelectric actuators, pneumatic actuators, servo mechanisms, solenoids, stepping motors, etc. This allows the system 500 to operate, for example, to control a manufacturing process, a robotic system, or an autonomous vehicle, depending on an estimate of the relative scale of an object in the image data.
[0059] In other embodiments (not shown in Figure 6), the system 500 may include an output interface to rendering devices such as a display, light source, speaker, and vibration motor, which may be used to generate a sensory-perceptible output signal that can be generated based on the output of the scale estimator. The sensory-perceptible output signal may directly represent patch-level scale estimates or scene geometry maps, or it may represent a sensory-derived, perceptible output signal. By using a rendering device, the system 500 can provide the user with sensory-perceptible feedback.
[0060] In general, each system described herein, including but not limited to System 100 in Figure 1 and System 500 in Figure 6, may be implemented as a single device or apparatus, such as a workstation or server. The device may be an embedded device. The device or apparatus may include one or more microprocessors running appropriate software. For example, the processor subsystem of each system may be implemented by a single central processing unit (CPU), or by a combination or system of such CPUs and / or other types of processing units. The software may be downloaded and / or stored in corresponding memory, such as volatile memory such as RAM or non-volatile memory such as flash. Alternatively, the processor subsystem of each system may be implemented in the form of programmable logic, for example, as a field-programmable gate array (FPGA), in the device or apparatus. In general, each functional unit of each system may be implemented in the form of a circuit. Each system may be implemented using a distributed approach, including different devices or apparatus, such as a distributed local server or a cloud-based server. In some embodiments, System 500 may be part of a control system configured to control a physical entity or a manufacturing process, or it may be part of a data analysis system. In some embodiments, System 500 may be part of a vehicle, a robot or similar computer-controlled entity, and / or it may represent a control system configured to control said entity.
[0061] Figure 7 shows an example of the above, where system 500 is a control system for a (semi)autonomous vehicle 610 operating within an environment 600. This autonomous vehicle 600 can incorporate system 500 to control aspects of the autonomous vehicle, such as steering and braking, based on sensor data acquired from a camera 622 integrated within the vehicle 600. For example, system 500 can control the electric motor 632 to perform (regenerative) braking when the autonomous vehicle 600 is expected to encounter traffic congestion, as can be detected from a scene geometry map.
[0062] Figure 8 shows a computer-implemented method 700 for estimating the relative scale of objects in an image. While this method 700 can correspond to the operation of system 500 in Figure 6, it may also be implemented using or by any other system, machine, apparatus, or device. It is shown that method 700 includes a step 710, in which a feature extractor is provided as described elsewhere in this specification, in a step titled “Providing a feature extractor,” and a step 720, in which a scale estimator is provided as described elsewhere in this specification, in a step titled “Providing a scale estimator.” The Method 700 further includes a step 730 in which, in a step titled “Applying Feature Extractor and Scale Estimator to Image Patches,” the feature extractor and scale estimator are applied to at least one image patch of an image to obtain a patch-level scale estimate for at least one image patch; a step 740 in which the step 730 is optionally repeated for other image patches; and a step 750 in which, in a step titled “Outputting Data Representation of Patch-Level Scale Estimates,” the data representation of the patch-level scale estimates is output, for example, in the form of a scene geometry map as described elsewhere in this Spec.
[0063] In some embodiments, the computer-implemented method 200 in Figure 2 and the computer-implemented method 700 in Figure 8 may be implemented by the same computer program or by the same system. In other embodiments, the computer-implemented method 200 in Figure 2 and the computer-implemented method 700 in Figure 8 may be implemented by different computer programs or by different systems.
[0064] In general, it will be understood that the operations or steps of the computer-implemented methods 200 and 700 in Figures 2 and 8, respectively, may be performed in any suitable order, for example, sequentially, simultaneously, or in combination thereof, subject to the specific order required by the input / output relationship, where applicable.
[0065] Each method, algorithm, or pseudocode described herein may be implemented on a computer as a computer-implemented method, either as dedicated hardware or a combination of both. As also shown in Figure 9, computer instructions, such as executable code, may be stored on a computer-readable medium 800, for example, in the form of an array of machine-readable physical marks 810, and / or as a series of elements having different electrical, such as magnetic or optical properties or values. The executable code may be stored by a temporary or non-temporary method. Examples of computer-readable media include memory devices, optical storage devices, integrated circuits, servers, and online software. Figure 9 shows an optical disc 800. In alternative embodiments, the computer-readable medium 800 may include data representations of feature extractors and / or scale estimators, as described elsewhere in this specification.
[0066] Examples, embodiments, or optional features, whether shown non-limitingly or otherwise, should not be understood as limiting the invention as described in the claims.
[0067] Mathematical symbols and notations are provided to facilitate the interpretation of the present invention and should not be construed as limiting the scope of the claims.
[0068] The embodiments described above are illustrative rather than limiting, and it should be noted that those skilled in the art can design many alternative embodiments without departing from the scope of the appended claims. Reference numerals in parentheses in the claims should not be construed as limiting the claims. The use of the verb “includes” and its conjugations does not preclude the existence of elements or steps other than those described in the claims. The articles “a” or “an” preceding an element do not preclude the existence of multiple such elements. Expressions such as “at least one” preceding a list or group of elements indicate a selection of all or any subset of elements from the list or group. For example, the expression “at least one of A, B, and C” should be understood to include A only, B only, C only, both A and B, both A and C, both B and C, or all of A, B, and C. The present invention may be implemented by hardware comprising multiple different elements and by a appropriately programmed computer. In an apparatus claim listing multiple means, some of these means may be embodied by hardware of the same component. The mere fact that certain means are described in different dependent claims does not imply that combinations of these means cannot be used advantageously.
Claims
1. 1. A computer-implemented method (200) for training a machine-learning capable model for estimating relative scale of objects in an image, comprising: The method comprises: - providing (210) a feature extractor (360), said feature extractor (360) comprising: - receiving as input an image patch of an image, - configured to detect a number of features in the image patch to obtain as output a number of feature maps, the number of features being related to one or more objects in the image, each feature map being generated by applying a filter to image data of the image patch, each feature map comprising a filter response across a set of different spatial scales along a scale dimension; - providing (220) a scale estimator (380) that processes the output of the feature extractor, said scale estimator including a machine learning capable model part; aggregating each feature map into a feature-level scale estimate, said aggregation comprising identifying maximum filter responses across different spatial scales, thereby obtaining a plurality of feature-level scale estimates; - configured to infer a patch-level scale estimate from the plurality of feature-level scale estimates using the machine learning capable model portion (220); Including, The method further comprises: - accessing (230) training data comprising a set of training images; - training (240) the machine learning capable model portion of the scale estimator based on the training data to infer the patch-level scale estimate from the plurality of feature-level scale estimates; Including, The training step includes: - spatially scaling (250) image data of an image patch (300) of said training image by at least two known scale factors to obtain at least two further image patches (310, 320); - applying (260) said feature extractor and said scale estimator to said at least two further image patches to obtain at least two patch level scale estimates; - optimizing (270) parameters of the machine-learnable model portion by minimizing an error term of a loss function, the error term expressing a mismatch between an actual relative scale and an estimated relative scale, the actual relative scale being determined as the difference between two known scale factors and the estimated relative scale being determined as the difference between at least two patch-level scale estimates; A computer-implemented method (200), comprising:
2. 2. The computer-implemented method of claim 1, wherein each feature map includes at least two spatial dimensions and a scale dimension, and the scale estimator is configured to aggregate the respective feature maps across the at least two spatial dimensions by averaging, weighting, or majority voting.
3. 2. The computer-implemented method of claim 1, wherein the scale estimator is configured to identify a spatial scale at which a filter response is maximum and aggregate the respective feature maps across the scale dimension by using an identifier of the spatial scale as or as part of a feature-level scale estimate.
4. The computer-implemented method of claim 1 , wherein the machine learning capable model portion of the scale estimator comprises a neural network.
5. 5. The computer-implemented method of claim 4, wherein the neural network is a shallow neural network having one hidden layer.
6. The computer-implemented method of claim 1 , wherein the error term defines a mean squared error or a mean squared deviation between the actual relative scale and the estimated relative scale.
7. 1. A computer-implemented method (700) for estimating relative scale of objects in an image, comprising: The method comprises: - providing (710) a feature extractor (360), said feature extractor (360) comprising: - receiving as input an image patch of an image, - a step (710) configured to detect a number of features in the image patch to obtain as output a number of feature maps, the number of features being related to one or more objects in the image, each feature map being generated by applying a filter to image data of the image patch, each feature map comprising a filter response across a set of different spatial scales along a scale dimension; - providing (720) a scale estimator (380) for processing the output of the feature extractor, the scale estimator comprising a machine-learned model part trained by the method according to any one of claims 1 to 6, the scale estimator comprising: aggregating each feature map into a feature-level scale estimate, said aggregation comprising identifying maximum filter responses across different spatial scales, thereby obtaining a plurality of feature-level scale estimates; - configured to infer a patch-level scale estimate from the plurality of feature-level scale estimates using the machine learning capable model portion (720); - applying (730) said feature extractor and said scale estimator to at least one image patch of said image to obtain a patch-level scale estimate for said at least one image patch; - outputting (750) a data representation of said patch level scale estimate; A computer-implemented method (700), comprising:
8. The method further comprises: generating a scene geometry map (420) indicative of a scene geometry for the image; - applying said feature extractor and said scale estimator to a number of image patches (410, 412) of said image to obtain a number of patch-level scale estimates; - generating a scene geometry map for said image as a representation of said patch level scale estimates associated with said image patches; 8. The computer-implemented method of claim 7, further comprising the step of generating by:
9. The computer-implemented method (700) of claim 8, further comprising applying the feature extractor and the scale estimator to overlap image patches of the images.
10. The method comprises: - subtracting the minimum of the plurality of patch level scale estimates from the plurality of patch level scale estimates in the scene geometry map; and - spatially scaling (430) the scene geometry map to the spatial resolution of the image; The computer-implemented method of claim 8 , further comprising at least one of:
11. The method comprises: - acquiring said images from sensors (620, 622) configured to sense the environment (600) of the computer-controlled entity (610); - analyzing a scene geometry map of said image; - generating control data for said computer-controlled entity, based on the results of said analysis, in order to adaptively control said computer-controlled entity in its environment; 10. The computer-implemented method of claim 8, further comprising:
12. 12. The computer-implemented method (700) of claim 11, wherein the computer-controlled entity is a robotic system or an autonomous vehicle (610).
13. A computer readable medium (800) comprising transitory or non-transitory data (810) representing instructions arranged to cause a processor system to perform a computer-implemented method according to any one of claims 1 to 12.
14. 1. A system (100) for training a machine-learning capable model for estimating relative scale of objects in an image, comprising: The system (100) comprises: an input interface subsystem (140), said input interface subsystem (140) comprising: training data (152) comprising a set of training images; a feature extractor (154, 360), said feature extractor (154, 360) comprising: - receiving as input an image patch of an image, a feature extractor (154, 360) configured to detect a number of features in the image patch to obtain as output a number of feature maps, the number of features being related to one or more objects in the image, each feature map being generated by applying a filter to image data of the image patch, each feature map comprising a filter response across a set of different spatial scales along a scale dimension; a scale estimator (156, 380) that processes the output of the feature extractor, the scale estimator including a machine learning capable model part; and aggregating each feature map into a feature-level scale estimate, said aggregation comprising identifying maximum filter responses across different spatial scales, thereby obtaining a plurality of feature-level scale estimates; a scale estimator (156, 380) configured to infer a patch-level scale estimate from the plurality of feature-level scale estimates using the machine learning capable model portion; is configured to access The system (100) further comprises: a processor subsystem (120) configured to train the machine learning capable model portion of the scale estimator based on the training data to infer the patch level scale estimate from the plurality of feature level scale estimates, the training comprising: - spatially scaling image data of an image patch (300) of said training image by at least two known scale factors to obtain at least two further image patches (310, 320); - applying said feature extractor and said scale estimator to said at least two further image patches to obtain at least two patch level scale estimates; - optimizing parameters of the machine-learnable model portion by minimizing an error term of a loss function, the error term representing a mismatch between an actual relative scale and an estimated relative scale, the actual relative scale being determined as the difference between two known scale factors and the estimated relative scale being determined as the difference between at least two patch-level scale estimates; A system (100).
15. A system (500) for estimating a relative scale of an object in an image, comprising: The system (500) comprises: an input interface subsystem (540), said input interface subsystem (540) comprising: - image (552), a feature extractor (554, 360), said feature extractor (554, 360) comprising: - receiving as input an image patch of an image, a feature extractor (554, 360) configured to detect a number of features in the image patch to obtain as output a number of feature maps, the number of features being related to one or more objects in the image, each feature map being generated by applying a filter to image data of the image patch, each feature map comprising a filter response across a set of different spatial scales along a scale dimension; - a scale estimator (556, 380) that processes the output of the feature extractor, the scale estimator comprising a machine-learned model part trained by the method according to any one of claims 1 to 6, the scale estimator comprising: aggregating each feature map into a feature-level scale estimate, said aggregation comprising identifying maximum filter responses across different spatial scales, thereby obtaining a plurality of feature-level scale estimates; a scale estimator (556, 380) configured to infer a patch-level scale estimate from the plurality of feature-level scale estimates using the machine learning enabled model portion; is configured to access The system (500) further comprises: - a processor subsystem (520), said processor subsystem (520) comprising: - applying said feature extractor and said scale estimator to at least one image patch of said image to obtain a patch level scale estimate for said at least one image patch; - a system (500) configured to output a data representation of said patch level scale estimate.