Method and apparatus
The method addresses the challenges of conventional training data generation for autonomous vehicles by transforming environment representations to create comprehensive and efficiently generated training data, enhancing the safety and efficiency of AV control software.
Patent Information
- Application Number
- JP2024521767
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-10-15
- Filing Date
- 2022-10-17
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-10-17
AI Technical Summary
Conventional methods for testing control software in autonomous vehicles are costly, time-consuming, and struggle to capture low-probability events, making it difficult to obtain comprehensive training data.
A computer-implemented method that generates training data by transforming a representation of an environment, which is partially synthesized using semantic information, into a set of transformed representations, allowing for accelerated and automated labeling of data, including low-probability events.
This method improves the efficiency and comprehensiveness of training data generation for autonomous vehicles, reducing costs and enhancing the safety of AV control software by simulating a wide range of scenarios.
Smart Images

Figure 0007696503000001 
Figure 0007696503000002 
Figure 0007696503000003
Abstract
Description
Technical Field
[0001] The present invention relates to an autonomous vehicle.
Background Art
[0002] For example, there are problems with conventional testing of control software for autonomous vehicles (AVs) (also known as the AV stack) according to SAE levels 1 to 5. For example, for testing the control software for autonomous vehicles, for example, for installation, warranty, verification, validation, regression, and / or progression testing, conventional methods for obtaining labeled training data typically include: 1. Collection of data based on specific requirements (scenarios, scene structure, scene appearance, weather, operating area, occurrence of rare events), and 2. Labeling of data (instance segmentation in images and / or LIDAR space, labeling of actions, etc.).
[0003] These conventional methods are not only extremely costly and time-consuming, but also require capturing low-probability events, which are often impossible.
[0004] Therefore, it is necessary to obtain training data.
Summary of the Invention
[0005] A first aspect is a computer-implemented method for generating training data, the method comprising: providing a representation of an environment, the representation of the environment having a defined structure and / or a defined shape; and generating training data including a set of transformed representations including a first transformed representation of the environment by transforming the representation of the environment into a set of transformed representations including the first transformed representation of the environment; and Provide a computer-implemented method, wherein the step of providing a representation of an environment includes at least partially synthesizing an image of the environment using semantic information.
[0006] The term "training data" can be extended to include training data, test data, validation data, and proof-of-concept data.
[0007] A second aspect provides a computer-implemented method for training a machine learning (ML) algorithm, the method comprising: generating training data including a set of transformed representations, including a first transformed representation of an environment, according to a first aspect; training the ML algorithm by classifying the set of transformed representations according to a set of classes including a first class. A computer-implemented method is provided that includes the above steps.
[0008] A third aspect provides a computer-implemented method for determining a class of a representation of an environment using a machine learning (ML) algorithm trained according to the second aspect, the method comprising: determining a class of a representation of the environment, including inferring the class of the representation of the environment, using the trained ML algorithm. A computer-implemented method is provided that includes the above steps.
[0009] A fourth aspect provides a computer-implemented method for testing an ego-vehicle, e.g., its control software, e.g., for installation, warranty, verification, proof-of-concept, regression, and / or progression testing, the method comprising: generating training data according to the first aspect; simulating a scenario including a first transformed representation of an environment having therein the ego-vehicle, a set of actors including a first actor, and optionally a set of objects including a first object. including a step of simulating a first scenario, where the step of simulating the first scenario includes a step of identifying a defect of the host vehicle in the scenario provides a computer-implemented method.
[0010] A fifth aspect provides a computer including a processor and a memory configured to execute a method according to the first aspect, the second aspect, the third aspect, and / or the fourth aspect.
[0011] A sixth aspect provides a computer program including instructions that, when executed by a computer including a processor and a memory, cause the computer to execute a method according to the first aspect, the second aspect, the third aspect, and / or the fourth aspect.
[0012] A seventh aspect provides a non-transitory computer-readable storage medium including instructions that, when executed by a computer including a processor and a memory, cause the computer to execute a method according to the first aspect, the second aspect, the third aspect, and / or the fourth aspect.
[0013] (Detailed Description of the Invention) According to the present invention, a method described in the appended claims is provided. Also provided are a computer program, a computer, a non-transitory computer-readable storage medium, and a vehicle. Other features of the present invention will become apparent from the dependent claims and the following description.
[0014] (Method for Generating Training Data) A first aspect is a computer-implemented method for generating training data, the method comprising: providing a representation of an environment, the representation of the environment having a defined structure and / or a defined shape; generating training data including a set of transformed representations including a first transformed representation of the environment by transforming the representation of the environment into a set of transformed representations including the first transformed representation of the environment. comprising, providing a computer-implemented method, wherein a step of providing a representation of an environment includes at least partially synthesizing an image of the environment using semantic information.
[0015] Since the training data is generated by transforming a representation of an at least partially synthesized environment, the generation of the training data is accelerated and can be automatically labeled if the ground truth is maintained. Since the representation of the environment is at least partially synthesized using semantic information, it can represent low-probability events, thereby providing a more comprehensive test and thus improving the safety of the AV control software. In this way, obtaining training data for AVs is improved.
[0016] Additionally and / or alternatively, a first aspect provides a computer-implemented method for generating training data for a machine learning model, wherein the method is based on image generation from an abstract representation (semantic information) and / or wherein the method is based on one or more learned or heuristic-based image transformations.
[0017] Examples of transformations include weather editing, partial or complete image synthesis, road surface manipulation, dynamic actor manipulation, and combinations thereof. The transformations can be chained.
[0018] In particular, the transformation of the data is designed to improve the performance of a model trained with such data and does not necessarily need to exactly resemble natural / realistic images.
[0019] In other words, the abstract representation of the structure of a scene is used as guidance for generating visual training data that maximizes the performance of a visual machine learning model trained thereon. Maximizing performance does not necessarily mean that the data lies on a manifold of realistic / natural images.
[0020] Existing solutions focus on photorealism and not on generating maximally informative training data. This approach focuses on generating optimal training data, which may not follow or lie on the manifold of natural / realistic images. This means that during the training process, it is possible to measure how informative the data is.
[0021] In contrast to conventional methods, the method according to the first aspect generates in silico training data by directly or by leveraging simulations and configurations in a simple space (e.g., semantic segmentation), and then synthesizes sufficiently realistic images from this representation (this aim is to synthesize optimal training data for a given task or model). Often, this can be achieved by following a distribution different from that of natural images. Furthermore, the inventors enable the composability of the transformation configurations.
[0022] Furthermore, actual data and synthetic data are transformed / adapted to follow different distributions, e.g., from day to night, again with the aim of obtaining optimal training data.
[0023] Furthermore, this method can directly transform the structure of existing (natural or synthetic) images. Examples of this include the movement / removal / placement of road actors (pedestrians, vehicles) and the manipulation of road surfaces and structures (road signs, lanes, etc.), as will be described below.
[0024] (Computer-implemented method) This method is computer-implemented by a computer including, for example, a processor and a memory. Suitable computers are known.
[0025] This method generates training data (i.e., multiple data, and note that the singular form is datum) for training a machine learning (ML) algorithm, for example, according to a second aspect. The ML algorithm can be as described with respect to the second aspect.
[0026] (Provision of an environmental representation) This method includes providing a representation of the environment. Generally, a scenario includes an environment having, therein, a host vehicle, a set of actors (i.e., at least one actor), and optionally a set of objects including a first object. The environment, also known as a scene, typically includes one or more roads having one or more lanes and optionally one or more obstacles, as understood by those skilled in the art. Generally, the host vehicle is the target connected vehicle and / or motor vehicle, and its behavior is a major concern in test, trial, or operation scenarios. It should be understood that the behavior of the host vehicle is defined by its control software (also known as the AV stack). In one example, the first actor is a road user, such as a vehicle, pedestrian, or cyclist. Other road users are also known. In one example, the first object includes and / or is infrastructure, such as a traffic signal, or a static road user. In one example, the set of actors includes A actors, where A is a natural number of 1 or more, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more. In one example, the set of objects includes O objects, where O is a natural number of 1 or more, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more. In one example, the environmental representation is defined with reference to an image, an encoded image, a synthesized image, an image defined with reference to other sub-images, a vector graphic, etc., saved, for example, in a raw data format (binary, bitmap, TIFF, MRC, etc.) or another image data format (PNG, JPEG).
[0027] The representation of the environment has a defined structure and / or a defined shape and thus provides ground truth. That is, the representation of the environment includes one or more roads having one or more lanes and optionally one or more obstacles, as understood by a person skilled in the art.
[0028] In one example, the step of providing a representation of the environment includes at least partially acquiring (also known as capturing) an image of the environment. Thus, the representation of the environment may be partially synthesized and partially acquired, like a mosaic.
[0029] In one example, the step of providing a representation of the environment includes at least partially semantically structuring an image of the environment. Semantic structuring is known. Thus, a target environment including low-probability events may be semantically structured.
[0030] In one example, the step of providing a representation of the environment includes inpainting an image of the environment. Thus, the image can be rendered for training.
[0031] (Generation of training data) This method includes generating training data including a set of transformed representations including a first transformed representation of the environment by transforming the representation of the environment into a set of transformed representations including the first transformed representation of the environment. In one example, the set of transformed representations includes T transformed representations, where T is a natural number greater than or equal to 1, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 200, 500, 1,000, 2,000, 5,000, 10,000, 20,000, 50,000 or more. Thus, a sufficient dataset for training can be provided. In particular, transforming the representation can be as described in https: / / arxiv.org / pdf / 1907.11004.pdf, which is hereby incorporated by reference in its entirety.
[0032] In one example, the step of generating training data by converting a representation of an environment into a set of transformed representations including a first transformed representation of the environment includes using each set of trained transformations (also known as adapters) including a first trained transformation to convert the representation of the environment into a set of transformed representations including a first transformed representation of the environment, thereby generating training data. In this way, the generation of training data is improved.
[0033] In one example, the step of generating training data by converting a representation of an environment into a set of transformed representations including a first transformed representation of the environment includes using each set of heuristics-based transformations including a first heuristics-based transformation to convert the representation of the environment into a set of transformed representations including a first transformed representation of the environment, thereby generating training data. Heuristics-based transformations are known.
[0034] In one example, the step of generating training data by converting a representation of an environment into a set of transformed representations including a first transformed representation of the environment includes using each set of augmentations including a first augmentation to convert the representation of the environment into a set of transformed representations including a first transformed representation of the environment, thereby generating training data. Image augmentation is known.
[0035] In one example, the step of generating training data by converting a representation of an environment into a set of transformed representations including a first transformed representation of the environment includes converting the representation of the environment into a set of transformed representations including a first transformed representation of the environment using each set of conditions including a first condition. In one example, the first condition is a weather condition (e.g., sun, rain, cloud, snow, drizzle, fog), a season condition (e.g., spring, summer, autumn, winter), a time condition (day, night), and a brightness condition (bright sun, streetlight, headlight). In one example, the set of conditions includes C conditions, where C is a natural number greater than or equal to 1, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 200, 500, or more. In one example, converting the representation of the environment into a set of transformed representations including a first transformed representation of the environment using each set of conditions including the first condition includes using a plurality of conditions of the set. Thus, the conditions may be combined.
[0036] In one example, the step of generating training data by converting a representation of an environment into a set of transformed representations including a first transformed representation of the environment includes mixing a set of representations of the environment. Thus, additional representations of the environment can be generated at low cost.
[0037] In one example, the first transformed representation of the environment has a defined structure and / or a defined shape of the representation of the environment. Thus, the ground truth of the representation of the environment is maintained with respect to the first transformed representation of the environment.
[0038] In one example, the step of generating training data by converting a representation of an environment into a set of transformed representations including a first transformed representation of the environment includes redefining a defined structure of the representation of the environment into a redefined structure of the first transformed representation of the environment. Thus, the ground truth of the representation of the environment is redefined with respect to the first transformed representation of the environment, for example, in a known manner.
[0039] (Synthesis of environmental images) The step of providing a representation of the environment includes at least partially synthesizing an image of the environment using semantic information.
[0040] In one example, at least partially synthesizing an image of the environment using semantic information includes obtaining an image or a portion thereof corresponding to the semantic information from a database or by learning.
[0041] (Method for training an ML algorithm) A second aspect is a computer-implemented method for training a machine learning (ML) algorithm, the method comprising generating training data including a set of transformed representations including a first transformed representation of the environment according to the first aspect; and training an ML algorithm including classifying the set of transformed representations according to a set of classes including a first class. A computer-implemented method is provided that includes.
[0042] Thus, the set of transformed representations is classified, for example, to test a particular scenario including the transformed representation of the environment.
[0043] (Computer-implemented method) This method is computer-implemented by a computer including, for example, a processor and a memory. Suitable computers are known.
[0044] This method trains an ML algorithm.
[0045] (Generation of training data) This method includes generating training data including a set of transformed representations including a first transformed representation of the environment according to the first aspect.
[0046] (Classification of representations) This method includes the step of training an ML algorithm that includes classifying a set of representations transformed according to a set of classes including a first class.
[0047] In one example, the set of classes including the first class is a set of conditions including a first condition, for example, as described with respect to the first aspect.
[0048] In one example, this method includes the step of identifying a set of characteristic features (also known as intermediate features) including a first characteristic feature associated with the first condition. In this way, the characteristic features or distinguishing features associated with the first condition can be identified and used for comparison such as discovering new conditions.
[0049] (Method for determining a class) A third aspect is a computer-implemented method for determining a class of a representation of an environment using a machine learning (ML) algorithm trained according to the second aspect, the method comprising: Determining a class of a representation of the environment, including inferring a class of a representation of the environment using the trained ML algorithm including, to provide a computer-implemented method.
[0050] (Computer-implemented method) This method is computer-implemented by a computer including, for example, a processor and a memory. Suitable computers are known.
[0051] This method determines a class of a representation of the environment, for example, as described with respect to the second aspect.
[0052] (Inference of a class of a representation) This method includes the step of determining a class of a representation of the environment, including inferring a class of a representation of the environment using a trained ML algorithm, for example, as described with respect to the second aspect.
[0053] In one example, the method includes calculating a confidence score for the inferred class. For example, the calculated confidence can be used during testing.
[0054] In one example, the method includes identifying a set of features including a first feature associated with a condition of the environmental representation.
[0055] In one example, the method includes comparing the identified set of features with a set of characteristic features associated with the condition of the environmental representation.
[0056] In one example, the method includes memorizing the environmental representation based on the result of the comparison.
[0057] In one example, the method includes training the transformation using the memorized environmental representation.
[0058] In one example, the method includes generating training data using the trained transformation, for example, according to a first aspect.
[0059] In one example, the method includes training an ML algorithm using the generated training data.
[0060] In one example, the method includes validating an ML algorithm using the generated training data.
[0061] In one example, the method includes performing an operation based on the result of the comparison. Thus, downstream tasks can be trained or adjusted, for example, by selecting its parameters.
[0062] (Testing method) A fourth aspect is a computer-implemented method for testing a host vehicle, for example, its control software, for example, a computer-implemented method for testing installation, warranty, verification, validation, regression, and / or progression, the method comprising: generating training data according to a first aspect; simulating a scenario comprising a first transformed representation of an environment having therein the host vehicle, a set of actors including a first actor, and optionally a set of objects including a first object; wherein the step of simulating the first scenario comprises: identifying a defect in the host vehicle in the scenario. A computer-implemented method is provided.
[0063] (Computer, computer program, non-transitory computer-readable storage medium) A fifth aspect provides a computer comprising a processor and a memory configured to execute a method according to the first, second, third, and / or fourth aspects.
[0064] A sixth aspect provides a computer program comprising instructions which, when executed by a computer comprising a processor and a memory, cause the computer to execute a method according to the first, second, third, and / or fourth aspects.
[0065] A seventh aspect provides a non-transitory computer-readable storage medium comprising instructions which, when executed by a computer comprising a processor and a memory, cause the computer to execute a method according to the first, second, third, and / or fourth aspects.
[0066] (Definition) Throughout this specification, the terms "comprising" or "comprises" mean including the specified element(s) but not excluding the presence of other elements. The terms "consisting essentially of" or "consists essentially of" mean including the specified element(s) but excluding other elements except for materials that are present as impurities, inevitable materials that result from the processes used to provide the elements, and components that are added for purposes other than achieving the technical effects of the present invention, such as colorants.
[0067] The terms "consisting of" or "consists of" mean including the specified element(s) but excluding other elements.
[0068] Whenever appropriate and depending on the context, the use of the term "include" or "comprising" can also be interpreted to include the meaning of "consisting essentially of" or "consists essentially of", and can also be interpreted to include the meaning of "consisting of" or "consists of".
[0069] Any feature presented in this specification can, where appropriate, be used individually or in combination with each other, particularly in the combinations presented in the appended claims. Any feature with respect to each aspect or exemplary embodiment of the present invention presented in this specification can, where appropriate, also be applicable to all other aspects or exemplary embodiments of the present invention. In other words, those skilled in the art reading this specification should consider any feature with respect to each aspect or exemplary embodiment of the present invention to be interchangeable and combinable between different aspects and exemplary embodiments.
Brief Description of the Drawings
[0070] For a better understanding of the present invention and to show how its exemplary embodiments can be implemented, reference is made, for illustrative purposes only, to the accompanying drawings.
[0071]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
[0072] Generally, FIGS. 1 - 14 schematically show a computer - implemented method for generating training data as described with respect to a first aspect. The method includes: providing a representation of an environment, wherein the representation of the environment has a defined structure and / or a defined shape; and generating training data including a set of transformed representations including a first transformed representation of the environment by transforming the representation of the environment into a set of transformed representations including the first transformed representation of the environment. The step of providing a representation of the environment includes at least partially synthesizing an image of the environment using semantic information.
[0073] FIG. 1 schematically shows in detail the method according to an exemplary embodiment.
[0074] In particular, FIG. 1 shows a corpus 1 including labeled data DL, where the data is manually or automatically labeled data DU and unlabeled data DU. The labeled data DL includes an original image IO of the environment and its respective semantic map / object position / depth SMOD. The unlabeled data DU includes a set of conditions including a first condition C1 (e.g., wet), a second condition C2 (e.g., snow), and a third condition C3 (e.g., night).
[0075] FIG. 2 schematically shows in detail method 2 according to an exemplary embodiment.
[0076] In this example, the step of generating training data by converting a representation of an environment into a set of transformed representations including a first transformed representation of the environment includes using each set of conditions including a first condition to generate training data by converting a representation of the environment into a set of transformed representations including a first transformed representation of the environment.
[0077] In this example, the step of generating training data by converting a representation of an environment into a set of transformed representations including a first transformed representation of the environment optionally includes using each set of trained transformations (also known as adapters) including a first trained transformation to generate training data by converting a representation of the environment into a set of transformed representations including a first transformed representation of the environment.
[0078] In this example, the step of generating training data by converting a representation of an environment into a set of transformed representations including a first transformed representation of the environment optionally includes using each set of heuristics-based transformations including a first heuristics-based transformation to generate training data by converting a representation of the environment into a set of transformed representations including a first transformed representation of the environment.
[0079] In one example, the step of generating training data by converting a representation of an environment into a set of transformed representations including a first transformed representation of the environment optionally includes using each set of augmentations including a first augmentation to generate training data by converting a representation of the environment into a set of transformed representations including a first transformed representation of the environment.
[0080] Referring to FIG. 2, data augmentation can be performed by executing an image enhancement task 20. The image enhancement task 20 can be executed using a Generative Adversarial Network (GAN). The original image IO and the semantic map / object / depth SMOD are inputs to the GAN. The GAN is trained to output an enhanced image based on the original image. The enhanced images IA1, IA2 can be the original image IO enhanced with conditions, such as attached rain IA1 or attached soil IA2. Since the structures of the enhanced images IA1, IA2 are the same as the original image, the ground truth, such as the semantic map, object position, depth, etc., remains valid.
[0081] The semantic map, object position, depth SMOD, original image IO, enhanced images IA1, IA2 may be labeled data. Thus, training the GAN can be supervised. Therefore, the GAN can be trained to enhance the original image under one or more conditions. The conditions can include weather conditions, brightness conditions, etc.
[0082] FIG. 3A shows the original image, FIG. 3B shows the enhanced image (attached rain), and FIG. 3C shows the enhanced image (attached rain) generated by the method described with respect to FIG. 2. Each ground truth is valid for the original image or the enhanced image (FIGS. 3B and 3C) including the conditions.
[0083] FIG. 4 schematically shows a method for generating a training data example.
[0084] In particular, FIG. 4 shows an augmentation 4 of the labeled data DL using unpaired translation methods. The image translation method can be the image translation task 20. The image translation task can translate an image to a different domain using a cycle-consistent GAN. More specifically, in addition to the semantic map, object position, and depth SMOD, the enhanced image IA and the original image OI generated from the method using the method of FIG. 2 are input into the cycleGAN. Various conditions, such as condition 1 (e.g., wet) C1, condition 2 (e.g., snow) C2, condition 3 (e.g., night) C3, can be input into the cycleGAN in the form of unlabeled data. Since the conditional data is unlabeled, the cycleGAN is trained in an unsupervised manner.
[0085] The cycleGAN can be trained to output the original image or the enhanced image translated to different conditions, such as condition 1 C1, condition 2 C2, or condition 3 C3. The original image or the enhanced image translated to different conditions IC1, IC2, IC3 may be referred to as the translated image. The translated images IC1, IC2, IC3 can retain the respective structures of the original image IO or the enhanced image IA, and thus the respective ground truths (semantic map, object position, depth) can still be valid.
[0086] FIGS. 5A and 5C show examples of the original image IO, and FIGS. 5B and 5D show examples of the translated images IC1, IC2 generated using the cycleGAN of FIG. 4. The exemplary examples of the translated images in FIGS. 5B and 5D are generated by translating the original image IO from FIGS. 5A and 5C at a low brightness level of illumination, which is artificial light from night-time conditions, such as street lights, buildings, vehicles, etc.
[0087] The original image IO in FIG. 5A shows a rural road without a median strip containing a bus during the day. FIG. 5B shows the translated image IC1 of a rural road without a median strip containing a bus at night.
[0088] The original image IO of FIG. 5C shows a road with a median strip in a city including automobiles and cyclists during the day. FIG. 5D shows the translated image IC2 of a road with a median strip in a city including automobiles and cyclists at night. Each ground truth is valid for the original image or the enhanced image (FIGS. 5B and 5D) including the condition.
[0089] Note that the discrete cycleGAN model trained with reference to FIG. 4 and FIGS. 5A - 5D can be trained with respect to a single condition, for example, during the day or at night. Thus, it may be necessary to train multiple discrete cycleGAN models, one for each of the known conditions, to generate training data translated for each of the known conditions.
[0090] FIG. 6 schematically shows a method for generating synthetic training data. More specifically, FIG. 6 shows a method for generating a synthetic training image based on a semantic map.
[0091] In this example, providing a representation of the environment includes inpainting an image of the environment.
[0092] Referring to FIG. 6, an inpainting and / or synthesis task 60 is performed. The inpainting and / or synthesis task can be performed by an inpainting model or a synthesis model that can take the form of an autoencoder AE such as a GAN, a cycle-consistent GAN, or a variational autoencoder VAE. Also, other models including vision transformers, diffusion models, etc. can be used.
[0093] The inpainting model or synthesis model can receive the semantic map SM1. The inpainting model or synthesis model can receive an image. The image can be any of the original image IO, the enhanced image IA, or the translated images IC1, IC2, IC3. The inpainting model or synthesis model can be trained to generate a synthesized image IS based on features from the semantic map SM1 in the style of the received image. A loss between the synthesized image IS and the target image IT is determined. The target image IT can be an image received by the inpainting model or synthesis model. Thus, the target image IT can be the original image IO, the enhanced image IA, or the translated images IC1, IC2, IC3. To reduce or minimize the loss between the synthesized image IS and the target image IT, the parameterisation of the inpainting model or synthesis model can be changed or optimized.
[0094] Figures 7A and 7C show the semantic map SM1. Figures 7B and 7D show the synthesized image IS generated using the synthesis model trained according to the method of FIG. 6.
[0095] Figure 7A shows the semantic map SM1 of a rural road including a car. Figure 7B shows the synthesized image IS of the rural road including a car during the day. In other words, Figure 7B shows an image in the style of the image received by the synthesis model showing features from the semantic map SM1 of Figure 7A.
[0096] Figure 7C shows a different semantic map SM1 of a road without a median strip in the suburbs, including trees and buildings but no other vehicles. Figure 7D shows the synthesized image IS of a road without a median strip in the suburbs during the day, including trees and buildings but no other vehicles, according to the semantic map SM1.
[0097] Figure 8 schematically shows a method for generating a training image from a semantic map.
[0098] In this example, providing a representation of the environment includes, at least in part, semantically constructing an image of the environment.
[0099] The method of FIG. 8 operates similarly to the method of FIG. 6 with the addition of a semantic map composer 80. The semantic map composer 80 may be a semantic map composer model and may include an AE such as a GAN, a cycle-consistent GAN, or a VAE. Other models including a vision transformer, a diffusion model, etc. may also be used.
[0100] The semantic map composer model can be trained to generate a new semantic map SM by combining features from a plurality of semantic maps from a corpus of labeled data LD.
[0101] FIGS. 9A and 9C show pairs of a semantic map SM1 and a synthetic image obtained using a synthetic model where the semantic map SM1 is a semantic map from a corpus of labeled data LD. FIGS. 9B and 9D show pairs of new semantic maps SM generated by a semantic map composer model based on the semantic maps of FIGS. 9A and 9C, respectively. The new semantic maps SM show the same roads as the respective semantic maps SM1 but have different road markings. In each case, the semantic map SM1 or the new semantic map SM is used by the synthetic model to generate an image in the style of the received image.
[0102] Each ground truth is valid for the original image or the enhanced image (FIGS. 9B and 9D) including the conditions.
[0103] FIG. 10 schematically shows in detail a method according to an exemplary embodiment and is generally as described with respect to FIG. 8.
[0104] In this example, the step of generating training data by converting the representation of the environment into a set of transformed representations including the first transformed representation of the environment includes mixing a set of representations of the environment.
[0105] In particular, FIG. 10 shows an extension 10 of labeled data DL for generating a final composite image IF derived from a semantic segmentation map SMn by directly combining or mixing any other image IK (real image, enhanced image, or synthetic image) corresponding to the semantic segmentation map SMn of distilled diverse data to accelerate the training of other tasks. The original image or enhanced images IC1 (including the first condition C1), IC2 (including the second condition C2), and IC3 (including the third condition C3) included in the labeled data LD similarly retain the respective structures of the original image IO and the enhanced image IA, and thus, each ground truth (semantic map / object position / depth SMOD) remains valid for the original image or the enhanced image including the conditions IC1, IC2, IC3.
[0106] FIG. 11A shows an original image (urban intersection, car, summer, building), FIG. 11B shows an enhanced image including conditions (urban intersection, car, winter, trees), FIG. 11C shows an original image (urban intersection, car, cloud cover 1 / 8, building), and FIG. 11D shows an enhanced image including conditions (urban intersection, car, cloud cover 7 / 8, trees) generated by the method described with respect to FIG. 10. Each ground truth is valid for the original image or the enhanced image including conditions (FIGS. 11B and 11D).
[0107] FIG. 12 schematically shows in detail a method according to an exemplary embodiment.
[0108] In this example, synthesizing at least partially an image of the environment using semantic information includes obtaining from a database an image or a portion thereof corresponding to the semantic information.
[0109] In particular, FIG. 12 shows that an actual image, enhanced image, or synthetic image (with and without a common segmentation map) can be decomposed into building blocks (referred to as blobs) using their segmentation maps. A plurality of blobs can be formed in a blob database (an “actual” image blob database IR DB and a “synthetic” image blob database IS DB). The scene composer 120 is used to combine blobs from the plurality of databases IR DB, IS DB into a new synthetic image IF and a segmentation map.
[0110] FIG. 13A shows an original image, and FIG. 13B shows an enhanced image including a blob, e.g., a loading platform storing rubble. FIG. 13C shows an original image, and FIG. 13D shows an enhanced image including a blob, e.g., a car, generated by the method described with respect to FIG. 12. Each ground truth is valid for the original image or the enhanced image including the conditions (FIGS. 13B and 13D).
[0111] FIGS. 14A - 14F show examples of the above-described transformations applied sequentially to obtain new conditions, appearances, structures, and training data. FIG. 14A shows an original fake (suburb, two lanes, cloud cover 7 / 8, winter, day), FIG. 14B shows a changed road marking (three lanes), FIG. 14C further shows a class switch (autumn), FIG. 14D further shows an added road user (cyclist), FIG. 14E further shows a changed condition (e.g., weather) (rain), and FIG. 14F further shows a changed time of day (night).
[0112] As described above, the final synthetic image may be partially synthesized and partially an actual image (as shown in FIGS. 10 - 13D).
[0113] As described above, the final composite image may be derived from an abstract representation (e.g., a semantic map or a bounding box, FIGS. 6-7D), and / or may be the result of a transformation on another image (FIGS. 2-5D).
[0114] As will be appreciated by those skilled in the art, the image synthesizer network (e.g., SPADE in this particular example) is interchangeable with other architectures.
[0115] The above process may be applied in an online (on-vehicle, on-platform) manner to improve downstream tasks in real-time or near real-time as the vehicle / platform explores new, changing, or unknown areas.
[0116] Generally, FIGS. 15-17 schematically show a computer-implemented method for training a machine learning (ML) algorithm, as will be described with respect to a second aspect, the method comprising generating training data including a set of transformed representations including a first transformed representation of an environment according to a first aspect; training an ML algorithm including classifying a set of representations transformed according to a set of classes including a first class and including.
[0117] FIG. 15 schematically shows in detail method 15 according to an exemplary embodiment. More specifically, FIG. 15 schematically shows a method for determining conditions of an image.
[0118] In particular, the condition classifier 150 is trained to detect and classify the condition or appearance of input data (i.e., an image including known conditions) ICK. The condition classifier 150 can be a neural network. The condition classifier 150 is trained to reduce or minimize the classification loss between the predicted condition PC and the actual condition AC. The current condition AC is a known condition and is the condition associated with the input data ICK.
[0119] Furthermore, condition-specific intermediate features (prediction condition features PCF) emitted or generated as part of the operation of the condition classifier 150 can be stored in the database CF DB. The term "feature" can be used in this context to mean an activation from within a neural network or an output from an activation function within a neural network. There can be multiple prediction condition features each associated with the respective output of the activation function of each node within the neural network. Thus, all activation outputs can be stored as prediction condition features PCF. The prediction condition features PCF can be stored in a database called the condition feature database CF DB.
[0120] The condition classifier 150 can optionally issue a confidence score Pr for the prediction. The prediction confidence Pr can be the probability of the output layer of the neural network that the condition of the acquired image is one of one or more unknown conditions of the image. For example, when a softmax output layer is used, the probabilities associated with each node of the output layer are taken in as the prediction confidence Pr.
[0121] In this example, an image having a known condition ICK is generated using a set including a first condition C1 (e.g., wet), a second condition C2 (e.g., snow), a third condition C3 (e.g., night), a fourth condition C4 (e.g., attached water droplets), …, an Nth condition included in the labeled data LD (actual data or synthetic data).
[0122] FIG. 16 schematically shows in detail a method 16 according to an exemplary embodiment and is generally as described with respect to method 15, but the repetition thereof is omitted for brevity. More specifically, FIG. 16 schematically shows a method for identifying new conditions.
[0123] In FIG. 16, an input image ICU including unknown conditions is acquired by one or more image sensors of an autonomous vehicle. The acquired image ICU is applied to a condition classifier 150. The conditions can be conditions such as weather conditions, time zones, brightness conditions, etc. The condition classifier 150 is configured to output a predicted condition feature PCF, a predicted condition PC, and a prediction confidence Pr based on the input image ICU.
[0124] Next, the method includes checking for similar features in a condition feature database CF DB. At 161, if the prediction confidence Pr is low, for example, below a confidence threshold, or if there are no similar features in the database, the condition is determined to be a new condition. The new condition is stored in a new condition image buffer CIB at 162.
[0125] For example, a known condition can be 100% brightness, for example, during the day, and another known condition can be 0 - 20% brightness, for example, at night. If the input image ICU is captured by a camera of an autonomous vehicle at a time zone in the evening with, for example, 50% brightness, the predicted condition feature PCF does not represent the predicted condition feature PCF for any of the known conditions. Any suitable matcher can be used to compare the condition features of the input image ICU with the condition features of the known conditions.
[0126] FIG. 17 schematically shows in detail a method 17 according to an exemplary embodiment, generally as described with respect to method 16, but the repetition is omitted for brevity.
[0127] More specifically, FIG. 17 schematically shows a method 17 for performing condition-specific downstream tasks, such as semantic segmentation, object detection, object recognition, etc.
[0128] Method 17 is the same as method 16 until it includes checking for similar features in the feature database at 160.
[0129] Next, at 171, if there are similar features in the conditional feature database CF DB and the prediction confidence Pr exceeds the threshold, then at 172, parameters are selected from the parameter feature database P DB. The parameters may be the parameters of a specific machine learning model used for a downstream task. For example, the parameters may include the weights of a neural network. The parameters are determined when training a machine learning model to perform a specific downstream task. For example, a model trained using a 100% brightness condition to perform semantic segmentation has specific weights. A method trained using a 20% brightness condition to perform semantic segmentation has different weights. Thus, there may be multiple parameter settings for a semantic segmentation model, or there may be one individual parameter setting for each condition for which the model is trained.
[0130] When a parameter setting is retrieved, a specific downstream task can be performed. For example, an image can be parameterized.
[0131] To do this, the method may further include comparing the prediction confidence with a confidence threshold and determining a similarity between one or more prediction confidence features and each of one or more confidence features of known conditions.
[0132] The above description is applicable when the new conditions are very close to, or exactly match, the known conditions with known parameter settings for downstream tasks. In this case, when the prediction confidence Pr exceeds the confidence threshold and the similarity of one or more prediction confidence features is greater than the matching threshold, the method includes the step of searching for a machine learning model from a parameter database, where the retrieved machine learning model has parameter settings obtained from training the machine learning model in an image that matches the acquired image, and the parameter setting database includes a plurality of machine learning models, each having different parameter settings derived from training the machine learning model using images containing different conditions, and further includes the step of executing the task by applying the acquired image to the retrieved machine learning model.
[0133] Similar conditions can use a similar approach. Such conditions are when the match between the features of the new condition and the features of the known condition is similar but not closely matching. For example, the difference between a first threshold and a second threshold. In this case, the parameters retrieved from the parameter database PDB at 172 can be interpolated from the closest known parameter settings. For example, the weights of the model for the closely matching parameter settings can be interpolated to generate a similar model with a new set of weights. The specific downstream task 174 can be executed using the model with the interpolated parameter settings.
[0134] In other words, in this case, when the prediction confidence exceeds the confidence threshold and when the similarity of one or more prediction confidence features is greater than the dissimilarity threshold and less than the matching threshold, the method includes the step of retrieving a machine learning model from a parameter database, where the retrieved machine learning model has parameter settings obtained from training of a machine learning model in an image containing conditions closest to the acquired image, and the parameter setting database includes a plurality of machine learning models, each having different parameter settings derived from training of a machine learning model using an image containing different conditions, the step of modifying the retrieved machine learning model by interpolating its parameter settings using the difference between the prediction condition features and the condition features of the conditions associated with the retrieved machine learning model, and the step of performing a task by applying the acquired image to the machine learning model having the interpolated parameter settings.
[0135] In any case, the method may further include controlling the autonomous vehicle to cross a path based on the result of performing the task.
[0136] Conversely, as in the method according to FIG. 16, when the prediction confidence is less than the confidence threshold and / or when the similarity of one or more prediction confidence features is less than the dissimilarity threshold, the method further includes the step of storing the retrieved image as an image containing unknown conditions and, optionally, the step of performing a minimum-risk operation by the autonomous vehicle. The minimum-risk operation can include, for example, an emergency stop or, for example, stopping at the side of the road.
[0137] As should be apparent from the above description, the task can be selected from a list including semantic segmentation, object detection, and object recognition.
[0138] As should be apparent from the above description, the image conditions can be selected from a list that includes the type of weather, the grade of the type of weather, brightness, the grade of brightness, time of day, and season. This list is not exhaustive. The conditions can likewise be characterized by features or a summary / statistics of features generated in a condition classifier.
[0139] As is apparent from the above description, FIG. 17 illustrates a method that can be summarized as a computer-implemented method for an autonomous vehicle to perform a task using a machine learning model. The computer-implemented method includes the steps of obtaining an image of the environment of the autonomous vehicle, applying the obtained image to a condition classifier, where the condition classifier is configured to generate one or more values associated with the conditions of the obtained image, determining a parameter setting of the machine learning model based on the one or more values, and performing the task by applying the input image to the machine learning model having the determined parameter setting.
[0140] The one or more values can include a predicted condition feature PCF, a predicted condition PC, and a prediction confidence Pr.
[0141] FIG. 18 schematically shows in detail method 18 according to an exemplary embodiment. More specifically, FIG. 18 schematically shows method 18 for storing new conditions in an image buffer (REMOTE). A vehicle-mounted (LOCAL) training data buffer CIBL can be wirelessly transferred by copy to a REMOTE training data buffer CIBR, for example located in a data center (180).
[0142] Figure 19 schematically shows in detail method 19 according to an exemplary embodiment. More specifically, Figure 19 schematically shows a method 19 of training an image enhancement model or an image translation model using an image containing new conditions. The image enhancement model (e.g., GAN) may be the image enhancement model 20 of Figure 2. The image translation model (e.g., cycleGAN) may be the image translation model 20 of Figure 4. Thus, training an image enhancement model or an image translation model may mean retraining each model previously trained under the closest matching conditions.
[0143] In the retraining of each model, an image containing new conditions is retrieved from a new condition image buffer CIB (LOCAL or REMOTE), and the training of each model is performed at 190. Further, the new condition image can be used to input new conditions as a style on the image. The newly trained model 20 can generate a new image, i.e., an original image containing new condition ICn.
[0144] Figure 20 schematically shows in detail method 20 according to an exemplary embodiment. More specifically, Figure 20 schematically shows a method 20 of training or retraining a downstream task, such as semantic segmentation. As described with reference to Figure 17, when the predicted features of a condition classifier processing an image containing unknown conditions match the features in the condition feature database, individual parameter settings of a downstream task model, such as a semantic segmentation model, can be selected from the parameter database. When the condition features are close to the features from the condition feature database CF DB, individual parameter settings from the parameter database P DB can be selected and interpolated accordingly. However, in situations where the features are dissimilar, e.g., outside the second threshold described above, the downstream task is not executed. In such cases, the downstream task model needs to be retrained with new parameter settings for the new conditions.
[0145] The downstream task model takes the parameter settings of a previously trained downstream task model and retrains it using the original images containing the new condition ICn, and can be retrained by reducing the loss between the predicted semantic map SMP and the known semantic map SMOD for the original images. This is possible because the ground truth is the same.
[0146] The method 20 of FIG. 20 can be summarized as a computer-implemented method for training a machine learning model of an autonomous vehicle that performs a task using an input image, the computer-implemented method including the steps of obtaining a plurality of images including unknown conditions, generating a predicted semantic map by applying the plurality of obtained images including unknown conditions to the machine learning model, optimizing the parameters of the machine learning model by minimizing the error between the predicted semantic map and the semantic map ground truth to generate parameter settings of the machine learning model for unknown conditions, and storing the generated parameter settings of the machine learning model in a parameter database, the parameter database being configured to store a plurality of machine learning models each having different parameter settings, and each parameter setting being associated with a specific condition.
[0147] Furthermore, the step of generating a predicted semantic map by applying a plurality of obtained images including unknown conditions to the machine learning model may include generating a predicted semantic map SMP by applying a plurality of obtained images ICn including unknown (or new) conditions to a machine learning model previously trained using images including conditions different from the unknown conditions.
[0148] As described above, the unknown conditions and the specific conditions are each selected from a list including the type of weather, the grade of the type of weather, brightness, the grade of brightness, time of day, and season. The term "grade" can be used to define the amount of a particular condition. For example, the grade of brightness can be set to 0% for a completely dark state, such as inside a tunnel at night without artificial lighting, and can be set to 100% for a completely bright state, such as during the day. The brightness can still be clear, but the evening time when it has decreased compared to early morning can be set to a brightness grade of 50%.
[0149] Furthermore, the task can be selected from a list including semantic segmentation, object detection, and object recognition.
[0150] Figure 21 schematically shows in detail method 21 according to an exemplary embodiment. Furthermore, and very importantly, the newly created data, together with the original ground truth (segmentation map, object bounding box, depth, etc.), can be used to check the performance of the existing task 210 by using the prediction performance as a proxy for the reliability 211 (global, local, instance, or pixel-wise). This represents an important aspect of continuous and long-term verification and validation, and is not only particularly useful, but also extremely important for the effective deployment of autonomous platforms in both existing, continuously changing / evolving, and new areas.
[0151] The processes shown in FIGS. 15, 16, 17, and 21 can be mainly performed on a vehicle.
[0152] The process shown in FIG. 18 can wirelessly transfer the training data buffer to a data center.
[0153] The processes shown in FIGS. 19 and 20 may be performed at a data center.
[0154] Alternatively, all the processing may be performed entirely on the vehicle or entirely at the data center.
[0155] When executed by one or more processors, this method can be embodied as a non-transitory or transitory computer-readable medium storing instructions that cause the one or more processors to execute the aforementioned computer-implemented method. Additionally, an autonomous vehicle including a storage, one or more processors, one or more image sensors, and one or more actuators is provided herein, and the storage includes a non-transitory or transitory computer-readable medium.
[0156] All processes can be performed continuously (where all parts of the data, including new conditions, are immediately used in the training process) or discretely (where the data is clustered based on predicted conditions or predicted condition features and used in training when a certain amount has accumulated) in real-time or near real-time.
[0157] Preferred embodiments have been shown and described, but it will be understood by those skilled in the art that various changes and modifications can be made without departing from the scope of the invention as defined in the appended claims, as described above.
[0158] At least some of the exemplary embodiments described herein can be constructed, in part or whole, using dedicated special-purpose hardware. Terms such as "component", "module", or "unit" as used herein can include, but are not limited to, hardware devices such as circuits in the form of individual or integrated components, field programmable gate arrays (FPGAs), or application specific integrated circuits (ASICs), which perform specific tasks or provide related functions. In some embodiments, the elements described can be configured to exist on a tangible, persistent, addressable storage medium and can be configured to execute on one or more processors. These functional elements can, in some embodiments, include, by way of example, components such as software components, object-oriented software components, class components and task components, processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. Exemplary embodiments are described with reference to the components, modules, and units discussed herein, but such functional elements can be combined into fewer elements or separated into additional elements. It will be understood that various combinations of any features are described herein and the described features can be combined in any suitable combination. In particular, the features of any one exemplary embodiment can be combined with the features of any other embodiment, as necessary, except where such combinations are mutually exclusive. Throughout this specification, the terms "including" or "comprising" mean including the specified component, but not excluding the presence of other components.
[0159] Attention is directed to all papers and documents filed in connection with this application and published in general with this specification, simultaneously or before, and the contents of all such papers and documents are hereby incorporated by reference into this specification.
[0160] All features disclosed in this specification (including any of the appended claims, abstract and drawings), and / or all steps of any method or process so disclosed, may be combined in any combination, except combinations where at least some of such features and / or steps are mutually exclusive.
[0161] Each feature disclosed in this specification (including any of the appended claims, abstract and drawings) may be replaced by an alternative feature serving the same, equivalent or similar purpose, unless expressly stated otherwise. Thus, each feature disclosed is only an example of a general series of equivalent or similar features, unless expressly stated otherwise.
[0162] The present invention is not limited to the details of the foregoing embodiments. The present invention extends to any novel one, or any novel combination, of the features disclosed in this specification (including any of the appended claims, abstract and drawings), or any novel one, or any novel combination, of the steps of any method or process so disclosed.
[0163] The subject matter can be understood with reference to the following clauses.
[0164] (Clause 1) A computer-implemented method of generating training data, the method comprising: providing a representation of an environment, the representation of the environment having a defined structure and / or a defined shape; generating training data including a set of transformed representations including a first transformed representation of the environment by transforming the representation of the environment into a set of transformed representations including the first transformed representation of the environment. comprising: A computer-implemented method, wherein the step of providing a representation of the environment includes at least partially synthesizing an image of the environment using semantic information. (Clause 2) The method according to clause 1, wherein the step of providing a representation of the environment includes at least partially obtaining an image of the environment. (Clause 3) The method according to clause 1 or 2, wherein at least partially synthesizing an image of the environment using semantic information includes obtaining an image or a part thereof corresponding to the semantic information from a database or by learning. (Clause 4) The method according to any one of clauses 1 to 3, wherein the step of providing a representation of the environment includes at least partially semantically constructing an image of the environment. (Clause 5) The method according to any one of clauses 1 to 4, wherein the step of providing a representation of the environment includes inpainting an image of the environment. (Clause 6) The step of generating training data by converting a representation of the environment into a set of transformed representations including a first transformed representation of the environment, using each set of trained transformations including a first trained transformation, to convert the representation of the environment into a set of transformed representations including a first transformed representation of the environment, the method according to any one of clauses 1 to 5. (Clause 7) The step of generating training data by converting a representation of the environment into a set of transformed representations including a first transformed representation of the environment, using each set of heuristics-based transformations including a first heuristics-based transformation, to convert the representation of the environment into a set of transformed representations including a first transformed representation of the environment, the method according to any one of clauses 1 to 6. (Clause 8) The step of generating training data by converting the representation of the environment into a set of transformed representations including a first transformed representation of the environment includes generating training data by using each set of augmentations including a first augmentation to convert the representation of the environment into a set of transformed representations including a first transformed representation of the environment, the method according to any one of clauses 1 to 7. (Clause 9) The step of generating training data by converting the representation of the environment into a set of transformed representations including a first transformed representation of the environment includes generating training data by using each set of conditions including a first condition to convert the representation of the environment into a set of transformed representations including a first transformed representation of the environment, the method according to any one of clauses 1 to 8. (Clause 10) The step of generating training data by converting the representation of the environment into a set of transformed representations including a first transformed representation of the environment includes mixing a set of representations of the environment, the method according to any one of clauses 1 to 9. (Clause 11) The first transformed representation of the environment has a defined structure and / or a defined shape of the representation of the environment, the method according to any one of clauses 1 to 10. (Clause 12) The step of generating training data by converting the representation of the environment into a set of transformed representations including a first transformed representation of the environment includes redefining a defined structure of the representation of the environment into a redefined structure of the first transformed representation of the environment, the method according to any one of clauses 1 to 10. (Clause 13) A computer-implemented method for training a machine learning (ML) algorithm, the method comprising generating training data including a set of transformed representations including a first transformed representation of the environment according to any of the previous clauses; and A step of training an ML algorithm, including a step of classifying a set of transformed representations according to a set of classes including a first class A computer-implemented method including (Clause 14) The method according to clause 13, wherein the set of classes including the first class is a set of conditions including a first condition (Clause 15) The method according to clause 14, including a step of identifying a set of characteristic features including a first characteristic feature associated with the first condition (Clause 16) A computer-implemented method for determining a class of a representation of an environment using a machine learning (ML) algorithm trained according to any of clauses 13 to 15, the method comprising A step of determining a class of a representation of an environment, including inferring a class of a representation of the environment using the trained ML algorithm A computer-implemented method including (Clause 17) The method according to clause 16, including a step of identifying a set of features including a first feature associated with a condition of a representation of an environment (Clause 18) The method according to clause 17, including a step of calculating a confidence score for the inferred class (Clause 19) The method according to clause 17 or 18, including a step of comparing the identified set of features with a set of characteristic features associated with a condition of a representation of an environment (Clause 20) The method according to clause 19, including a step of storing a representation of an environment based on the result of the comparison (Clause 21) The method according to clause 20, including a step of training a transformation using the stored representation of the environment (Clause 22) The method according to clause 21, including a step of generating training data using the trained transformation (Clause 23) The method according to clause 22, comprising the step of training an ML algorithm using the generated training data. (Clause 24) The method according to clause 22, comprising the step of validating an ML algorithm using the generated training data. (Clause 25) The method according to clause 19, comprising the step of performing an operation based on the result of the comparison.
Claims
1. A computer-implemented method for an autonomous vehicle to perform a task selected from a list including semantic segmentation, object detection, and object recognition using a machine learning model, the computer-implemented method comprising: obtaining an image of the environment of the autonomous vehicle; applying the obtained image to a condition classifier, the condition classifier being configured to generate one or more values associated with the conditions of the obtained image; determining a parameter setting of the machine learning model based on the one or more values; executing the task by applying the obtained image to the machine learning model using the determined parameter setting; comprising: the condition classifier includes a neural network, the one or more values include one or more prediction condition features and a prediction confidence level; the computer-implemented method further comprising: comparing the prediction confidence level with a confidence threshold; determining a similarity between the one or more prediction condition features and each of one or more condition features of known conditions; further comprising: when the prediction confidence level exceeds the confidence threshold and the similarity of the one or more prediction condition features is greater than a dissimilarity threshold and less than a matching threshold, the method comprising: searching for a machine learning model from a parameter database, the searched machine learning model having a parameter setting obtained from training of the machine learning model in an image including conditions closest to the obtained image, the parameter database including a plurality of machine learning models each having a different parameter setting derived from training of the machine learning model using images including different conditions; Modifying the retrieved machine learning model by interpolating its parameter settings using the difference between the predicted condition features and the condition features of the conditions associated with the retrieved machine learning model; Executing the task by applying the obtained image to the machine learning model having the interpolated parameter settings; A computer-implemented method further comprising.
2. The computer-implemented method according to claim 1, wherein the predicted condition feature or each predicted condition feature includes an output of an activation function.
3. The computer-implemented method according to claim 1, wherein the prediction confidence includes a probability from an output layer of the neural network that the condition of the obtained image is one of one or more known conditions of the image.
4. When the prediction confidence exceeds the confidence threshold and when the similarity of the one or more predicted condition features is greater than a matching threshold, the method comprises: Searching a parameter database for a machine learning model, wherein the retrieved machine learning model has parameter settings obtained from training of the machine learning model in an image including a condition that matches the obtained image, and the parameter database includes a plurality of machine learning models each having different parameter settings derived from training of the machine learning model using images including different conditions; Executing the task by applying the obtained image to the retrieved machine learning model; The computer-implemented method according to claim 1, further comprising.
5. When the prediction confidence is less than the confidence threshold and / or when the similarity of the one or more predicted condition features is less than a dissimilarity threshold, the method comprises: Storing the retrieved image as an image including an unknown condition; The computer-implemented method according to claim 1, further comprising
6. The computer-implemented method according to claim 1, further comprising controlling the autonomous vehicle to cross a path based on a result of executing the task.
7. The computer-implemented method according to claim 1, wherein the condition is selected from a list including a type of weather, a grade of the type of weather, brightness, a grade of brightness, a time zone, and a season.
8. A transient or non-transient computer-readable medium having stored instructions that, when executed by one or more processors, cause the one or more processors to execute the computer-implemented method according to claim 1.
9. An autonomous vehicle including a storage, one or more processors, one or more image sensors, and one or more actuators, wherein the storage includes the transient or non-transient computer-readable medium according to claim 8.
Citation Information
Patent Citations
Discrimination method, discrimination device, discriminator generation method and discriminator generation device
JP2018081404A
Scene classification
JP2020047267A
Image processing system and image processing method
JP2020126432A