Method and apparatus for testing a machine
By selecting a subset of sample trajectories and utilizing similarity metrics and an autoencoder framework, the problem of time-consuming machine testing is solved, resulting in an efficient and reliable testing method suitable for both machine simulation and real-world testing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-24
- Publication Date
- 2026-03-10
AI Technical Summary
Machine testing is time-consuming and expensive during the development cycle, especially when all relevant parameters need to be tested.
By providing a set of sample trajectories, selecting a subset of them, and utilizing similarity metrics and an autoencoder framework, the machine's output can be determined, and simulated or real-world testing can be used to determine test results, reducing testing costs and time.
This approach reduces testing costs and time while improving testing reliability and efficiency, and effectively identifies relevant sample trajectories to ensure reliable machine operation.
Smart Images

Figure CN113051808B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present invention relates to a method and apparatus for testing a machine. BACKGROUND
[0002] In the development cycle of a machine, such as a robot or a vehicle, machine testing requires a lot of time. A vehicle or a robot can be tested in the real world or by simulation. However, testing all relevant parameters in all relevant cases is both cumbersome and expensive.
[0003] Therefore, a practical method for reliable testing is desirable. SUMMARY
[0004] The method and apparatus for testing a machine produce highly reliable results at a reasonable cost.
[0005] The method for testing a machine comprises providing a set of sample trajectories, selecting a subset of the set, determining an output of the machine for a movement of the machine in accordance with a sample trajectory of the subset, determining a result of the test in dependence of the output, wherein the subset is selected in dependence of a score characterizing a difference between the set and the subset, wherein the score is determined in dependence of a similarity measure characterizing a difference between a first sample trajectory of the set and a second sample trajectory of the set. The test can comprise determining an emission distribution by generating an output for each trajectory of the subset of sample trajectories using a simulation model for the machine, or by moving the machine over the sample trajectories and measuring the output that occurs. The set of sample trajectories can be a base set of a plurality of trajectories containing a real distribution of real world sample trajectories of the machine. Testing the machine using all trajectories of the real distribution of real world sample trajectories is very expensive or not possible. While testing the machine only on the base set of a plurality of trajectories containing the real distribution is less expensive, testing only selected sample trajectories of the subset is more cost efficient and less time consuming. The score is used to select from the set only a subset of sample trajectories that will result in an output distribution that is similar to the output distribution achieved by testing the trajectories of the base set. It is assumed here that the distribution of the output for the trajectories of the base set approximates the output produced by the trajectories of the real distribution of real world trajectories. In order to effectively select the subset, the similarity measure is used to identify sample trajectories that are relevant for the test.
[0006] In one aspect, the output characterizes emissions of an internal combustion engine operable to move a machine, wherein the result indicates a test failure when the output value is greater than a threshold value, and wherein the result indicates a test success otherwise. The test can be based on determining an emissions profile of the machine, either by generating emissions as output of a simulation model of the machine for each of a subset of sample trajectories, or by moving the machine over sample trajectories and measuring emissions that occur. The machine can be moved over real-world sample trajectories. Simulating sample trajectories is more cost-efficient and less time-consuming.
[0007] In one aspect, the output characterizes emissions of an internal combustion engine operable to move a machine, wherein an output statistic, in particular an expectation value or a 99% quantile, is determined depending on a distribution of the output values, wherein the result indicates a test failure when the output statistic is greater than a threshold value, and wherein the result indicates a test success otherwise. The test can also be based on the entire distribution of the output values.
[0008] In one aspect, the output characterizes emissions of an internal combustion engine operable to move a machine, wherein a first distribution of output values is determined for a first quantification of the machine or a simulated environment, and wherein in a comparison, the first distribution of output values is compared to a second distribution of output values characterizing a second quantification of the machine or the simulated environment, and wherein depending on a result of the comparison, either the first quantification is used for the test or the second distribution is used for the test. The comparison identifies which of the two quantifications of the machine or the simulated environment results in a better output statistic, e.g. a better expectation value or a better 99% quantile, over the two simulated or real-world output distributions. The result of the comparison is significantly improved by taking into account the entire output distribution, e.g. by estimating the entire output distribution by representative samples.
[0009] Preferably, at least one sample trajectory of the set of sample trajectories is defined by a sequence of route features. Route features can be used to determine relevant trajectories when the machine is a vehicle moving over a route.
[0010] Preferably, the route characteristics characterize a geographical profile, in particular an absolute height or road slope profile; and a traffic flow characteristic, in particular a time-dependent average traffic speed; and a road characteristic, in particular a number of lanes, a road type and / or a road curvature; and a traffic control characteristic, in particular and / or a speed limit characteristic, a number of traffic lights, a number of traffic signs of a specific type, a number of stop signs, a number of yield signs and / or a number of pedestrian crossing signs; and a weather characteristic, in particular a rainfall at a predetermined time, a wind speed, a fog, a route distance, a time of day or week while driving the route, and / or a vehicle type for driving the route. Different vehicle types have different driving needs. For example, a vehicle type distinguishes between a car in a carpool and a delivery truck; between a light vehicle and a heavy vehicle. These characteristics are particularly relevant for the emissions produced by the vehicle.
[0011] The sample trajectory can be determined depending on a sequence of route characteristics, wherein the sequence of route characteristics is defined along a route, wherein the route is defined by a pair of coordinate sequences, in particular a pair of coordinate sequences characterizing a latitude and a longitude at one step, and a start time. This way, the sample trajectory can be provided in a pre-processing step from data describing the route.
[0012] Preferably, the output value is determined by a simulation of the machine depending on the sample trajectories in the subset. This is a more cost-efficient and less time-consuming way of determining the output value of the test.
[0013] The method comprises: providing a set of sample trajectories representing an unbiased sample of a true distribution of real-world operations of the machine, in particular the true distribution being conditioned on a set of machine classes characterizing different variants of the machine, a set of regions of interest characterizing the sample trajectories in a geographical space, and a set of time intervals characterizing the set of sample trajectories in time, wherein a first sample trajectory is sampled from the set of sample trajectories, wherein a second sample trajectory is sampled from the set of sample trajectories, wherein a first feature vector having a dimension is determined from the first sample trajectory, and / or wherein a second feature vector having the same dimension is determined from the second sample trajectory, and wherein a similarity measure is determined for the first feature vector and the second feature vector. The sample trajectories can have the same or different lengths. Determining vectors having the same dimension further improves efficiency.
[0014] Preferably, the score is determined as an approximation, in particular a random approximation, of pairwise similarities between samples in the set of sample trajectories and samples in the subset, and pairwise similarities between samples in the subset, wherein the pairwise similarity of one pair is determined depending on the similarity measure.
[0015] Preferably, the similarity measure is determined depending on a trajectory kernel, in particular a Triangular Global Alignment kernel.
[0016] In training, a sample trajectory is provided as an input sequence to an autoencoder comprising a first recurrent neural network, RNN, wherein the first RNN is configured as an encoder, wherein a second RNN is configured as a decoder and receives, in particular, a feature vector of a fixed dimension as an initial hidden state, which is extracted as an output and / or hidden state of the last time step of the first RNN, and wherein the second RNN determines an output sequence, wherein at least one parameter of the first RNN and / or the second RNN is determined, which reduces a reconstruction loss between the input sequence and the output sequence, and wherein the first RNN of the thus trained autoencoder is used for testing or for determining a subset of the test, in particular without the second RNN. In this way, the autoencoder is trained and then the encoder of the autoencoder can be used to extract a first feature vector from a first sample trajectory and / or a second feature vector from a second sample trajectory. The thus trained neural network can be used to efficiently implement the test. BRIEF DESCRIPTION OF DRAWINGS
[0017] Further embodiments can be derived from the following description and the attached drawings. In the drawings:
[0018] Figure 1 A device for testing a machine is schematically depicted,
[0019] Figure 2 Steps in a method for testing a machine are depicted,
[0020] Figure 3 An output distribution is depicted,
[0021] Figure 4 Aspects of input and output distributions are depicted,
[0022] Figure 5 Alternative aspects of input and output distributions are depicted,
[0023] Figure 6 Alternative aspects of input and output distributions are depicted,
[0024] Figure 7 Steps in a method for training are depicted. DETAILED DESCRIPTION
[0025] Figure 1 A device 100 for testing a machine 102 is schematically depicted.
[0026] Device 100 is adapted to perform the methods described below. Device 100 may include processor 104 or multiple processors. Device 100 may include input 106 for input data 108 and output 110 for output 112. Input data 108 for device 100 may be sampled from storage device 116. Output 112 may indicate whether a test on machine 102 has passed or failed. Output 112 may include instructions 114 for operating machine 102 or a simulation environment for testing.
[0027] From the set of sampling trajectories X The sampled input data is 108, as described below. Sample trajectory. X It can be part of a data set It is stored in storage device 116. The data set in the example. Includes: real-world data that can be captured from machine 102 or a similar machine 102. D base For example, the entire dataset, such as all routes traveled in Germany over a year using all vehicle categories. Specifically, X It can be a subset of the data set ,or X It can be the entire dataset In one example, for instance, in In cases where the size is too large to handle, relative to To filter this dataset, or from A small set X is randomly sampled from the data. The real-world data in this example is captured from the real-world operations of machine 102. The output values of machine 102 can be... y s The distribution is analyzed using distribution analysis. Output values y s For this purpose, it can be stored in database 118.
[0028] Machine category C can be defined for machine 102, for example:
[0029] C base C
[0030] Lightweight+
[0031] Heavy-
[0032] Among them, column C base The available categories are indicated, and column C uses a + sign to indicate the category as... C base Applicable to machine 102, and the category is indicated by –. C base Not applicable to machine 102.
[0033] Area of Interest may be defined in a geographical space R base Area of Interest R characterizes a sample trajectory X .
[0034] set of time intervals may be defined as the time interval between a shortest time and a longest time on which I base a sample trajectory X is characterized. In an example, a start time and a stop time define a set of time intervals I . The set of time intervals can be all rush hours within a certain time range.
[0035] In an example, the machine 102 is a vehicle, and the output y s characterizes the emissions of an internal combustion engine operable to move the vehicle along an arbitrary trajectory. In this example, the sample trajectory X may influence the actuation of the engine, and thus the emissions of the engine.
[0036] The goal of the output analysis is for example to test whether a legal emission threshold is met. For example, the test fails when the output value y s is larger than the threshold, and otherwise the test succeeds. In case the output is y s a set, then it is possible to only analyze the entire distribution by summarizing the simulated outputs according to a selected subset. In this regard, one goal is to compare two different parameterizations of the machine 102 at a distribution level, and to select the system parameterization which is better in the sense that the expected value (or some other distribution statistic) is better, e.g. smaller expected emissions in the real world.
[0037] geographical space R base may be defined by map data, e.g. by longitude and latitude coordinates. In this example, the Area of Interest R characterizes a set of geographical areas (e.g. map areas) defining sample trajectories X .
[0038] In another example, a robot can move along a trajectory by an engine. In this case, the output y s may characterize the energy consumption of the engine, and the trajectory can describe a path through a geographical spaceR base Defined and constrained by the area of interest that limits robot movement R A sequence of poses in an environment.
[0039] In the following description, the set of sample trajectories X Sample trajectory x 1:T Based on route characteristics sequence Definition. Route characteristics. Characterized by, for example, geographical features, particularly absolute elevation or road gradient; and traffic flow features, particularly time-related average traffic speed; and road features, particularly the number of lanes, road type and / or road curvature; and traffic control features, particularly and / or speed limit features, the number of traffic lights, the number of specific types of traffic signs, the number of stop signs, the number of yield signs and / or the number of pedestrian crossing signs; and weather features, particularly rainfall at a predetermined time, wind speed, fog, route distance, time of day or week on the route, and / or vehicle type on the route.
[0040] In this example, route features sequence It is along the route Defined. Route It consists of a pair of coordinates l j sequence Specifically, the pair of coordinate sequences representing the latitude and longitude at step j, and the start time. Defined.
[0041] Dataset The unbiased set of samples representing the observed true distribution A set of machine categories representing different variants of machine 102. To characterize geographic space R base Sample trajectory X Area of Concern And in terms of time I base Characterizing sample trajectory X time interval set To define a data set .
[0042] This method is based on a three-step process, including: creating an experimental design to determine the input data; measuring the output from machine 102 or its simulation therein; processing the input data and output measurements to learn the input-output relationship between the input data determined by the experimental design and the corresponding measured output; and, instead of learning the input-output relationship, using analysis of the output distribution.
[0043] The following reference Figure 2 This describes the computer-implemented method used for testing machine 102.
[0044] The method includes step 202, which provides a set of sample trajectories. X .
[0045] For example, depending on route characteristics sequence To represent or describe sample trajectories x 1:T .
[0046] In the example, route features sequence It is along the route Defined.
[0047] route It consists of a pair of coordinates l j sequence and start time Defined. Certain route features (e.g., traffic speed) depend on time, and in order to extract those route features, the time must be at least approximately known when passing the corresponding coordinates.
[0048] In the example, coordinates l j It is defined by the latitude and longitude at step j. This identifies the orientation of machine 102 along the trajectory.
[0049] In the example, along the route route characteristics sequence length T Possibly compared to the coordinates at step j l j sequence length Longer.
[0050] The method includes step 204, which selects a set. X subset of X s .
[0051] subset is dependent on a score MMD selected, the score characterizing a difference between the set X of first sample trajectories X s and the set of second sample trajectories
[0052] MMD is dependent on a similarity measure determined, the similarity measure characterizing a difference between the set X of first sample trajectories x 1:T and the set X of second sample trajectories x’ 1:T .
[0053] In an example, the first sample trajectory of length T x 1:T is sampled from the set of sample trajectories and a first feature vector of dimensionality is determined from the first sample trajectory of length T x 1:T .
[0054] In this example, the second sample trajectory of length T' T’ is sampled from the set of sample trajectories and a second feature vector of dimensionality is determined from the second sample trajectory of length T' (T' and T can indicate the same or different lengths).
[0055] The first feature vector may be determined as an output of a recurrent neural network , in particular trained using an autoencoder framework to extract the first feature vector x 1:T from the first sample trajectory . The autoencoder framework is as follows: the trajectory is an input sequence to a recurrent neural network RNN. The output and / or hidden state of the last time step of the RNN is the extracted feature vector. This feature vector has a fixed dimensionality. A second RNN takes this output as a first hidden state and now tries to reconstruct the input sequence in turn. The first RNN is the encoder g(.), the second RNN is the decoder. The decoder is not used after training. The autoencoder is trained to minimize a reconstruction loss between the input sequence and the output sequence.
[0056] The second feature vector The last time step of the recurrent neural network g(. ) can be determined as output and / or hidden state, in particular using an autoencoder framework to train the recurrent neural network to extract a second feature vector from a second sample trajectory .
[0057] This step can be obsolete when sample trajectories of the same dimension are processed and the feature vector is defined as the sample trajectory. Another alternative is to define a fixed number of features to extract from the sequence, e.g. mean value, zero-crossing number, histogram, etc. These “expert features” can be used in the same way as the extracted learned features.
[0058] Using and , a similarity measure ) is determined for the first feature vector and the second feature vector .
[0059] The similarity measure may be determined as the output of a kernel function (e.g. a radial basis function kernel ) or a (triangular) global alignment kernel .
[0060] The similarity measure is determined, for example, as:
[0061] ,
[0062] where RBF is a radial basis function kernel, e.g.
[0063]
[0064] where is a free parameter.
[0065] In an example, the similarity measure is for different and for the same . I.e. if according to the similarity measure the two routes represented by are identical, both routes result in the same emission or emission profile.
[0066] In an example, the score MMD is evaluated and minimized by maximizing the average difference:
[0067]
[0068] With which and where Indicates the similarity between two sample trajectories, e.g. by a similarity measure The two sample trajectories are determined.
[0069] In an alternative objective function, the 1 / m2 factor can be replaced by 1, while the 2 / nm factor can be replaced by a parameter λ. λ can be set by the user or picked on a validation set. The advantage of this is that it is possible to weight coverage versus diversity with λ. The disadvantage of this is that the parameter can be difficult to tune. Therefore, MMD is preferred over the parametric version using λ.
[0070] Maximum Mean Discrepancy MMD is a measure of the difference between the base dataset and the selected subset X of this dataset X s The goal is to minimize this difference with respect to the subset X s .
[0071] The MMD or square MMD distance between two datasets can be written in terms of kernel similarities between individual elements within a dataset and across datasets. The first term in the following expression is constant in the optimization problem at hand, since the optimization is over the subset X s , but the first term is only over the base dataset.
[0072] Therefore, this term is ignored for the optimization, and the optimization problem is formulated as a maximization.
[0073] To find the subset with cardinality X s , where , in the example is minimized with respect to X s .
[0074]
[0075] Given that is constant in X s , the estimate X s of the subset is determined as
[0076]
[0077] where .
[0078] In this example, if then the subset X s represents q% of the set X and also represents q% of the distribution y s of the output
[0079] where the empty set represents 0% of the set X
[0080] To determine a solution to the optimization problem, an optimization algorithm can be used. The optimization problem above is NP-hard because it is a selection problem with cardinality constraints. In an exemplary selection problem for finding a solution, if a data point is selected for the subset X s then an indicator variable for each data point in the base set X is introduced to be 1, while if a data point is not selected for the subset X s then the indicator variable is introduced to be 0. This selection problem can be solved to global optimality using integer linear program solvers, but this can take exponential time due to NP-hardness.
[0081] Instead, approximate optimization is more efficient.
[0082] Approximate optimization alternative 1:
[0083] The relaxation of the selection problem is easier to solve if the indicator variables are not binary but continuous values between 0 and 1. The resulting quadratic optimization problem can then be solved efficiently.
[0084] This relaxation has been applied to MMD minimization, for example, in [Wang and Ye, 2013, Section 3.4, http: / / chbrown.github.io / kdd-2013-usb / kdd / p158.pdf]. To enforce the cardinality constraint in the optimization problem, m samples are picked for the subset with the highest indicator variable score.
[0085] Approximate optimization alternative 2:
[0086] The authors in [Kim et al., 2016, Section 3, https: / / people.csail.mit.edu / beenkim / papers / KIM2016NIPS_MMD.pdf] solve the same optimization problem for different applications. The problem is optimized using a greedy selection through the submodularity properties of the objective under the conditions of Lemma 3 disclosed therein. Even if the conditions are not fully met, near-submodular quantities of the objective allow the use of a greedy optimization scheme with less guarantees with respect to reaching the true optimal value. This yields the following optimization scheme.
[0087] Maximize a monotone submodular function with cardinality constraint by the following greedy algorithm is (approximately) 0.63 optimal.
[0088] Greedy optimization algorithm:
[0089]
[0090] It can be shown that if the parametrization of the similarity measure is chosen such that for all there are and then is a monotone submodular function.
[0091] Alternative 2 is preferred because it can give optimization guarantees under certain conditions and the optimization scheme is fast due to the greedy optimization.
[0092] The optimization problem includes the following sum over the entire base set X
[0093]
[0094] This term takes linear time for each candidate sample of the subset X s and thus quadratic time with respect to the size of the base set X .
[0095] The score evaluation includes the double sum over the entire base set MMD X
[0096] .
[0097] This computation takes quadratic time with respect to the size of the base set X .
[0098] To speed up these computations, the following formula for a randomized approximation can be applied as follows:
[0099] .
[0100] Computational term Pairwise similarities between all samples in the set are needed. To approximate this, if X + then a sparse indicator matrix can be defined and only the similarity measure .
[0101] This means that is approximated by
[0102] .
[0103] This term can be computed offline, i.e. outside the method steps. The term is computed in the method, i.e. online. It is worth noting that the term can be computed quickly because it only involves a relatively small subset X s of terms. And since X s increases incrementally, there is no need to store the pairwise similarities .
[0104] Instead, the similarity measure can be determined in step 204 depending on a trajectory kernel. An exemplary triangular global alignment kernel is described in Cuturi, ICML 2011 (http: / / www.icml-2011.org / papers / 489_icmlpaper.pdf), where .
[0105] A trajectory kernel tries to find the minimum cost alignment between two sequences. In the example, the two sequences are trajectories represented by time series and The kernel can determine the minimum cost to transform one time series into the other. The kernel depends on this cost and yields 1 if the two trajectories are identical.
[0106] .
[0107] where is all allowed curved paths that satisfy the monotonicity constraint, is the curved path that maps the i-th time step to the aligned time step in the curved trajectory, and where, is an element-wise kernel (e.g. RBF kernel) or is a local similarity function that measures the similarity between the feature vector of the i-th time step of the first curved trajectory and the feature vector of the i-th time step in the second curved trajectory.
[0108] The method comprises a step 206 of determining the output of the machine 102 for a movement of the machine 102 from a subset of sample trajectories X s . x s The output value y s .
[0109] The output value y s is determined in step 206 from a subset of sample trajectories X s . x s The sample trajectories x s are describing representative routes for a distribution of real-world traffic demand.
[0110] The method comprises a step 208 of determining the result of the test from the output y s .
[0111] In examples, the output y s characterizes emissions of an internal combustion engine operable to move the machine 102.
[0112] The result indicates a failure of the test when the output value y s is greater than a threshold value, and wherein the result indicates a success of the test otherwise. In case the output is y s a set, then the entire distribution can be analyzed by summarizing the outputs of the simulations from the selected subset only.
[0113] Additionally or alternatively, a distribution statistic, e.g. an expected value or a 99% quantile, can be determined in step 208 from the distribution of the output values y s . In this case, the result indicates a failure of the test when the distribution statistic, e.g. the expected value, is greater than a threshold value, and the result indicates a success of the test otherwise.
[0114] In one aspect, in step 206, a first distribution of the output values y s is determined for a first parameterization of the machine 102 or the simulated environment.
[0115] In this respect, in step 208, in the comparison, the output values y s of the first distribution are compared with a second distribution representing the second parametrization of the output values of the machine 102 or of the simulated environment.
[0116] Depending on the result of the comparison, the first parametrization is or is not used for testing in the subsequent tests (e.g. in step 206).
[0117] It is assumed that all routes occurring during real-world driving are independent random samples from a true underlying route distribution, and that they are a large sample set, i.e. a set D base has been collected from this distribution as an underlying sample set.
[0118] To estimate the emissions distribution (i.e. the target distribution), a simulation model of the vehicle is used to generate the emissions of routes from the input distribution or the entire underlying sample set. Alternatively, instead of simulating the emissions, it is also possible to drive routes in the real world to measure the occurring emissions, but this is more time-consuming and costlier. Therefore, simulation is preferred.
[0119] With reference to Figure 3 a distribution analysis is described. Figure 3 An exemplary emissions distribution 302 is depicted, in the example, an emissions density of a vehicle resulting from a simulation based on an input distribution. Also depicted is an emissions limit 304, which can be, for example, a legal limit or an engineering target. The emissions limit 304 is a threshold of the test. In the example of at least one simulation cycle 306, the output y s exceeds the threshold. This means that the test fails. However, when estimating the emissions distribution only through some scenarios and / or in certain extreme cases, as depicted in Figure 3 the resulting distribution 308 can remain completely within the legal limit 304 of the emissions. A test based on the distribution 308 would not fail.
[0120] The distribution 308 and the distribution 302 show distributions of simulated emissions with two different parametrizations of the machine 102 (e.g. a vehicle). The expected value of the distribution 308 is higher than the expected value of the distribution 302, although some simulated routes of the distribution 302 produce emissions higher than the maximum emissions 304 of the distribution 308. In this case, we pick the system parametrization of the distribution 302 because it reduces the impact of the emissions on the environment when considering the entire distribution. This comparison enables to pick the distribution that is best suited to the needs.
[0121] The goal of the experimental design is to select routes that should be simulated for vehicle emissions such that, given an infinite sample of routes from the input distribution, the emission distribution on the selected routes has minimal deviation from the true emission distribution.
[0122] Reference Figure 4 In this case, the true input distribution and real-world emission distribution 404 for route 400 are unknown. The true input distribution for route 400 can be conditioned specifically on region R, vehicle class set C, and time interval set I. For a sample trajectory x (i.e., in this case, route), the emission distribution 406 that occurs when driving the sample trajectory an infinite number of times is unknown. In Figure 4 , an example depicts the emission distribution y conditioned on route x_0, x_1,..., x_99. For a set , it is assumed that it has no sample bias and that it itself represents the input distribution for route 400. The set is depicted as a histogram 402 under the distribution for route 400. Thus, whether the selected route represents the sample set or the true input distribution, it is treated equivalently. In this example, we pick . Since the sample set is assumed to be too large to be completely measured or simulated, the goal is to find a smaller subset, i.e., a subset that has the same target variable statistics, i.e., a representative subset that increases the efficiency of the measurement process. Figure 4 The true unknown input distribution is depicted in Figure 4 and labeled 408.
[0123] As described above, this is implemented based on MMD. In the example, a representative route set 410 is depicted for MMD = 0.07.
[0124] Example routes x are selected as representative routes 410, which are depicted as beams. As Figure 4 depicted in Figure 4 , there can be more or fewer routes selected as representative routes 410. As described above, these representative routes 410 determine the resulting approximate emission distribution 412. In , the resulting approximate emission distribution 412 is depicted next to the true real-world emission distribution 404.
[0125] The cloud 414 symbolizes in Figure 4 that the emissions for each selected representative route 410 are unknown. In the above concept, it is assumed that similar routes lead to similar emissions. The similarity measure The similarity is used to determine such a similarity as described above.
[0126] In Figure 4 a real-world distribution of emissions 404 is also depicted and the corresponding estimate from the approximated emission distribution 412 Interesting statistical quantities about this distribution can be e.g. the center of gravity and or some higher quantile and In a better approximation, i.e. if a fully representative subset of routes is found, for all possible f there is a pair E
[0127] Trajectories (in this case routes) can be represented by a sequence of features x 1:T such that the desired target y (i.e. in this case emissions) is related to the trajectory representation x 1:T This means that the sequence of features x 1:T has a direct or indirect influence on the emissions y
[0128] For example, these features include: slope, elevation, speed limit, road type, road curvature, road surface, number of lanes, number of crossings of pedestrians or trains, traffic density distribution or weather conditions for each path point, route length.
[0129] In this case, is the real distribution of emissions for a given specific route x For example, when driving route x, y is a measured emission statistic, e.g. NOx per km.
[0130] The variance in x comes from all direct or indirect emission causing factors that are not observed in the route features
[0131] The driver can cause variance, e.g. due to attention, experience, aggressiveness of driving. The speed can cause variance, e.g. due to varying environment (e.g. state of traffic lights, pedestrian or train crossings, other drivers). The engine load can cause variance, e.g. due to vehicle load, operating state of air conditioning or number of passengers. The engine state or exhaust treatment state can cause variance, e.g. due to NH3 load or start temperature.
[0132] In this case, is the true distribution of emissions that occur. The true distribution can depend on a given vehicle type, i.e. a category C , a region R and / or a time span I .
[0133] Hence, the representative route selection has the following goals:
[0134] From the experience sample , M routes are selected that, when simulated, yield a similar distribution of emissions as all routes X in the real world.
[0135] The conditional emissions are unknown, but they are assumed to be similar to x . This is due to the selection of route features that are correlated with the target variable (in this example, emissions).
[0136] A success criterion for a good approximation of the true emissions distribution is met when the true emissions distribution has similar statistics as the true distribution.
[0137] In this example, the success criterion is a good approximation, i.e. if a completely representative subset of routes is found, e.g. for all possible E there is a pair E .
[0138] An alternative to representative route selection is random selection or clustering.
[0139] In random selection, the similarity between trajectories x and x’ is ignored. The subset X s can be randomly selected, e.g. with MMD = 0.22. Figure 5 The true underlying input distribution, the true real-world emissions distribution 506 and the resulting approximated emissions distribution 506 of the routes 502 and the exemplary selected subset 504 are depicted.
[0140] In this case, redundant information arises and in common scenarios, exemplary statistical values of the true real-world emissions distribution 504 are and the corresponding estimate from the resulting approximate emissions distribution 506 deviates. In addition, characteristics or patterns can be lost in the true distribution on the route p(x).
[0141] In clustering, the focus is on variance rather than frequency, and each selected route only represents its cluster. Instead, MMD is used to consider similarity to the whole dataset. Preferably, no clusters are used in case there are more clusters than selected routes. Figure 6 The true underlying input distribution, the true real-world emissions distribution 604 and the resulting approximate emissions distribution 606 of routes 602 and an exemplary selected subset 604 are depicted.
[0142] In this case, redundant information arises, and in common scenarios, exemplary statistical values of the true real-world emissions distribution 506 and the corresponding estimate from the resulting approximate emissions distribution 508 deviates. In addition, characteristics or patterns can be lost in the true distribution on the route p(x).
[0143] With reference to Figure 7 The training of the autoencoder framework is described.
[0144] In step 702, the trajectory is provided as an input sequence to a recurrent neural network, RNN. The output of the RNN is an extracted feature vector. The feature vector has a fixed dimensionality. The first RNN is an encoder g(.).
[0145] In step 704, a second RNN takes this output as a first hidden state and now tries to reconstruct the input sequence in turn. The second RNN is a decoder.
[0146] In step 706, the second RNN determines an output sequence.
[0147] In step 708, at least one parameter of the first RNN and / or the second RNN is determined which reduces a reconstruction loss between the input sequence and the output sequence.
[0148] The encoder thus trained can then be used for testing or for determining a subset X s for testing. The decoder is not used after training.
Claims
1. A computer-implemented method for testing a machine (102), characterized by: Provide a set of sample trajectories ( X ); Select set ( X A subset of the sample trajectories in the subset; the output of the machine (102) for the movement of the machine (102) is determined based on the sample trajectories in the subset. y s ); depends on the output ( y s The result of the test is determined by the set of characteristics ( X The subset is selected based on the score of the difference between the subset and the representation set, where the subset depends on the score of the difference between the subset and the representation set. X The first sample trajectory of ) x ) and sets ( X The score is determined by a similarity measure of the difference between the second sample trajectories. wherein the sample trajectory is provided (702) as an input sequence to an autoencoder comprising a first recurrent neural network, RNN, wherein the first RNN is configured as an encoder, wherein a second RNN is configured as a decoder and receives (704) a feature vector as an initial hidden state, which is extracted as an output and / or hidden state of the last time step of the first RNN, and wherein the second RNN determines (706) an output sequence, wherein at least one parameter of the first RNN and / or the second RNN is determined (708) which reduces a reconstruction loss between the input sequence and the output sequence, and wherein the first RNN of the thus trained autoencoder is used for testing or for determining a subset of tests.
2. The method of claim 1, wherein, output ( y s ) characterizes emissions of an internal combustion engine operable for a mobile machine (102), wherein when the output value ( y s ) is greater than a threshold value, the result indicates a test failure, and wherein otherwise the result indicates a test success.
3. The method according to one of the preceding claims, characterized in that, determining distribution statistics depending on a distribution of output values (x) y s , wherein the result indicates a test failure when the distribution statistics are greater than a threshold, and wherein the result indicates a test success otherwise.
4. The method of claim 3, wherein, The distribution statistic is the mean value or the 99% quantile.
5. The method according to claim 1 or 2, characterized in that, Determine the output value for the first parameter quantization of the machine (102) or simulation environment. y s The first distribution of ), and where, in the comparison, the output value ( y s The first distribution is compared with the second distribution, which represents the output value of the second parametric quantization of the machine (102) or the simulation environment, and wherein, depending on the result of the comparison, the first parametric quantization is used for testing, or the second distribution is used for testing.
6. The method of claim 1 or 2, wherein, Sample trajectory set ( X At least one sample trajectory in ) x 1:T A sequence of route features definition.
7. The method of claim 6, wherein, The route features characterize a geographical characteristic; and a traffic flow characteristic; and a road characteristic; and a traffic control characteristic; and a weather characteristic, a route distance, a time of day or week while driving the route, and / or a vehicle type of driving the route.
8. The method of claim 7, wherein, The geographical characteristic is an absolute height or a road slope characteristic; the traffic flow characteristic is a time-dependent average traffic speed; the road characteristic is a number of lanes, a road type, and / or a road curvature; the traffic control characteristic is a speed limit characteristic, a number of traffic lights, a number of traffic signs of a specific type, a number of stop signs, a number of yield signs, and / or a number of pedestrian crossing signs; and the weather characteristic is a rain amount, a wind speed, a fog at a predetermined time.
9. The method of claim 6, wherein, Sequence of route features to determine a sample trajectory x 1:T wherein the sequence of route features is defined along a route, wherein the route is defined by a pair of coordinate sequences and a start time 10. The method of claim 9, wherein, A pair of coordinate sequences characterizes a latitude and a longitude at a step.
11. The method of claim 6, wherein, The output value is determined by simulation of the machine (102) depending on the sample trajectories in the subset y s .
12. The method of claim 1 or 2, wherein, Provides a set of (302) sample trajectories ( X ), which represents an unbiased sample of the real-world distribution of the machine's (102) operations, where, ( X The first sample trajectory in ) x 1:T () is sampled from the set of sample trajectories The second sample trajectory is sampled from the set of sample trajectories. Among them, according to the first sample trajectory ( x 1:T A first feature vector with dimension is determined, and / or a second feature vector with the same dimension is determined based on a second sample trajectory, and a similarity measure is determined for the first feature vector and the second feature vector.
13. The method of claim 12, wherein, The real distribution is conditioned on a set of machine classes characterizing different variants of the machine (102), a set of regions of interest characterizing sample trajectories (108) X in a geographical space, and a set of time intervals (110) I base characterizing a set of sample trajectories (108) X in the geographical space.
14. The method of claim 1 or 2, wherein, Fraction( MMD ) was determined as the sample trajectory set ( X The pairwise similarity between samples in a subset and samples in a subset, and the approximation of the pairwise similarity between samples in a subset, wherein the pairwise similarity of a pair depends on the similarity metric used to determine the pairwise similarity.
15. The method of claim 14, wherein, The approximation is a stochastic approximation.
16. The method of claim 1 or 2, wherein, The similarity measure is determined depending on a trajectory kernel.
17. The method of claim 16, wherein, The trajectory kernel is a triangle global alignment kernel.
18. The method of claim 1, wherein, The feature vector has a fixed dimension, and the first RNN of the thus trained autoencoder is used for testing or for determining a subset of tests without the second RNN.
19. Apparatus for testing a machine, characterised in that, The device is adapted to carry out the steps according to one of claims 1 to 18.
20. A computer program product, characterised in that, The computer program product comprises computer readable instructions which, when executed by a computer, cause the computer to carry out the steps according to any one of claims 1 to 18.
Citation Information
Patent Citations
Method for determining a driving cycle for driving tests for determining exhaust emissions of motor vehicles
DE102017107271A1