Method for determining a sensor configuration
By determining whether a virtual sensor can be replaced with a real sensor in the vehicle and optimizing the vehicle sensor configuration, the problem of increasing weight, complexity and cost of a large number of sensors in the vehicle is solved, and a more efficient and safe sensor configuration is achieved.
Patent Information
- Application Number
- CN202080070877.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-08-09
- Filing Date
- 2020-08-06
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2040-08-06
AI Technical Summary
A large number of sensors in modern vehicles add to the weight, complexity and cost of the vehicle, and prior art is difficult to effectively replace physical sensors.
By establishing a preliminary sensor configuration, it is determined whether the real sensor can be replaced with a virtual sensor, thereby optimizing the sensor configuration and reducing the number of real sensors. The method involves recording and evaluating real sensor signals using artificial intelligence and machine learning techniques and simulating real sensor signals through virtual sensors.
It realizes the reduction of the number of real sensors, the weight and cost of the vehicle, while improving the redundancy and safety of sensor configurations without reducing the accuracy of vehicle behavior.
Smart Images

Figure CN114556248B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for determining a sensor configuration in a vehicle including a plurality of sensors. Background Art
[0002] Modern vehicles include a large number of sensors for detecting a variety of state variables, such as, for example, the rotational speed, temperature, force, torque, voltage, current, acceleration about the roll axis, pitch axis, and yaw axis of wheels, axles, gears, etc. In addition, vehicles sometimes include sensors for determining the position of the vehicle or the distance between the vehicle and other vehicles or obstacles. Other sensors are cameras that detect visual or non-visual images, such as rear-view cameras, infrared cameras, etc. The sensors are based on a variety of different technologies, such as rotary encoders, temperature probes, voltmeters, radar transmitters and receivers, CCD chips, etc.
[0003] The large number of sensors in a vehicle contribute to the weight, complexity, and cost of the vehicle.
[0004] “A study on the use of virtual sensors in vehicle control” by Canale M. et al. (Decision and Control, 2008. CDC 2008. 47th IEEE Conference on, IEEE, Piscataway, NJ, USA December 9, 2008 - ISBN 978-1-4244-3123-6) relates to the design of direct virtual sensors (DVS), where DVS technology can be considered to replace physical sensors.
[0005] US2008 / 312756A1 relates to a method for providing sensors for a machine. The method may include obtaining a data record including data from a plurality of sensors of the machine, and determining a virtual sensor corresponding to one of the plurality of sensors. The method may further include: establishing a virtual sensor process model of the virtual sensor based on the data record, the virtual sensor process model indicating the interrelationship between at least one sensed parameter and a plurality of measured parameters; and obtaining a set of values corresponding to the plurality of measured parameters. In addition, the method may include substantially simultaneously calculating a value of at least one sensed parameter based on the set of values corresponding to the plurality of measured parameters and the virtual sensor process model, and providing the value of at least one sensed parameter to a control system.
[0006] The present invention aims to solve at least part of the above problems. Summary of the Invention
[0007] The above object is achieved by a method for determining a sensor configuration in a vehicle comprising a plurality of sensors, the method comprising the steps of: establishing a preliminary sensor configuration for the vehicle, the sensor configuration comprising a first number of real sensors, each real sensor outputting a real sensor signal; determining whether at least one real sensor can be replaced by a virtual sensor; changing the preliminary sensor configuration into a final sensor configuration, the final sensor configuration comprising a second number of real sensors and at least one virtual sensor, wherein the second number is less than the first number.
[0008] A real sensor is a piece of hardware that measures a certain state variable, the certain state variable being a physical entity such as, for example, rotational speed, force, torque, light, etc.
[0009] A virtual sensor is a software module that receives at least one measurement signal from a real sensor and optionally other parameters and / or variables or signals, and preferably calculates a physical target value in real time from these inputs.
[0010] The basic idea of the present invention is to find an optimal sensor configuration among both the real and virtual sensors of the vehicle, i.e., to replace as many real sensors as possible with virtual sensors, and preferably to find an optimum between the accuracy achievable by the virtual sensors and the cost incurred by the real sensors.
[0011] The step of determining whether at least one real sensor can be replaced by a virtual sensor preferably comprises using artificial intelligence, in particular using machine learning techniques.
[0012] In many cases, the real sensor signals are recorded and then evaluated. The recording of the real sensor signals can be carried out during a test run of the vehicle, wherein the evaluation of the recorded real sensor signals is subsequently carried out on a stationary evaluation computer.
[0013] In an alternative embodiment, the recording of the real sensor signals is carried out during a test run of the vehicle, wherein the evaluation of the recorded real sensor signals and the replacement of at least one real sensor by a virtual sensor are carried out on a mobile evaluation computer during the test run.
[0014] Furthermore, it is possible to carry out at least one of the recording of the real sensor signals, the evaluation of the recorded real sensor signals, and the replacement of at least one real sensor by a virtual sensor on an analog computer.
[0015] Using a mobile evaluation computer has the advantage that the impact of replacing a real sensor on the vehicle behavior can be experienced immediately. On the other hand, using an analog computer has the advantage that a real test drive can be dispensed with.
[0016] In some evaluation examples, the step of determining whether at least one real sensor can be replaced by a virtual sensor is not performed for each real sensor. Instead, for example, due to security considerations, some real sensors can be classified as "irreplaceable". Secondly, some sensors are very cheap and have low weight. Therefore, only when a real sensor has a significant weight and / or a significant cost, one may consider evaluating whether a certain real sensor can be replaced by a virtual sensor. In addition, some real sensors in certain environments can be defined as "must be replaced". This applies, for example, to a development environment where the initial sensor configuration includes not only the sensors that will be implemented in the vehicle to be produced. Instead, such a development environment can include sensors that are set up and connected only for development purposes. These "development sensors" are no longer available in mass-produced vehicles and are therefore considered "must be replaced".
[0017] In addition, the accuracy of the virtual sensor and the time delay that the virtual sensor may have compared to the real sensor can be relevant considerations. The time delay may be caused by complex calculations based on the input to the virtual sensor. On the other hand, some virtual sensors may not be as accurate as the real sensor they replace. The loss of accuracy and the time delay may have an impact on the vehicle behavior, and in some cases, the vehicle behavior also needs to be analyzed and evaluated to determine whether the replacement of the real sensor is possible. Therefore, the question of whether a real sensor can be replaced by a virtual sensor is usually not a clear yes or no question, but a question considering several boundary conditions, and the several boundary conditions may furthermore be weighted in order to achieve an optimal final sensor configuration.
[0018] In addition, although real sensors may be replaced for cost reasons, the present invention can also be used to avoid replacing real sensors, but instead create secondary virtual sensors for real sensors in order to improve the redundancy and possible safety of the sensor configuration.
[0019] This object is fully achieved.
[0020] In a preferred embodiment, the determining step includes recording real sensor signals of at least a subset of a first number of real sensors, evaluating the recorded real sensor signals to determine whether at least the first real sensor can be replaced by a first virtual sensor that receives at least one real sensor signal from a second real sensor and outputs a virtual sensor signal that simulates the real sensor signal of the first real sensor.
[0021] As discussed above, the evaluating step can be performed in multiple different steps.
[0022] In a preferred embodiment, the evaluation step includes using a Boltzmann machine having a plurality of visible nodes and a plurality of hidden nodes, each visible node representing a real sensor, and the hidden nodes being calculated by utilizing combinations of the nodes.
[0023] Using a Boltzmann machine in the evaluation step is a powerful method. A Boltzmann machine is an undirected generative stochastic neural network that can learn the probability distribution of its input set. It can always generate different states of the system.
[0024] Given infinite training data, a Boltzmann machine can represent any system with many states. In the current case, the system first represents a preliminary sensor configuration. The visible nodes are the features / inputs to the system, which are the real sensors in the vehicle. The hidden nodes are the nodes to be trained, which will identify and utilize combinations of the visible nodes. Essentially, the Boltzmann machine tries to learn how the nodes affect each other by estimating the weights in its edges (the edges are similar to conditional probability distributions).
[0025] In theory, once the model is trained, the Boltzmann machine can reconstruct all sensors given only one sensor. In other words, this theoretical approach would lead to a concept where only a single physical sensor is required to construct all the other sensors in the vehicle.
[0026] Although the Boltzmann machine is a great model in theory and can solve many problems, it is very difficult to implement in practice. This is due to computational power, as an increase in the number of nodes leads to an exponential increase in the number of edges / connections. If a preliminary sensor configuration uses 200 sensors in a vehicle and if an additional 400 hidden nodes are added, the number of edges will be 600 x (600 - 1) / 2 = 179,700 edges.
[0027] Therefore, a restricted Boltzmann machine (RBM) is usually used, where nodes of the same type do not connect to each other. This concept is used to trade off performance against the ability to run the calculations. The contrastive divergence algorithm is used to train and predict the RBM in the same way as the Boltzmann machine.
[0028] Boltzmann machines and restricted Boltzmann machines are structures that may not consider the time dependence in a time series. Therefore, the Boltzmann machine used is preferably a recurrent time-restricted Boltzmann machine. That is, when processing signals and time series, this Boltzmann machine can be used. The recurrent time-restricted Boltzmann machine (RTRBM) uses recurrent neurons as storage units for memory paths and uses backpropagation through time in the contrastive divergence algorithm to train the model. A very advanced and powerful type of RTRMB is the RNN-Gaussian dynamic Boltzmann machine, which is preferably used to model sensor configurations.
[0029] Generally and as explained above, the step of determining whether at least one real sensor can be replaced by a virtual sensor is a function of the accuracy of the virtual sensor and / or the driving behavior of the vehicle and / or the cost of the real sensor to be replaced.
[0030] It is preferred herein if the accuracy of the virtual sensor and / or the driving behavior of the vehicle and / or the cost of the real sensor to be replaced are weighted and calculated as a target value.
[0031] The above concept of determining whether at least one real sensor can be replaced by a virtual sensor is more or less based on brute-force methods. However, there are also ways to perform this step based on causal relationship analysis.
[0032] Therefore, it is preferred if the determination step includes detecting and recording the outputs of at least a subset of real sensors for a predetermined number of temporary subsequent sampling steps, and performing a causal relationship analysis that determines the causal relationships between the recorded outputs of the real sensors.
[0033] The term causal relationship or causality should be understood as the relationship between cause and effect. The basic problem in this method is whether and to what extent one real sensor causes another real sensor. The terms causal relationship, causality, and correlation are used interchangeably within this application. The broadest interpretation will be applied to each of those terms.
[0034] Furthermore, in this application, if one sensor output causes another sensor output, this basically means that the other sensor outputs depend on one sensor output.
[0035] Preferably, the outputs of all sensors of the preliminary sensor configuration are passed to an algorithm that will determine whether a sensor can be replaced, and the algorithm is preferably capable of constructing a model of the sensor (virtual sensor) that replaces the real sensor.
[0036] In a preferred embodiment, it can be assumed that a preliminary sensor configuration forms a sensor space X. A dependency graph reflecting dependencies or causal relationships between real sensors has a plurality of edges, which can be constructed by the following statement:
[0037] In this statement, C is a binary causal relationship function, f is a penalty factor taking into account aspects such as cost or safety (like redundancy), and E x,y is the coefficient of the edge.
[0038] Some measures said to measure causality are Granger causality, transfer entropy, transfer cross mapping, and mutual information, while correlation can be estimated by the Pearson autocorrelation algorithm.
[0039] For each of the mentioned measures, one must specify a maximum lag (shift). The lag typically corresponds to a certain number of time - space samples. In one lag, the number of samples of different sensors may be different because different sensors may have different sampling frequencies. For example, if one signal has a sampling time of 10 ms (corresponding to a sampling frequency of 100 Hz), and if another signal has a sampling time of 100 ms, then a lag of 10 will mean a period of 100 ms being considered for the first signal and a period of 1 s for the second signal. For any causality calculation, it does not matter whether the signals have the same time base. However, this may ultimately be relevant for the subsequent training of virtual sensors.
[0040] The result of the causality metric is a matrix, which preferably includes the relationships between sensors. The relationships or values or causality of the matrix are preferably normalized or standardized such that the maximum causality has a value of 1 and the minimum causality has a value of 0.
[0041] In a preferred embodiment, the causal relationships between the recorded outputs of real sensors are determined for at least a subset of the samples, where the causal relationships determined for the sample subset are post - processed in order to determine a final set or matrix of causal relationships between the recorded outputs of real sensors.
[0042] The sample subset can be defined as a single vector of the form [lag; now], where lag < or = maximum lag.
[0043] Furthermore, in a preferred embodiment, a directed cyclic graph (DCG) is established based on the determined causal relationships. The weights in the directed edges explain "how much one sensor causes / relates to other sensors". One can use a depth - first search (DFS) algorithm to detect cycles in the graph.
[0044] To find out which real sensors can be best replaced, it is preferred if the DCG is converted into a directed acyclic graph (DAG), where the real sensor with the highest causality or the real sensor with the lowest causality is taken as the root of the DAG.
[0045] One can use several algorithms to convert the DCG into a DAG. Another strategy would be to take only the most dependent sensors as the root and build a tree from there.
[0046] A directed acyclic graph is a tree that has a root and a stem and final leaves. For example, sensors can be replaced by removing the leaves at the nodes of the tree. Each sensor at a leaf will identify the pipeline through the model, where the target value is the sensor signal to be reconstructed and where the input is the corresponding parent in the tree. Additionally, one can remove more levels, but it should be borne in mind that the more levels are removed, the less accurate the reconstruction will be.
[0047] As another example, it may be possible to identify replaceable sensors by using a graph sorting algorithm on the DCG in order to identify the most important sensors based on their outgoing and incoming causal edges. One such sensor is the personalized PageRank algorithm.
[0048] Correspondingly, it is preferred if at least one real sensor that forms a leaf or a root in the DAG is determined to be replaceable.
[0049] When establishing the DCG, a preferred method of replacement is to calculate a sorting matrix based on the DCG, where at least one real sensor is determined to have a low rank and is thus replaceable.
[0050] The sorting matrix can be calculated based on a sorting algorithm. An example of such a sorting algorithm is the page sorting algorithm used in search engines.
[0051] As another preferred example of using the DCG, it may be possible to generate a random probability process based on the DCG, where the state of at least one real sensor can be reached with the state of another real sensor and can thus be determined to be replaceable.
[0052] The random probability process can be implemented by a probability algorithm such as, for example, Markov chain Monte Carlo (MCMC).
[0053] Additionally, it is preferred if the mathematical model of a real sensor that has been determined to be replaceable is determined based on statistical or deterministic methods (algorithms).
[0054] Specifically, for each pruned leaf from the previous step, the leaf is taken as the label and the causal branch (starting from the root to the leaf) is taken as the feature to train the model. As a favorable way to identify the model that can replace a given sensor (i.e., turn a real sensor into a virtual sensor), any statistical or deterministic algorithm that can learn the signal representation can be used to build the model that will be used to reconstruct the sensor.
[0055] For example, a neural network architecture called Time Delay Neural Network (TDNN) can be used.
[0056] Such a network is a feed-forward neural network that can be applied to time series. The general architecture will be used for all pruned leaves. However, an optimization algorithm should be used to optimize the hyperparameters of the model to help the general algorithm architecture be specialized for a given problem such as grid search, random search, or Bayesian hyperparameter optimization.
[0057] For the example of TDNN, the following hyperparameters can be optimized: the number of neurons, the number of layers, the dropout rate, etc.
[0058] Once the model is trained and evaluated against the test set, it is preferred if the weights and parameters of the model are extracted and if the predictions of the model are calculated through feed-forward computation.
[0059] When using the method of causal analysis, there is an important aspect that causality can depend on the current system state. For example, the speed related to an automotive transmission may have the following causality: when the starting clutch is closed, there is a high causality between the engine speed and the wheel speed. On the other hand, when the clutch is open, the causality is low. There are several ways to solve this problem, such as:
[0060] 1. People put expertise into the computational routine, for example, by increasing a predefined penalty factor f, or by defining the engine speed sensor as irreplaceable.
[0061] 2. People can utilize event recognition to trigger state differences. For example, in an automatic transmission like a dual-clutch transmission, there are signals that give information about the clutch state, enabling people to distinguish between the slipping and sticking phases. Thus, people will obtain the causality values for each of these states. Depending on these values, people can then determine whether the sensor can be replaced in its "entirety" (i.e., its meaning is: by treating the time series as a unique signal again, rather than a series of events, and by evaluating that one of the two causality values is too low, such that people can say that the given sensor is actually irreplaceable).
[0062] 3. One can perform the proposed causality calculation and identify the optimal lag. Once this is done, one can again go through the time series (by maintaining the calculated optimal lag) and calculate a relative error for each individual data point. Based on this, one can plot the relative error distribution of all the points. In this way, one can see whether the errors are always concentrated in a given region or whether there is dispersion. In the second case, one might reject replacing the sensor and impose a penalty on it.
[0063] 4. One can keep the causality calculation as it is and not care about possible state dependencies, such that one will train the model for the virtual sensor anyway. At the end of the training, one will evaluate the accuracy of the generated model. If this is not good enough, one might then decide that the sensor is actually not replaceable (in any case, the final decision on sensor replaceability is made after the model is trained). If the model is good enough, the sensor can be replaced and event dependencies might not be crucial.
[0064] The whole method is highly parallelizable, where building the graph, building models for each pruned leaf, and model optimization can be multi-threaded.
[0065] In addition, there are many ways to set the maximum lag. Schwert (1989) proposed a heuristic method as a rule of thumb. It is calculated as follows:
[0066] Max._lag = [12x(T / 100) 0.25
[0067] where T is the number of observations in the signal, i.e., the length of the signal. As mentioned, the Schwert rule of thumb is an ad hoc method, and getting the lag value correct is challenging because too small a lag value will bias the statistical test. However, too large a value will magnify the power of the statistical test. There are many publications that have proposed the following: it is best to have an error (type 2 error) on the side of including too many lags.
[0068] In another preferred aspect of the present invention, the determination step includes detecting and recording the outputs of at least a subset of the real sensors, and performing a causality analysis that determines the causal relationships between the recorded outputs of the subset of real sensors, where the causality analysis includes building a component-wise neural network CWNN, where each real sensor of the subset of real sensors corresponds to one of the components of the CWNN, and where each component is formed by a virtual sensor that is trained to simulate the corresponding real sensor.
[0069] The virtual sensor is preferably a sub-model of a neural network. In a preferred embodiment, the training step uses the outputs of some or each of the other real sensors of a subset of real sensors. In addition, the past outputs of the real sensors to be simulated (so-called targets) can also be used to train the virtual sensor.
[0070] The virtual sensors (sub-models) of the neural network can be trained individually or all together.
[0071] Preferably, the training step includes applying a sparsity-inducing penalty to the corresponding first hidden layer of at least some of the virtual sensors.
[0072] Preferably, the sparsity-inducing penalty is applied to the corresponding first hidden layer of each virtual sensor. When applying the sparsity-inducing penalty, the similar features are grouped together using the parameter tying technique, and the features that do not Granger-cause the target are zeroed out.
[0073] In one embodiment, it is preferred if the sparsity-inducing penalty is selected from the group lasso regularization family.
[0074] In another preferred embodiment, the sparsity-inducing penalty is selected from the group-ordered weighted lasso (GrOWL) rule family.
[0075] In addition, it is preferred if a sparsity-inducing optimizer is used to optimize the sparsity-inducing penalty in order to generate a sparse model.
[0076] Here, it is preferred if the semi-random approximate gradient descent SPGD algorithm is used to optimize the sparse model.
[0077] In an alternative preferred embodiment, the Follow-the-Regularized-Leader FtRL algorithm is used to optimize the sparse model.
[0078] In a preferred aspect of the present invention, it is generally preferred if a causality vector is calculated for each trained virtual sensor (sub-model), and the causality vectors are concatenated to generate a causality matrix.
[0079] In this case, it is preferred if calculating the causality vector of the corresponding sub-model includes the following:
[0080] - Converting the weight matrix of the first layer of the virtual sensor into an affinity matrix,
[0081] - Clustering the affinity matrix to group similar features together,
[0082] - Sorting the clusters according to importance,
[0083] - Sorting the features in each cluster according to importance,
[0084] - Calculate the global ranking of features by considering the ranking of clusters and the ranking of features.
[0085] - Use the global ranking as the causal relationship vector.
[0086] Here, it is advantageous if the ranking of clusters by importance is done by a permutation test method.
[0087] In an alternative embodiment, the ranking of clusters by importance is done by a zeroing method.
[0088] It will be understood that, without departing from the scope of the present invention, the features of the present invention mentioned above and those to be explained below can be used not only in the indicated corresponding combinations, but also in other combinations or alone. BRIEF DESCRIPTION OF THE DRAWINGS
[0089] Exemplary embodiments of the present invention are explained in more detail in the following description and shown in the drawings, wherein:
[0090] Figure 1 is a schematic diagram of a motor vehicle having a sensor configuration;
[0091] Figure 2 is Figure 1 the output of several sensors of the sensor configuration over a certain period of time;
[0092] Figure 3 is a schematic diagram of the sensor configuration process;
[0093] Figure 4 is an example of a causal relationship matrix;
[0094] Figure 5 is based on Figure 4 an example of a directed cyclic graph of the causal relationship matrix;
[0095] Figure 6 is based on Figure 5 an example of a directed acyclic graph of the directed cyclic graph;
[0096] Figure 7 is another example of a causal relationship matrix;
[0097] Figure 8 is based on Figure 7 an example of a directed cyclic graph of the causal relationship matrix;
[0098] Figure 9 is based on Figure 8 an example of a directed acyclic graph of the directed cyclic graph;
[0099] Figure 10It is a schematic diagram of a Boltzmann machine;
[0100] Figure 11 It is a schematic diagram of a restricted Boltzmann machine;
[0101] Figure 12 It is another example of a restricted Boltzmann machine for sensors used in a vehicle driveline;
[0102] Figure 13 It is Figure 12 The actual speed values of the driveline of the RBM within three hysteresis;
[0103] Figure 14 It is based on Figure 12 The restricted Boltzmann machine in [[ ]] with two sensors identified as replaceable;
[0104] Figure 15 It is Figure 14 Another example of the outputs of two real sensors of the RBM of [[ ]];
[0105] Figure 16 It is a flowchart of a method for determining a sensor configuration according to a preferred aspect of the present invention;
[0106] Figure 17 It is Figure 16 An embodiment of the virtual sensor (sub - model) architecture of the method of [[ ]];
[0107] Figure 18 It is Figure 16 Another embodiment of the virtual sensor (sub - model) architecture of the method of [[ ]]; and
[0108] Figure 19 It is according to Figure 16 The causality analysis concept of the method of [[ ]]; Detailed Description of the Invention
[0109] In [[ ]], a vehicle 10 is shown, which can be a motor vehicle such as a passenger car, for example. The vehicle 10 has a body 12, front wheels 14L, 14R and rear wheels 16L, 16R. Figure 1 The rear wheels 16L, 16R are driven wheels driven by a driveline 18.
[0110] The driveline 18 includes an internal combustion engine 20 and a transmission 24. The internal combustion engine 20 and the transmission 24 are preferably connected via a clutch device 22 such as a starting clutch.
[0111] Generally, the transmission 24 includes a plurality of shiftable gear stages 25 for establishing a plurality of gear stages.
[0112]
[0113] The output of the transmission 24 is connected to a differential 26 that is adapted to distribute driving force to the driven rear wheels 16L, 16R.
[0114] The vehicle 10 includes a plurality of sensors, such as an engine speed sensor 30 for detecting the rotational speed Seng of the internal combustion engine 20.
[0115] In addition, the transmission 24 includes a first transmission speed sensor 32 that detects the speed of the input shaft of the transmission. In addition, the transmission 24 includes a second transmission speed sensor 34 that detects a second transmission speed, such as the rotational speed ST, Strn of the output shaft of the transmission 24.
[0116] In addition, the driveline 18 may include a left driven wheel sensor 36 for measuring the rotational speed SL, Swl of the left driven rear wheel 16L, and a right driven wheel sensor 38 for detecting the rotational speed SR, Swr of the right driven wheel 16R.
[0117] The sensors 30 to 38 are connected to a controller 40, which may be the driveline controller 18. The controller 40 may be a multi-system controller, including, for example, a transmission controller, an internal combustion engine controller, etc.
[0118] In addition, the vehicle 10 may include additional sensors, such as an engine torque sensor 42 for detecting the torque provided by the internal combustion engine 20. The additional sensors may include a clutch position sensor 44 for detecting the clutch position of the clutch device 22, and one or more temperature sensors 46 for measuring, for example, the fluid temperature in the transmission 24.
[0119] The vehicle 10 may include a large number of additional sensors for measuring, for example, the rotational speed of an electric motor for adjusting the tilt of the vehicle seat, a temperature sensor for measuring the temperature inside the passenger compartment, a radar sensor (such as LIDAR) for measuring distance, a camera sensor for detecting the vehicle's surrounding environment, an acceleration sensor for detecting roll movement, pitch movement, and / or yaw movement. In addition, a plurality of electrical sensors for measuring voltage, current, etc. may be provided.
[0120] At least some of the sensors, preferably each sensor, are connected to a controller of the vehicle, which may include the driveline controller 40 mentioned above.
[0121] In addition, any controller (such as controller 40) may be connected via wireless communication 48 to a network 46 external to the vehicle 10, such as the Internet, a GPS network, a cellular phone network, a wireless local area network (WLAN, Wifi) network, etc.
[0122] Figure 1An evaluation computer 50 is also shown.
[0123] The evaluation computer 50 is connected to at least one controller of the vehicle (such as controller 40) and is adapted to perform a method for determining a sensor configuration in a vehicle 10 including a plurality of sensors. The method includes: a step of determining a preliminary sensor configuration of the vehicle, the preliminary sensor configuration including a first number of real sensors, each real sensor outputting a real sensor signal; and a step of determining whether at least one real sensor can be replaced by a virtual sensor, and including a step of changing the preliminary sensor configuration into a final sensor configuration, the final sensor configuration including a second number of real sensors and at least one virtual sensor, wherein the second number is less than the first number.
[0124] The method can be carried out according to a plurality of different embodiments, some of which are explained below. The following embodiments mainly relate to the sensor configuration for the powertrain 18. However, the embodiments currently applied to the powertrain 18 can also be applied to other parts of the vehicle 10, such as applied to the navigation system configuration, temperature control configuration, etc.
[0125] Figure 2 There is shown Figure 1 three diagrams of the outputs of several sensors of the shown sensor configuration over a certain period of time. In particular, the first diagram shows the second transmission speed ST measured by sensor 34 over time; the second diagram shows the left driven wheel speed SL measured by sensor 36; and the third diagram shows the right driven wheel speed SR measured by sensor 38.
[0126] In the diagrams, it is assumed that the current time is t. In addition, it is assumed that each sensor has a similar sampling frequency corresponding to the same sampling period, although this is not necessary.
[0127] Figure 2 There is shown a window 54 corresponding to a plurality of sampling time periods. In Figure 2 it, one sampling time period is indicated as a single lag 56. The window 54 generally consists of a plurality of single lags 56. Figure 2 The window 54 shown in
[0128] corresponds to the maximum lag 58. The maximum lag 58 corresponds to the maximum number of single lags 56, which is used in the process of determining whether at least one real sensor can be replaced by a virtual sensor.
[0129] The optimal lag can be determined by one or more of the following:
[0130] - Using statistical and information criteria such as the Akaike or Bayesian information criteria (AIC, BIC);
[0131] - The first minimum of the mutual information between the time series and its shifted version; and
[0132] - Trial and error.
[0133] In the Figure 2 diagram, one can see that ST is almost constant from the start up to t-15.
[0134] In the time period from t-25 to t-20, SL deviates from ST and is greater than ST. Similarly, SR is less than ST during the time period from t-25 to t-20.
[0135] At t-15, the transmission output speed ST starts to decrease to zero. The output transmission speed is achieved as zero at t-10.
[0136] At this point, the vehicle is at a stop. Correspondingly, SL and SR are also zero.
[0137] If the driver wishes to start the vehicle again, he may experience a situation on the μ bypass road where, for example, the right follower speed SR remains zero within a few samples while the other follower wheel speed SL increases.
[0138] From t-10 to t-5, the right follower wheel speed SR remains at zero and then accelerates again, for example due to the braking effect applied to the right follower wheel by the anti-skid control.
[0139] At t, the speeds ST, SL, and SR are equal again.
[0140] One can see from Figure 2 that the follower wheel speeds SL, SR have a certain relationship with the output transmission speed ST. During some time windows, they are the same. At other times, they may deviate quite significantly from the output transmission speed ST.
[0141] However, one can say that at least for some cases and preferably for most of the time, the follower wheel speeds SL, SR are causing the output transmission speed ST.
[0142] The question arises: Can any one of these three sensors (the real sensor in the Figure 1 example) be replaced by a virtual sensor that receives at least one real sensor signal from the real sensor and outputs a virtual sensor signal that simulates the real sensor signal of the replaced real sensor?
[0143] Figure 3 It is a schematic diagram of the sensor configuration process.
[0144] The sensor configuration process includes using a so-called causal stage 72, into which the true sensor signals and optionally other parameters are input. The causal stage 72 includes a causal relationship matrix 74 established based on the true sensor signals, looking at the true sensor signals for a certain lag, ideally the optimal lag.
[0145] The causal relationship matrix 74 is based on the causal relationships between the recorded outputs of the true sensors, which are determined for at least a subset of the samples (e.g., the optimal lag), and wherein the causal relationships determined for the sample subset are post-processed in order to determine the final set or matrix of causal relationships between the recorded outputs of the true sensors.
[0146] In other words, the causal relationship matrix is a representation of the result of a causal relationship analysis that determines the causal relationships between the recorded outputs of the true sensors.
[0147] The causal relationship matrix 74 is used to establish a directed cyclic graph (DCG).
[0148] In block 76 of the causal stage 72, a conversion process is performed in order to convert the DCG into a directed acyclic graph (DAG), where the true sensor with the highest causal relationship or the true sensor with the lowest causal relationship is taken as the root of the directed acyclic graph.
[0149] In the directed acyclic graph (DAG), at least one true sensor that forms a leaf or the root of the graph is determined to be replaceable.
[0150] In other words, the causal stage 72 determines which true sensors can be replaced with virtual sensors.
[0151] The output of the causal stage 72 is input into a modeling stage 78, which is used to model the virtual sensors that will replace the true sensors. The modeling stage 78 includes a model construction process 80, where the model of the virtual sensor is constructed. In addition, the modeling stage 78 includes a model optimization process 82, where the model of process 80 is optimized.
[0152] Finally, the virtual sensors are included in the final sensor configuration, which is shown at 84, and based on this final sensor configuration, code is generated for implementing the virtual sensors.
[0153] Figure 4 It is an example of a causal relationship matrix 74’, which shows an example of the causal relationships between six sensors X1 to X6.
[0154] In the first row, it is shown that sensor X1 causes sensor X2 with a factor of 0.8 (causality 75a) and sensor X5 with a factor of 0.7, while X1 does not cause any of sensors X3, X4, X6 at all.
[0155] Causality is typically a value between 0 and 1, where 0 means that a sensor does not cause another sensor at all. On the other hand, "1" means that a sensor completely causes another sensor, such that the other sensor is redundant or even superfluous. In any case, the other sensor can be replaced by the first sensor.
[0156] Another example is that, for instance, sensor X4 causes sensor X3 with a value of 0.4 (causality 75b), while sensor X3 causes X4 with a value of 0.7. These two sensors X3, X4 do not cause any other sensors.
[0157] In Figure 4 sensor configuration X is shown as X = {X 1 , X 2 , X 3 , X 4 , X 5 , X 6}.
[0158] Figure 5 A directed cyclic graph established based on the causality determined in the causality matrix 74' is shown. Figure 5 The directed cyclic graph 88 of Figure 5 shows that X1 causes X2 and X5. In addition, it is shown that only X5 causes X1, and X5 also causes X2. Figure 4 It is also shown that X6 does not cause any other sensors and is not caused by any other sensors, as can be seen from the last row and the last column of the causality matrix 74' of
[0159] In a directed cyclic graph (DCG), the weight in the directed edge explains "how much the corresponding sensor causes / relates to other sensors". One can use the depth-first search (DFS) algorithm to detect cycles in the graph. In order to identify the dependency order of sensors, the DCG must be transformed into a directed acyclic graph (DAG).
[0160] Figure 5 The directed cyclic graph DCG 88 of Figure 6 can be transformed into a directed acyclic graph (DAG) as shown as 90 in
[0161] Figure 6 The DAG 90 in
[0162] On the other hand, X2 causes X3, and X3 causes X4. Figure 6 It is shown that the DAG has four levels, where level 0 corresponds to the root of the DAG, and level 3 corresponds to the leaves of the DAG.
[0163] Figure 6 The graph or tree in is a four-level tree. It will be noted that X4 is the son of X3 and not the son of X2 because X3 has a higher causality. From the DAG, one can start replacing sensors by removing the leaves at the nodes of the tree. Each sensor at the leaf will identify the pipeline through the model, where the target value is the sensor signal that one wishes to reconstruct, and the input is the corresponding parent in the tree. Additionally, one can remove more levels, but it should be borne in mind that the more levels are removed, the less accurate the reconstruction will be.
[0164] In Figure 6 's DAG, it is quite clear that X4 can be replaced by a virtual sensor.
[0165] The above example illustrates a simple sensor space or configuration, where the DAG tree is constructed based on the understanding that all sensors in the configuration are usually replaceable, and then the least dependent sensors are removed. However, in a dynamic system like a car, for safety reasons, there are many redundant sensors that should neither be removed nor replaced. Therefore, it is important for the algorithm to distinguish these important non-replaceable sensors in the configuration / space. Thus, for example, one can associate each sensor with a flag variable indicating whether it is replaceable, or utilize a high penalty factor f to reflect non-replaceability. Then, the algorithm that converts the DCG to a DAG can take non-replaceability into account when constructing the tree. For example, in the former method (flag), the algorithm designates the sensor Xi that has a flag value of "irreplaceable" and is the most dependent sensor as the root and constructs the tree from there, or it can be excluded from the process.
[0166] In some cases, the combination of two or more sensors (any time signal operation, such as some summation, subtraction, or dynamic scaling) can cause another sensor. If the resulting sensor is worthy of replacement, then one combines that sensor into a new hybrid sensor, which will be added to the sensor configuration. A simple example can be seen when reconstructing the output speed of a transmission: due to the differential, during a turn, the speed of an individual left or right wheel does not cause the transmission output speed. However, the mean or average of the speeds of the two wheels directly causes and can be used to infer the transmission output speed ST. When two or more sensors are combined, then all the involved sensors in the combination should be marked as "irreplaceable".
[0167] Figure 7 Another example of a causal relationship matrix 74″ is shown, which shows the rotational speed Seng measured by, for example, sensor 30, the rotational speed Stran measured by, for example, sensor 32, the rotational speed Swl measured by, for example, sensor 36, the rotational speed Swr measured by, for example, sensor 38, and the rotational speed Swl, wr corresponding to the sum of sensor 36 and sensor 38.
[0168] From the above analysis, it is known that the combination of the two wheel speeds Swr, Swl directly causes the transmission speed ST(Swl, wr). Therefore, one can create a new linear combination sensor Swl, wr, add it to the sensor space, and switch the reproducible flags of Swl and Swr. This has been done in Figure 7 the causal relationship matrix.
[0169] When one assumes that a causal relationship graph is constructed with a maximum lag of 50, the values in the causal relationship matrix may differ based on the lag value (in the above example, the values were assumed based on domain knowledge and not actually calculated).
[0170] Figure 7 The causal relationship matrix 74” can be calculated, for example, using Granger causality with a maximum lag of 83. The maximum lag can be calculated by “Schwert”, and the optimal lag can be selected using the Akaike Information Criterion (AIC), alternatively based on the Bayesian Information Criterion (BIC).
[0171] Figure 8 shows an example of a directed cyclic graph 88” established based on Figure 7 the causal relationship matrix 74”.
[0172] Figure 8 The DCG 88” shows that the corresponding sensors cause each other depending on the thickness of the corresponding lines and the path length from the root.
[0173] Figure 9 shows a directed acyclic graph 90” converted from Figure 8 the DCG 88”.
[0174] It can be seen that the speeds Swl, wr corresponding to the output transmission speed ST are taken as the roots, and the true sensor signals cause Seng, Swl, and Swr.
[0175] Furthermore, Seng causes Stran (corresponding to ST) to a certain extent.
[0176] In view of the above, it can be seen that Stran forms the leaf of the directed acyclic graph 90” and thus indicates that the corresponding true sensor 32 may be replaced by a virtual sensor.
[0177] Figure 9 shows the most suitable tree that can be generated from the DCG 88” considering the causality values and the flags as mentioned above. Figure 8
[0178] In Figure 9 , the sensors in level 0 and level 1 cannot be replaced. Therefore, using the combination of the driven rear wheel speed and the engine speed as the features of the model used in the modeling phase 78, only the output transmission speed sensor Stran can be replaced.
[0179] In the modeling phase 78, the causal branches of the DAG are taken as the features for training the model. As a favorable way to identify the models that can replace a given sensor (i.e., turning a real sensor into a virtual sensor), any statistical or deterministic algorithm for learning signal representation can be used to construct the model that will be used to reconstruct the sensor. For example, a neural network architecture called the Time Delay Neural Network (TDNN) can be used. This is a feedforward neural network suitable for time series. This general architecture can be used for all pruned leaves. However, an optimization algorithm (in 82) should be used to optimize the hyperparameters of the model to help the general algorithm architecture be specialized for a given problem, such as grid search, random search, or Bayesian hyperparameter optimization. For an example of TDNN, the following hyperparameters can be optimized: the number of neurons, the number of layers, the dropout rate, etc.
[0180] Once the model is trained and evaluated against the test set, the final sensor configuration of process 70 will extract the weights and parameters of the model and calculate the predictions of the model through feedforward calculation.
[0181] In Figures 10 to 15 , another method for replacing the sensors of the preliminary sensor configuration is shown. The method shown in Figures 10 to 15 is a brute - force method using a Boltzmann Machine (BM). The BM is an undirected generative stochastic neural network that can learn the probability distribution of its input set. It can always generate different states of the system.
[0182] For example, the Boltzmann Machine shown at 100 in Figure 10 can represent any system with many states given infinite training data. In the current case, it will represent the above - mentioned sensor configuration.
[0183] Figure 10 Illustrates the architecture of the BM. The visible nodes 102 are the features / inputs to the system, which in our case are all the sensors in the vehicle. On the other hand, the hidden nodes 104 are the nodes to be trained, which will identify and utilize the combinations of the visible nodes. Essentially, the BM tries to learn how the nodes affect each other by estimating the weights in their edges (the edges are similar to conditional probability distributions).
[0184] Boltzmann machines and their variants use the contrastive divergence algorithm to train the model. Briefly, the training works as follows:
[0185] 1. Randomly initialize the weights between the nodes;
[0186] 2. Feed the sample input vector to the visible nodes;
[0187] 3: Calculate the hidden nodes based on the weights and the global bias (feed-forward method);
[0188] 4. Reconstruct the visible nodes from the hidden nodes;
[0189] 5. Compare the visible nodes with the reconstructed visible nodes using, for example, the Kullback divergence;
[0190] 6. Update the weights using gradient descent based on, for example, the Kullback divergence loss function; and
[0191] 7. Repeat steps 2 to 6 for all feature samples until convergence.
[0192] Although, in theory, Figure 10 the Boltzmann machine is a good model and can solve many problems, it is actually very difficult to implement due to the required computational power.
[0193] Therefore, one can use a variant of the Boltzmann machine called the restricted Boltzmann machine (RBM), where nodes of the same type do not connect to each other, as Figure 11 shown at 100’ in
[0194] Here, the visible node 100” is connected to the hidden node 104’ by an edge 106’, but the visible nodes 100” do not connect to each other, nor do the hidden nodes 104’ connect to each other.
[0195] In the above example, for different speeds, namely the engine speed Seng, the transmission speed Stran, the speed Swl of the left driven wheel, and the speed Swr of the right driven wheel, a restricted Boltzmann machine (RBM) as shown at 100” in Figure 12 can be established.
[0196] Here, the visible node 100” corresponds to the four speeds mentioned above. In addition, multiple hidden nodes are established, where the number of hidden nodes is preferably greater than the number of visible nodes.
[0197] During training, the model requires as Figure 12The large dataset of all sensors shown, as discussed previously. For each sample, the model will use the contrastive divergence algorithm with backpropagation through time to update the weights. Here, in Figure 12 each node uses a recurrent neuron.
[0198] Figure 13 illustrates an example corresponding to the Figure 12 speed setting.
[0199] After training the model, one can provide information about physical sensors (real sensors) that one does not want to replace. These sensors help identify the current state of the system and calculate the values of missing sensors, as Figure 14 shown.
[0200] Here, the engine speed and transmission speed are measured by real sensors, and the speeds FL (corresponding to Swl) and FR (corresponding to Swr) are calculated by the RBM, as Figure 15 shown.
[0201] The advantage of this powerful algorithm is that once the model is trained, one can remove or add sensors at any time without having to retrain or reconfigure the model. This is very useful if a physical sensor fails, as then the model will remain fully operational.
[0202] As mentioned above, in theory, the BM might be the best concept to represent the system, especially its recurrent version. However, due to lack of computing power, it is difficult to implement. However, this might be easier in the future.
[0203] There are additional variants of the Boltzmann machine, such as the deep Boltzmann machine. But the intuition is the same. The only difference is that it will require more effort and resources to compute in an attempt to better generalize the given problem.
[0204] The Markov chain Monte Carlo (MCMC) algorithm inspired the Boltzmann machine. More specifically, the training algorithm "contrastive divergence" is based on Gibbs sampling, which is used in MCMC to obtain an observed sequence approximated from a specified multivariate probability distribution.
[0205] Using the Boltzmann machine approach, the trained model of a vehicle is only applicable to that vehicle. There is no guarantee that it will be applicable to other vehicles even from the same series.
[0206] Even though the BM sounds appealing, one might still prefer the Figures 4 to 9 above graph network approach because the BM is more theoretical and difficult to implement and compute.
[0207] In the above description, several terms have been used and will be defined as follows:
[0208] Lag refers to a past point of a time signal that has passed.
[0209] The maximum lag is the maximum past point that one can look at.
[0210] Optimal lag. This is a past point in time that occurs between the observation time and the maximum lag. This is the optimal lag because the sliding window from the observation time until this lag time produces the best causality value, which in turn will potentially be the best value for modeling the desired observed value.
[0211] Sliding window. This is a way of reconstructing a time series into windows with a lag size. Then the window is shifted by one step. For example, here is a time series:
[0212] Signal 78 74 22 17 82 10 23 Time T-6 T-5 T4 T-3 T-2 T-1 T .
[0213] When the lag is selected as 3 and the step is 1, the sliding window will be:
[0214] Window 1 23 10 82 Window 2 10 82 17 Window 3 82 17 22 Window 4 17 22 74 Window 5 22 74 78 .
[0215] The term "feature" is a term from machine learning. These are the inputs to an algorithm for training and fitting the predicted output (label or target). In other words, features are the input variables used in making predictions.
[0216] Label is also a term used in machine learning terminology. It is the output of the algorithm. Additionally, it is the prediction that the fitted model will produce given the features (inputs).
[0217] Hyperparameters are also used in machine learning. These are parameters whose values are set before the start of the learning process. For example, the number of neurons in the hidden layer of a neural network is a hyperparameter. Another example is the number of decision trees in a random forest.
[0218] In Figures 16 to 19 , another embodiment for determining a sensor configuration is shown, which is based on Granger neural causality.
[0219] Figure 16 is a flowchart of method 120 for determining a sensor configuration.
[0220] Method 120 includes a first step D2 that occurs after the start of the method. In Figure 16 step D2 in Figure 19In step D2', the outputs of at least a subset of the vehicle's multiple real sensors are detected and recorded. Preferably, the outputs of each real sensor of the vehicle are detected and recorded.
[0221] The recorded output of the real sensor is a sampled time series of real sensor data.
[0222] The recorded output of the real sensor ( Figure 19 X in 0 , X 1 ,... X N ) is input into a neural Granger causality (hereinafter simply referred to as "neural GC") D4 (or Figure 19 D4' in Figure 19 ). The neural GC is implemented as a component-by-component neural network, where each real sensor corresponds to one of the components of the neural GC, and where each component is formed by a virtual sensor sub-model (which itself is a neural network, as shown in
[0223] at NN in Figure 19 C in 1 , C 2 ,... C N ). Each sub-model receives and uses the outputs of each other real sensor. For example, C 2 receives the recorded outputs X 0 , X 2 ,... X N .
[0224] The neural GC is trained. In particular, each virtual sensor (sub-model) of the neural GC is trained. The virtual sensors can be trained individually or together, as described later.
[0225] The neural GC is a non-sequential neural network that branches into several internal neural networks (sub-models). Given all other sensors as input (the recorded outputs of the other real sensors except for the real sensor to be predicted by that particular sub-model), each of those sub-models can be trained individually to predict the sensor (real sensor). In an alternative method, each sub-model is trained together by summing their losses and backpropagating them to optimize the weights of the sub-model.
[0226] In other words, the neural GC is a component-by-component model, where each component can be regarded as an independent neural network, which is labeled as a sub-model or a virtual sensor (or component).
[0227] Training of the neural GC is shown in D6, where the question of whether the neural GC is suitable arises. If not, the training must be restarted ( Figure 16 the word "no" in Figure 16 ). If the neural GC is suitable and is maintained to include at least one virtual sensor that emulates a corresponding real sensor, then Figure 19 the method 120 of i proceeds to step D8, where the weights of the first layer (the first hidden layer) of each submodel are extracted. For example, in
[0228] In Figure 16 the subsequent step D10 (or Figure 19 D10' in i ), the extracted weights W Figure 19 are interpreted to extract the relevant causal relationships (similar to the causal relationships described in the earlier embodiments above). In N the causal stage is shown at 72”'. Here, for example, the causal relationship of submodel C 2 includes that X N causes X Figure 4 or Figure 7 at a value of 0.4 (shown at 75a”'). The causal relationships can be used to generate a causal relationship matrix, as shown, for example, in
[0229] the above embodiments or Figure 19 In Figure 6 or Figure 9 shown in
[0230] Here, the causal relationship vectors of each submodel will be calculated. The causal relationship vectors are then concatenated to generate a causal relationship matrix.
[0229] As described earlier, the causal relationship matrix can then be converted into a directed cyclic graph (DCG), as shown, for example, at 88”' in Figure 19 As in the earlier embodiments, the DCG can then be converted into a directed acyclic graph (DAG), as shown, for example, above in Figure 6 or Figure 9 shown in
[0230] Subsequently, at least one real sensor that forms a leaf or a root of the DAG can be determined to be replaceable and preferably replaced with a virtual sensor in the final sensor configuration of the vehicle.
[0231] In Figure 16 these final steps preferably include, in step D12, which is Figure 16 the final step before the end of the method 120 of
[0232] As described above, step D6 determines whether the neural GC is suitable. If all sub-models (virtual sensors) of the neural GC are fitted, the neural GC is fitted. A definition of "fitted" is provided below.
[0233] As previously mentioned, all sub-models preferably have the same architecture, which is basically shown in Figure 17 the flowchart of Figure 17 It is also shown how to train the corresponding sub-models.
[0234] In step T2, the sub-model receives the output of each real sensor as input (in one embodiment, except for the real sensor corresponding to the actually trained sub-model).
[0235] In steps T4 and T6, the input is split into a continuous time series (T4) and a categorical time series (T6). A categorical time series is a time series where the value at each time point is a category rather than a measurement, and the sampled values of the categorical time series can be, for example, integer values. For example, the categorical time series is the output of an ignition switch key sensor (ignition switch on or ignition switch off) or a gear number sensor.
[0236] Before the categorical time series is concatenated with the continuous time series T4 and fed to the first hidden layer, the categorical time series is transformed into their corresponding embedding layers (shown at T8). The layers of the sub-model neural network are shown at T10-1 to T10-N. The first layer T10-1 is the first hidden layer. All subsequent layers are preferably 1D convolutional layers. Such 1D convolutional layers work well for time series. However, the layers can also be recurrent or dense layers.
[0237] For the first hidden layer T10-1, it is preferred if the grouped lasso or group-ordered weighted lasso (GrOWL) regularization penalty is used to group similar features together using the parameter bundling technique, and those features that do not Granger-cause the target are zeroed out with the help of PGD (approximate gradient descent) or another sparsity-inducing optimizer.
[0238] In other words, as shown at T24-1 to T24-N, a sparsity-inducing penalty is used only for the weights T24-1 of the first hidden layer T10-1 to establish the weights of the corresponding layer.
[0239] Layers T10-1 to T10-N result in a prediction of the output of the real sensor to be simulated. This is shown at T14.
[0240] T18 is the input of the real value to be predicted / simulated (the output of the real sensor).
[0241] In T16, the loss function is calculated. In other words, the loss between the predicted value and the real value is calculated. The loss isFigure 17 shown at T20 in
[0242] The loss T20 is used to optimize the weights T24-1 to T24-N based on using a sparsity-inducing optimizer, as shown at T22 in Figure 17 shown at T22 in
[0243] The sparsity-inducing optimizer can be PGD, semi-stochastic PGD (SPGD), or FtRL ("Follow the Regularized Leader").
[0244] The approximation operator in the sparsity-inducing optimizer needs to be optimized to work with the regularization penalty. The submodel is suitable if one of the following conditions is met:
[0245] - Early stopping when the loss does not decrease after K iterations;
[0246] - Reaching the target sparsity percentage in the weights of the first hidden layer; the desired sparsity percentage is a hyperparameter of the submodel neural network.
[0247] Similarly, as shown in Figure 17 each submodel can be trained individually.
[0248] On the other hand, the submodels can be trained together, as shown in Figure 18 shown in
[0249] In this case, the loss T20 will be accumulated, as shown in T26, where the accumulated loss is used to optimize the weights using the sparsity-inducing optimizer at T28. The output of the sparsity-inducing optimizer (T28) is backpropagated to each other submodel and their corresponding weights in this case, and not just the weights T24-1 to T24-N of the current submodel.
[0250] If all submodules of the neural GC are fitted, then the neural GC is fitted.
[0251] Once the neural GC is fitted, the weights of the corresponding first layer of each submodel should be sparse (where the features assigned zeros do not Granger-cause the target (prediction) of that submodel).
[0252] To generate the causality matrix as shown in Figure 4 and Figure 7 shown in, it is preferred if the causality vectors of each submodel are calculated and then joined to generate the causality matrix.
[0253] The transformation is performed as follows:
[0254] 1. The first-layer weight matrix is converted into a kinship matrix.
[0255] 2. The affinity matrix is clustered to group similar features together.
[0256] 3. The clusters are sorted by their importance, preferably in descending order.
[0257] 4. The features within each cluster are sorted by their importance, preferably in descending order.
[0258] 5. Considering the rankings of the clusters and the rankings of the features, the global ranking of the features is calculated.
[0259] 6. Then, the global ranking is considered to be normalized and used as the causal vector for eliciting the target (the prediction of the sub-model).
[0260] In the first step, a pairwise similarity metric like cosine similarity is used to convert the weight matrix into an affinity matrix (similarity matrix). Subsequently, the features are clustered using the generated affinity matrix and any clustering algorithm that works with the affinity matrix (like the affinity propagation algorithm). In step 3, feature importance measurements like permutation tests or null tests are used to sort the clusters by importance. For example, in a permutation test, the original dataset (the output data of other real sensors recorded) - i.e., the dataset used to train the corresponding model - is randomly shuffled and fed again for prediction. The cluster that results in a higher loss means it has higher importance than the remaining clusters. Similarly, each feature is sorted.
[0261] For example, in step 5, the absolute global ranking of feature j found in cluster Pi can be calculated by the following equation (other equations can also be used):
[0262]
[0263] where:
[0264] - is the error after permuting cluster Pi.
[0265] - is the error after permuting cluster Pi without permuting feature j.
[0266] -e j : is the error after permuting feature j.
[0267] -e orig : is the original error without any permutation.
[0268] Finally, in step 6, the ranking is normalized such that all rankings sum to 1.
[0269] Based on these causal relationships or causality vectors, a causality matrix can be generated by connecting them. Based on the causality matrix, a directed cyclic graph (such as the DCG 88”' in Figure 19 can be constructed and converted into a directed acyclic graph (such as those shown in Figure 7 and Figure 8 ).
[0270] List of reference numerals:
[0271] 10 Vehicle
[0272] 12 Body
[0273] 14L, 14R Front wheels
[0274] 16L, 16R Right wheels
[0275] 18 Powertrain
[0276] 20 Internal combustion engine
[0277] 22 Clutch device
[0278] 24 Transmission
[0279] 25 Shiftable gear stage
[0280] 26 Differential
[0281] 30 Engine speed sensor (Seng)
[0282] 32 First transmission speed sensor
[0283] 34 Second transmission speed sensor (ST, Stran)
[0284] 36 Left driven wheel sensor (SL, Swl)
[0285] 38 Right driven wheel sensor (SR, Swr)
[0286] 40 Controller
[0287] 42 Engine torque sensor
[0288] 44 Clutch position sensor
[0289] 46 Temperature sensor
[0290] 46 Network unit
[0291] 48 Wireless communication
[0292] 50 Evaluation computer
[0293] 54 Window
[0294] 56 hysteresis
[0295] 58 maximum hysteresis
[0296] 60 optimal hysteresis
[0297] 70 sensor configuration process
[0298] 72 causal stage
[0299] 74 causality matrix
[0300] 76 conversion process
[0301] 78 modeling stage
[0302] 80 model construction
[0303] 82 model optimization
[0304] 84 sensor configuration
[0305] 88 directed cyclic graph
[0306] 90 directed acyclic graph
[0307] 100 Boltzmann machine
[0308] 102 visible node
[0309] 104 hidden node
[0310] 106 edge
[0311] 108 missing node (sensor)
[0312] 120 method
Claims
1. A method for determining a sensor configuration in a vehicle including a plurality of sensors, comprising the following steps: - Establish a preliminary sensor configuration for the vehicle, the sensor configuration including a first number of real sensors, each real sensor outputting a real sensor signal; - Determine whether at least one real sensor can be replaced by a virtual sensor; - Change the preliminary sensor configuration into a final sensor configuration, the final sensor configuration including a second number of real sensors and at least one virtual sensor, the at least one virtual sensor being determined to replace at least one real sensor, wherein the second number is less than the first number, wherein the determining step includes: - Detecting and recording the outputs of at least a subset of the real sensors, and - Performing a causality analysis that determines the causal relationships between the recorded outputs of the subset of real sensors, wherein the causality analysis includes constructing a component-wise neural network CWNN, wherein each real sensor of the subset of real sensors corresponds to one of the components of the CWNN, and wherein each component is formed by a virtual sensor that is a trained sub-model of the neural network to simulate the corresponding real sensor, wherein, for each trained sub-model, a causality vector is calculated, wherein calculating the causality vector of the corresponding sub-model includes: - Converting the weight matrix of the first layer of the virtual sensor into a kinship matrix, - Clustering the kinship matrix to group similar features together, - Ranking the clusters according to importance, - Ranking the features in each cluster according to importance, - Calculating the global ranking of the features by considering the ranking of the clusters and the ranking of the features, - Using the global ranking as the causality vector.
2. The method according to claim 1, wherein, the training step includes applying a sparse-inducing penalty to the corresponding first hidden layer of at least some of the virtual sensors.
3. The method according to claim 2, wherein, the sparse-inducing penalty is selected from the group lasso regularization family.
4. The method according to claim 2, wherein, the sparse-inducing penalty is selected from the group-ordered weighted lasso (GrOWL) rule family.
5. The method according to any one of claims 2 to 4, wherein, a sparse-inducing optimizer is used to optimize the sparse-inducing penalty to generate a sparse model.
6. The method according to claim 5, wherein, the semi-stochastic approximate gradient descent SPGD algorithm is used to optimize the sparse model.
7. The method according to claim 5, wherein, the Follow-the-Regularized-Leader FtRL algorithm is used to optimize the sparse model.
8. The method according to any one of claims 1 to 4, wherein, the causality vectors are concatenated to generate a causality matrix.
9. The method according to claim 1, wherein, ranking the clusters according to importance is accomplished by a permutation test method.
10. The method according to claim 1, wherein, ranking the clusters according to importance is accomplished by a zeroing method.
11. A method for determining a sensor configuration in a vehicle including a plurality of sensors, including the following steps: - Establish a preliminary sensor configuration for a vehicle, the sensor configuration including a first number of real sensors, each real sensor outputting a real sensor signal; - Determine whether at least one real sensor can be replaced by a virtual sensor; - Change the preliminary sensor configuration into a final sensor configuration, the final sensor configuration including a second number of real sensors and at least one virtual sensor, wherein the second number is less than the first number, wherein the determining step includes: - Record the real sensor signals of at least a subset of the first number of real sensors, and - Evaluate the recorded real sensor signals to determine whether at least a first real sensor can be replaced by a first virtual sensor, the first virtual sensor receiving at least one real sensor signal from a second real sensor and outputting a virtual sensor signal that simulates the real sensor signal of the first real sensor, wherein a causal relationship between the recorded outputs of the real sensors is determined for at least a subset of the samples, and wherein the causal relationships determined for the sample subsets are post-processed to determine a final set or matrix of causal relationships between the recorded outputs of the real sensors, and wherein a directed cyclic graph (DCG) is established based on the determined causal relationships, wherein the DCG is converted into a directed acyclic graph (DAG), wherein the real sensor with the highest causal relationship or the real sensor with the lowest causal relationship is taken as the root of the directed acyclic graph, and wherein at least one real sensor that forms a leaf or a root in the DAG is determined to be replaceable.
12. The method according to claim 11, wherein, the evaluating step includes using a Boltzmann machine, the Boltzmann machine having a plurality of visible nodes and having a plurality of hidden nodes, each visible node representing a real sensor, and the hidden nodes being calculated by utilizing combinations of the nodes.
13. The method according to claim 12, wherein, the Boltzmann machine is a recurrent time-limited Boltzmann machine, which is implemented by an RNN-Gaussian dynamic Boltzmann machine.
14. The method according to claim 11, wherein the determining step includes: - Detect and record the outputs of at least a subset of the real sensors for a predetermined number of temporary subsequent sampling steps, and - Perform a causal relationship analysis that determines the causal relationships between the recorded outputs of the real sensors.
Citation Information
Patent Citations
Virtual sensor system and method
US20080312756A1
Virtual sensor system and method
CN101681155A