An engine simulation method, a sound wave synthesis method, and related devices
By obtaining and predicting the driving information of new energy vehicles, obtaining equivalent working information of virtual engines, the problem of inaccurate engine simulation in the existing technology is solved and the user's driving experience is improved.
Patent Information
- Application Number
- CN202310600576.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-25
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2043-05-25
AI Technical Summary
The existing engine simulation methods are not accurate enough for the engine simulation of new energy vehicles, which affects the user's driving experience.
By obtaining the driving information of the target vehicle, including driving speed, electric gate pedal depth and brake pedal depth, and predicting based on this information, the equivalent working information of the virtual engine is obtained, including the motion state, gear state and speed data.
It improves the accuracy of new energy vehicle engine simulation and enhances the user's driving experience.
Smart Images

Figure CN116572831B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of automobiles, and particularly to an engine simulation method, a sound wave synthesis method, and related devices. Background Art
[0002] With the booming development of new energy vehicles, the power system driven by new energy has become one of the important ways in the current design of automotive power systems. Since new energy vehicles do not have a traditional fuel engine system, users cannot experience the sound wave feedback of driving a traditional fuel vehicle when driving new energy vehicles, which to a certain extent affects the driving experience of users.
[0003] The applicant of the present application has found in the long-term research and development process that the existing engine simulation methods are not accurate enough for the engine simulation of new energy vehicles. However, the sound wave simulation of new energy vehicles also depends to a large extent on the accuracy of its engine simulation. In view of this, how to improve the accuracy of the engine simulation of new energy vehicles has become an urgent problem to be solved. Summary of the Invention
[0004] The main technical problem to be solved by the present application is to provide an engine simulation method, a sound wave synthesis method, and related devices, which can improve the accuracy of the engine simulation of new energy vehicles.
[0005] To solve the above technical problem, a technical solution adopted by the present application is: to provide an engine simulation method, the method includes: obtaining driving information of a target vehicle; wherein, the driving information includes driving speed data, first depth data of an accelerator pedal, and second depth data of a brake pedal collected at the same time point, and the driving speed data, the first depth data, and the second depth data have the same target length; predicting based on the driving information of the target vehicle to obtain equivalent working information of the virtual engine of the target vehicle at the time point; wherein, the working information includes at least one of a motion state and a gear state and predicted rotational speed data with a target length.
[0006] Wherein, predicting based on the driving information of the target vehicle to obtain equivalent working information of the virtual engine of the target vehicle under the driving information includes: extracting a first feature representation including power feature information based on the first depth data, and extracting a second feature representation including braking feature information based on the second depth data, and extracting a third feature representation including vehicle speed feature information based on the driving speed data; predicting based on the first feature representation, the second feature representation, and the third feature representation to obtain the working information.
[0007] Among them, predictions are made based on the first feature representation, the second feature representation, and the third feature representation to obtain working information, including: fusing the first feature representation, the second feature representation, and the third feature representation to obtain a first fused feature; performing temporal modeling on the first fused feature to obtain a second fused feature; where the temporal modeling includes at least fine-grained modeling and coarse-grained modeling; splitting the second fused feature to obtain a splitting result, where the splitting result includes at least one of a first split feature and a second split feature and a third split feature, and the first split state contains feature information related to the motion state, the second split feature contains feature information related to the gear state, and the third split feature contains feature information related to the rotational speed; mapping each split feature in the splitting result respectively to obtain the working information.
[0008] Among them, performing temporal modeling on the first fused feature to obtain a second fused feature includes: extracting fine-grained features from the first fused feature to obtain a short-term fused feature, and processing the first fused feature based on an attention mechanism to obtain the attention weights of the elements in the short-term fused feature; where the dimension of the short-term fused feature is greater than that of the first fused feature, and the attention weight of the element represents the importance of the element for predicting the working information; multiplying each element in the short-term fused feature by the attention weight of the element to obtain a third fused feature; performing coarse-grained feature extraction on the third fused feature to obtain the second fused feature.
[0009] Among them, fusing the first feature representation, the second feature representation, and the third feature representation to obtain a first fused feature includes: concatenating the first feature representation, the second feature representation, and the third feature representation in the feature dimension to obtain the first fused feature.
[0010] Among them, splitting the second fused feature to obtain a splitting result includes: splitting the second fused feature in the feature dimension to obtain the splitting result.
[0011] Among them, obtaining the driving information of the target vehicle includes: respectively collecting the driving speed, the first depth of the accelerator pedal, and the second depth of the brake pedal from the bus information of the target vehicle at the same time point; performing time alignment on the collected driving speed, the first depth, and the second depth to obtain the driving information.
[0012] Among them, the working information is obtained by the engine simulation model predicting the driving information, and the engine simulation model is trained based on the sample driving information of the sample vehicle. The sample driving information includes the sample driving speed data, the first sample depth data of the throttle pedal, and the second sample depth data of the brake pedal when the sample vehicle is actually driving, and the sample driving information is labeled with sample working information. The sample working information includes at least one of the sample motion state of the real engine, the sample gear state, and the sample rotational speed data when the sample vehicle is actually driving.
[0013] Among them, the training steps of the engine simulation model include: predicting based on the sample driving information of the sample vehicle to obtain the sample prediction working information equivalent to the virtual engine of the sample vehicle, where the sample prediction working information includes at least one of the sample prediction motion state and the sample prediction gear state and the sample prediction rotation speed data; adjusting the model parameters of the engine simulation model according to the difference between the sample prediction working information and the sample working information.
[0014] To solve the above technical problems, another technical solution adopted by this application is: to provide a sound wave synthesis method, including: predicting the working information equivalent to the virtual engine of the target vehicle under the current driving information; where the working information is predicted based on the engine simulation method in any of the above items; synthesizing the sound wave data of the target vehicle based on the working information.
[0015] To solve the above technical problems, another technical solution adopted by this application is: to provide an engine simulation device, which includes an acquisition module and a prediction module. The acquisition module is used to acquire the driving information of the target vehicle; where the driving information includes the driving speed data, the first depth data of the throttle pedal, and the second depth data of the brake pedal collected at the same time point, and the driving speed data, the first depth data, and the second depth data have the same target length; the prediction module is used to predict based on the driving information of the target vehicle to obtain the working information equivalent to the virtual engine of the target vehicle at the time point; where the working information includes at least one of the motion state and the gear state and the prediction rotation speed data with the target length.
[0016] To solve the above technical problems, another technical solution adopted by this application is: to provide a sound wave synthesis device, which includes a prediction module and a synthesis module. The prediction module is used to predict the working information equivalent to the virtual engine of the target vehicle under the current driving information; where the working information is predicted based on the engine simulation method in any of the above items; the synthesis module is used to synthesize the sound wave data of the target vehicle based on the working information.
[0017] To solve the above technical problems, another technical solution adopted by this application is: to provide an electronic device, including a memory and a processor coupled to each other. The processor is used to execute the program data stored in the memory to implement any of the above engine simulation methods or sound wave synthesis methods.
[0018] To solve the above technical problems, another technical solution adopted by this application is: to provide a vehicle, which at least includes a vehicle body and an intelligent device carried on the vehicle body, and the intelligent device is the above electronic device.
[0019] To solve the above technical problems, another technical solution adopted in this application is: to provide a computer-readable storage medium, on which program data is stored, and when the program data is executed by a processor, any of the above engine simulation methods or sound wave synthesis methods is implemented.
[0020] In the above solution, prediction is performed using driving information including driving speed data, first depth data of the accelerator pedal, and second depth data of the brake pedal to obtain the working information equivalent to the virtual engine of the target vehicle. On the one hand, during the engine simulation process, since it is based on various aspects of driving information such as driving speed, accelerator pedal depth, and brake pedal depth, it helps to improve the prediction accuracy. On the other hand, the driving speed, accelerator pedal depth, and brake pedal depth in the driving information are ensured to be collected at the same time point and have the same target length, which can further improve the temporal consistency of the driving speed, accelerator pedal depth, and brake pedal depth in the driving information, so that temporal information can be further introduced for prediction, which helps to further improve the prediction accuracy. On the other hand, since the obtained working information includes at least one of the motion state and the gear state and the predicted rotational speed with the same target length, it can accurately characterize the state of the virtual engine from multiple aspects. Therefore, the accuracy of engine simulation can be improved. Description of the Drawings
[0021] Figure 1 is a flowchart of an embodiment of the engine simulation method of this application;
[0022] Figure 2 is a flowchart of another embodiment of step S120 of this application;
[0023] Figure 3 is a flowchart of another embodiment of step S222 of this application;
[0024] Figure 4 is a flowchart of another embodiment of the engine simulation method of this application;
[0025] Figure 5 is a flowchart of an embodiment of the sound wave synthesis method of this application;
[0026] Figure 6 is a framework diagram of an embodiment of the engine simulation device of this application;
[0027] Figure 7 is a framework diagram of an embodiment of the sound wave synthesis device of this application;
[0028] Figure 8 is a framework diagram of an embodiment of the electronic device of this application;
[0029] Figure 9 is a framework diagram of an embodiment of the computer-readable storage medium of this application;
[0030] Figure 10 It is a schematic diagram of the framework of an embodiment of the vehicle in this application. Detailed implementation manners
[0031] To make the objectives, technical solutions and effects of this application clearer and more definite, the following further describes this application in detail with reference to the accompanying drawings and by way of examples. In the following description, specific details such as specific system architectures, interfaces, technologies, etc. are presented for the purpose of illustration rather than limitation, so as to thoroughly understand this application.
[0032] The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after. In addition, "multiple" in this article means two or more than two. In addition, the term "at least one" in this article represents any one of multiple or any combination of at least two of multiple. For example, including at least one of A, B, and C can represent including any one or more elements selected from the set composed of A, B, and C.
[0033] Please refer to Figure 1 , Figure 1 It is a schematic flowchart of an embodiment of the engine simulation method in this application. Specifically, the method may include the following steps:
[0034] Step S110: Obtain the driving information of the target vehicle.
[0035] It should be noted that the engine simulation method provided in this application can be applied to the target vehicle, and the target vehicle can be a new energy vehicle (such as an electric vehicle, a hydrogen energy vehicle, etc.), and is used to obtain the equivalent working information of the virtual fuel engine of the target vehicle. For example, for a certain electric vehicle, the equivalent working information of the virtual engine corresponding to the electric vehicle under the current driving information can be obtained through engine simulation.
[0036] In addition, the equivalent working information of the virtual engine obtained by using the engine simulation method can be used in the processing work related to the target vehicle. For example, it can be used for engine engine sound simulation, comparing the target vehicle with a fuel vehicle, and serving as a reference for data processing related to the target vehicle, etc.
[0037] Among them, the driving information is collected for the target vehicle and may include the driving speed data, the first depth data of the accelerator pedal, and the second depth data of the brake pedal collected at the same time point.
[0038] In a specific application scenario, the driving information of the target vehicle can be collected according to a preset sampling frequency. For example, the driving speed, the first depth of the accelerator pedal, and the second depth of the brake pedal are collected at several time points within a period of time, thereby forming sequence data. It should be noted that the period of time for collecting the above driving information can be the period of time corresponding to an audio frame, so that the working information equivalent to the time points of the target vehicle's virtual engine within this period of time can be used to synthesize the sound wave data corresponding to an audio frame.
[0039] The collected driving speed data, first depth data, and second depth data are all discrete sequences, and the three can have the same target length, that is, they are collected at several time points within the same period of time. The sequence length refers to the number of elements contained in the sequence data, which is the same as the number of time points included in this period of time, and the sequence length can be any positive integer. Further, the elements at the same position in different sequences are collected at the same time point. For example, the driving speed, the first depth of the accelerator pedal, and the second depth of the brake pedal are collected at n time points within a period of time, obtaining three sequence data with a length of n. Sequence one is the driving speed data sequence, including A1, A2, A3, …, A n , sequence two is the first depth sequence of the accelerator pedal, including B1, B2, B3, …, B n , sequence three is the second depth sequence of the brake pedal, including C1, C2, C3, …, C n , where the elements at the same position in different sequences are collected at the same time point. Specifically, for example, A1, B1, and C1 are collected at the same time point.
[0040] In a specific application scenario, the driving information of the target vehicle can be collected within a time period between 120 ms and 300 ms, the sampling frequency can be set to 8 kHz, and the sequence length of the driving speed data, first depth data, and second depth data is denoted as L, and the value of L can be between 960 and 2400. Further, the typical value of L can be 1024 or 2048.
[0041] In some embodiments, the driving information may further include other data, such as driving road condition data, vehicle attribute data, and so on.
[0042] In some embodiments, obtaining the driving information of the target vehicle may include: respectively collecting the driving speed, the first depth of the accelerator pedal, and the second depth of the brake pedal from the bus information of the target vehicle at the same time point. Perform time alignment on the collected driving speed, first depth, and second depth to obtain the driving information.
[0043] In some embodiments, before making a prediction, the driving information of the target vehicle can also be preprocessed, and the preprocessing can include, but is not limited to, normalization processing.
[0044] Step S120: Based on the driving information of the target vehicle, make a prediction to obtain the equivalent working information of the virtual engine of the target vehicle at the time point.
[0045] Among them, the working information characterizes the equivalent working condition of the virtual engine at the above time point, and can include at least one of the motion state and the gear state and the predicted rotational speed data with a target length. The motion state can include three types: acceleration, deceleration, and idling. The gear state represents the gear in which the engine is working and can include several preset gear types.
[0046] The predicted rotational speed data represents the rotational speed data at the above acquisition time point. Since the driving information is a discrete sequence sampled at several time points within a time period, the predicted rotational speed data is also a sequence of the same length, representing the predicted rotational speeds at each time point. It should be noted that for the sake of reducing the amount of computation and other considerations, a single state value can be used to represent the motion state / gear state of the virtual engine during the time period corresponding to the acquisition of the driving information.
[0047] In a specific application scenario, the gear state can include neutral and gears 1-6, a total of seven types.
[0048] It can be understood that the equivalent working information of the virtual engine corresponds to the driving state of the target vehicle at the above time point, that is, when the virtual engine is in the working state of this equivalent working information, the driving state of the vehicle is the same as the driving state of the target vehicle at the above time point. The driving information of the target vehicle characterizes the driving state of the target vehicle at the above time point. Then, based on the driving information of the target vehicle, the working information of the virtual engine in this driving state can be predicted, and thus the equivalent working information of the virtual engine is obtained.
[0049] Furthermore, the driving information including the driving speed data of the target vehicle, the first depth data of the accelerator pedal, and the second depth data of the brake pedal can characterize the driving state of the target vehicle at the above time point. According to the driving state of the target vehicle at the above time point, the working information of the virtual engine in this driving state can be predicted. Specifically, it can include at least one of the motion state and the gear state, as well as the predicted rotational speed data.
[0050] In the above solution, on the one hand, during the engine simulation process, since it is based on various aspects such as driving speed, accelerator pedal depth, and brake pedal depth to characterize driving information, it helps to improve the prediction accuracy. On the other hand, the driving speed, accelerator pedal depth, and brake pedal depth in the driving information are ensured to be collected at the same time point and have the same target length, which can further improve the temporal consistency of the driving speed, accelerator pedal depth, and brake pedal depth in the driving information. Thus, temporal information can be further introduced for prediction, which helps to further improve the prediction accuracy. On the other hand, since the predicted working information includes at least one of the motion state and gear state and the predicted rotational speed with the same target length, it can accurately characterize the state of the virtual engine from multiple aspects. Therefore, the accuracy of engine simulation can be improved.
[0051] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of another embodiment of step S120 of the present application. Specifically, step S120 may include the following steps:
[0052] Step S221: Based on the first depth data, extract the first feature representation containing power feature information, based on the second depth data, extract the second feature representation containing braking feature information, and based on the driving speed data, extract the third feature representation containing vehicle speed features.
[0053] Among them, the first depth data is the depth of the accelerator pedal, which can reflect the power situation of the target vehicle. Based on the first depth data, the first feature representation containing the power feature information of the target vehicle can be extracted. Similarly, the second depth data is the depth of the brake pedal, which can reflect the braking situation of the target vehicle. Based on the second depth data, the second feature representation containing the braking feature information of the target vehicle can be extracted. Based on the driving speed data, the third feature representation containing vehicle speed feature information can be extracted.
[0054] Step S222: Based on the first feature representation, the second feature representation, and the third feature representation, perform a prediction to obtain the working information.
[0055] It should be noted that the first feature representation, the second feature representation, and the third feature representation respectively characterize the driving state of the target vehicle from the three aspects of power, braking, and vehicle speed. Based on these three aspects of feature representations, a prediction can be made to obtain the equivalent working information of the virtual engine in this driving state.
[0056] In some embodiments, since power and braking are closely related to the power situation of the vehicle's forward movement, it is also possible to first fuse the first feature representation and the second feature representation, and then combine them with the third feature representation for prediction.
[0057] Please refer to Figure 3 ,Figure 3 It is a schematic flowchart of another embodiment of step S222 of the present application. Specifically, step S222 may include the following steps:
[0058] Step S3221: Fuse the first feature representation, the second feature representation, and the third feature representation to obtain a first fused feature.
[0059] It should be noted that the fusion method can be diverse. For example, the first feature representation, the second feature representation, and the third feature representation can be fused according to corresponding fusion weights.
[0060] In some embodiments, fusing the first feature representation, the second feature representation, and the third feature representation may be to splice the first feature representation, the second feature representation, and the third feature representation in the feature dimension to obtain a first fused feature.
[0061] Step S3222: Perform temporal modeling based on the first fused feature to obtain a second fused feature.
[0062] Among them, temporal modeling at least includes fine-grained modeling and coarse-grained modeling.
[0063] Furthermore, since the driving information includes driving information corresponding to several time points, the fine-grained modeling can be based on each time point, and the coarse-grained modeling can be based on between several time points.
[0064] In some embodiments, performing temporal modeling based on the first fused feature to obtain a second fused feature includes: extracting fine-grained features from the first fused feature to obtain a short-term fused feature, and processing the first fused feature based on the attention mechanism to obtain the attention weights of each element in the short-term fused feature, multiplying each element in the short-term fused feature by the attention weight of the element to obtain a third fused feature, and extracting coarse-grained features from the third fused feature to obtain a second fused feature.
[0065] Among them, the dimension of the short-term fused feature is greater than that of the first fused feature, and the attention weight of the element represents the importance degree of the element to the predicted work information.
[0066] In a specific application scenario, the fine-grained feature extraction can be implemented by a convolutional network, and the coarse-grained feature extraction can be implemented by a long short-term memory network.
[0067] Step S3223: Split the second fused feature to obtain a split result.
[0068] It should be noted that the first feature representation, the second feature representation, and the third feature representation used for fusion are not in one-to-one correspondence with the split result.
[0069] Among them, the splitting result includes at least one of the first splitting feature and the second splitting feature, and also includes the third splitting feature. The first splitting feature contains feature information related to the motion state, which can be used to obtain the motion state as working information. The second splitting feature contains feature information related to the gear state, which can be used to obtain the gear state as working information. The third splitting feature contains feature information related to the rotational speed, which can be used to obtain rotational speed data as working information.
[0070] In some embodiments, the splitting result includes the first splitting feature and the third splitting feature. In some embodiments, the splitting result includes the second splitting feature and the third splitting feature. In some embodiments, the splitting result includes the first splitting feature, the second splitting feature, and the third splitting feature.
[0071] In some embodiments, splitting the second fusion feature may be splitting the second fusion feature in the feature dimension to obtain a splitting result.
[0072] Step S3224: Map each splitting feature in the splitting result respectively to obtain working information.
[0073] In some embodiments, mapping the first splitting feature obtains the motion state. The value range of the mapping result is 0, 1, 2, representing acceleration, deceleration, and idling respectively.
[0074] In some embodiments, mapping the second splitting feature obtains the gear state. The value range of the mapping result is 0 - 6, representing neutral gear and gears 1 - 6 respectively.
[0075] In some embodiments, mapping the third splitting feature obtains predicted rotational speed data with a target length.
[0076] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of another embodiment of the engine simulation method of the present application.
[0077] As Figure 4 shown, the working information equivalent to the virtual engine of the target vehicle can be obtained by predicting the driving information with the engine simulation model. Among them, the engine simulation model is trained based on the sample driving information of the sample vehicle.
[0078] Specifically, the engine simulation model is trained with the sample driving information of the sample vehicle. After the model training is completed, the engine simulation model is used to predict the driving information of the target vehicle to obtain the working information equivalent to the virtual engine of the target vehicle.
[0079] Among them, the sample driving information includes the sample driving speed data of the sample vehicle during actual driving, the first sample depth data of the accelerator pedal, and the second sample depth data of the brake pedal, and the sample driving information is marked with sample working information, where the sample working information includes at least one of the sample motion state of the real engine and the sample gear state of the sample vehicle during actual driving, and also includes sample rotation speed data.
[0080] Furthermore, the training steps of the engine simulation model include: making a prediction based on the sample driving information of the sample vehicle to obtain the sample predicted working information equivalent to the virtual engine of the sample vehicle, and adjusting the model parameters of the engine simulation model according to the difference between the sample predicted working information and the sample working information. Among them, the sample predicted working information includes at least one of the sample predicted motion state and the sample predicted gear state, and also includes sample predicted rotation speed data. Among them, adjusting the model parameters of the engine simulation model according to the difference between the sample predicted working information and the sample working information can be based on at least one of the first difference between the sample predicted motion state and the sample motion state, the second difference between the sample predicted gear state and the sample gear state, and the third difference between the sample predicted rotation speed data and the sample rotation speed data to adjust the model parameters.
[0081] It should be noted that during the model training process, the sample vehicle is a fuel vehicle, and its driving information includes the first sample depth data of the accelerator pedal. During the process of using the model for inference, the target vehicle is a new energy vehicle, and its driving information includes the first depth data of the throttle pedal, corresponding to the first sample depth data.
[0082] In some embodiments, during the training of the engine simulation model, the sample driving information of the sample vehicle is input into the engine simulation model, including the sample driving speed data, the first sample depth data of the accelerator pedal, and the second sample depth data of the brake pedal. The sample driving information is marked with sample working information, including at least one of the sample motion state of the real engine and the sample gear state of the sample vehicle during actual driving and the sample rotation speed data.
[0083] Specifically, the training data can be obtained by real-time acquiring the real motion information of the sample vehicle during actual driving and the working state information of the real fuel engine through the CAN signal (bus signal) of the in-vehicle OBD system (On Board Diagnostics, an in-vehicle system for monitoring and reporting vehicle status).
[0084] The real motion information during the driving of the sample vehicle includes but is not limited to the accelerator pedal depth, the brake pedal depth, and the actual running speed. The working state information of the real fuel engine includes but is not limited to the acceleration, deceleration, or idle state, the engine working gear, and the engine working speed.
[0085] The sample driving information of the sample vehicle used in one training can be intercepted from the throttle pedal data sequence (SeqPedal), the brake pedal data sequence (SeqBrake), and the vehicle speed data sequence (SeqSpeed) in the training data. Specifically, several time points are intercepted with the same sequence length L, and the three sequence data obtained after interception are time-aligned and used as input data.
[0086] Preprocessing is performed on the three input sequences respectively, where the preprocessing can include normalization. The three normalized input sequences are input into the engine simulation model.
[0087] The engine simulation model can include an extraction module, a fusion module, a time series modeling module, and a splitting module. Among them, the time series modeling module includes a fine-grained modeling module and a coarse-grained modeling module.
[0088] In some embodiments, the engine simulation model may further include a mapping module.
[0089] In this embodiment, taking the FuseBN module (fusion BN module) as an example of the extraction module for illustration, the normalized data sequence passes through the fusion BN layer (FuseBN module) to obtain high-dimensional feature parameters f1, f2, and f3 respectively.
[0090] The fusion BN module mainly includes three sub-modules: a Conv convolutional layer, a BN batch normalization layer, and an Act activation layer. The specific operations implemented by the fusion BN module are shown in the following formula:
[0091]
[0092] Among them, x and y respectively represent the featureized input and output vectors, μ is the mean within a Batch, is the variance within a Batch, and γ and β are the parameters that can be learned by this fusion network. During the training process, learning is performed by means of gradient descent.
[0093] In a specific application scenario, considering each training data sequence, it is necessary to be able to include as many virtual engine ignition frequency cycles in the full frequency band as possible. Therefore, the time length value corresponding to the input sequence is set to be between 120 ms and 300 ms. Corresponding to an 8 kHz sampling rate, the value of the sequence length L is between 960 and 2400, and the typical value of L is preferably 1024 or 2048.
[0094] The convolutional kernels in the fusion BN layer can be set to different output dimensions for the three input features of the throttle sequence, brake sequence, and speed sequence respectively. The consideration for setting different output dimensions is that the fuel injection amount (i.e., output power) corresponding to the throttle sequence and the change in engine gear speed corresponding to the speed sequence have a more significant high-dimensional feature mapping for the time series prediction of the final speed. A typical preset parameter setting is that when the input of the convolutional kernel is all L, the output dimensions of the convolutional kernels corresponding to f1, f2, and f3 are 150, 50, and 800 respectively.
[0095] In this embodiment, taking the feature parameter concatenation module (Concatenate) as an example of the fusion module, f1, f2, and f3 are concatenated with feature parameters to obtain the feature parameter f with superimposed dimensions. all 。
[0096] In a specific application scenario, the corresponding feature parameter concatenation (Concatenate) obtains the feature parameter f with superimposed dimensions. all The output dimension is 1000.
[0097] In this embodiment, taking the fine-grained modeling module as an example of the convolutional module based on the attention mechanism, the fine-grained modeling module includes a convolutional branch and an attention branch.
[0098] For the feature parameter f all , it is respectively input into the convolutional branch and the attention branch. Fine-grained feature extraction is performed through the convolutional branch to obtain the short-term fusion feature, and the attention weights of each element in the short-term fusion feature are obtained through the attention branch. Multiply the short-term fusion feature obtained by the convolutional branch by the attention weights obtained by the attention branch to obtain the short-term fusion feature after weight adjustment.
[0099] In a specific application scenario, the attention weights and the short-term fusion feature can be multiplied in an element-wise multiplication manner.
[0100] For the vehicle model of the sample vehicle, usually under the corresponding relationship of specific throttle, brake, and real-time speed, the upshift, downshift, deceleration, or idle state of the engine is triggered. At this time, the mapping ability of the high-dimensional time series features to the output label is often more important. Therefore, through the attention mechanism, the convolutional branch can better learn the amplitude expression of the best speed signal in a fine-grained manner for different time series features (motion state and gear state). Specifically, the attention branch designs a larger output overlapping step to expand the receptive field, and directly converts the high-dimensional feature f all sequence into query elements, key elements, and value elements, performs attention calculation internally, and captures the internal time series feature correlation to judge the importance degree of this local time series feature.
[0101] The corresponding formula is expressed as follows:
[0102]
[0103] Among them, represents the result of the convolutional branch output, represents the result of the attention branch output, represents element-wise multiplication, is used for input into the coarse-grained modeling module. The above convolutional module based on the attention mechanism can more effectively achieve subsequence feature extraction, thereby removing interference information.
[0104] In this embodiment, the coarse-grained modeling module is taken as an LSTM (Long Short-Term Memory) module as an example for illustration.
[0105] The calculation of the long short-term memory module is as follows:
[0106]
[0107] In the above formulas, W xi 、W xf 、W xo 、W xc respectively represent the weight vectors from the input layer to the input gate, forget gate, output gate, and cell state; W hi 、W hf 、W ho 、W hc respectively represent the weight vectors from the hidden layer to the input gate, forget gate, output gate, and cell state; b i 、b f 、b o 、b c respectively represent the offsets of the input gate, forget gate, output gate, and cell state; σ and tanh respectively represent the corresponding activation functions. is the f’ all output by the coarse-grained modeling module. Through the LSTM, the high-dimensional features of short sequences are integrated for time series prediction to learn the time series expression of the optimal rotational speed signal.
[0108] In a specific application scenario, based on the output dimensions of the convolutional module and the long short-term memory module based on the attention mechanism and the final corresponding model effects, the overall is related to the input overlapping step size, attention depth, convolutional depth, time step size, and depth of the LSTM. A typical preset parameter setting is: for the high-dimensional feature parameter f with a dimension of 1000 all, the convolutional branch and the parallel attention branch can achieve the abstraction and extraction of short - term sequence features with an output dimension of 2000. Further, the LSTM module further extracts coarse - grained features from the fine - grained high - dimensional features extracted from the previous segment for feature classification of gear states and motion states and time - series prediction of rotational speed, corresponding to the feature parameter f’ all The output dimension is set to 600.
[0109] In this embodiment, the splitting module is used to split the second fusion feature, and the splitting result is obtained by splitting the second fusion feature in the feature dimension.
[0110] In a specific application scenario, the high - dimensional feature parameter f’ all is dimensionally split to obtain output sub - feature parameter sequences of dimensions M, N, and W respectively. In particular, a typical preset parameter for M, N, and W is set to 44, 44, and 512.
[0111] The above - mentioned sub - feature parameter sequences are respectively mapped to obtain the classification of gear states and motion states and the regression prediction of the time series of rotational speed.
[0112] In a specific application scenario, three features with different dimensions are used as the input of the fully - connected layer, and the classification result of the gear state with an output dimension of 1, the classification result of the motion state with a dimension of 1, and the regression prediction result of the time series of rotational speed with a dimension of L (corresponding to the selection of sampling time points of the training data, typically taking values of 1024 or 2048) are obtained respectively.
[0113] Specifically, the three final output data are respectively: motion state (StateMotion, dimension 1x1, value range 0, 1, 2, representing acceleration, deceleration, and idle speed respectively); gear state (StateGear, dimension 1x1, value range 0 - 6, representing neutral gear and gears 1 - 6 respectively); virtual engine rotational speed sequence data (dataRPM, dimension 1xL, sequence elements are normalized to values < 1.0f, representing the rotational speed sequence data within the corresponding time - frame length L).
[0114] During the process of using the engine simulation model for inference, the input to the engine model is the driving information of the target vehicle, including driving speed data, the first depth data of the throttle pedal, and the second depth data of the brake pedal. The processing process can refer to the relevant description of the foregoing embodiment.
[0115] Using the driving information of different vehicle models as training data, the working information of virtual engines for different vehicle models can be trained to meet the needs of engine simulation for different vehicle models. By adopting a deep neural network model, a complete and effective simulation of the virtual engine for a specific vehicle or vehicle model is achieved.
[0116] Please refer toFigure 5 , Figure 5 is a schematic flowchart of an embodiment of the engine sound synthesis method of the present application. Specifically, the engine sound synthesis method may include the following steps:
[0117] Step S510: Predict the equivalent working information of the virtual engine of the target vehicle under the current driving information.
[0118] Among them, the above working information can be predicted by using any of the engine simulation methods in the foregoing embodiments.
[0119] Step S520: Synthesize the engine sound data of the target vehicle based on the working information.
[0120] Among them, the working information includes at least one of a motion state and a gear state and predicted rotational speed data with a target length.
[0121] It should be noted that during driving, the working information of the target vehicle can be continuously predicted, and then the engine sound data of the target vehicle can be synthesized according to the working information obtained by each prediction. In a specific application scenario, during driving, taking the duration corresponding to an audio frame as the unit duration of the time period for collecting driving information, the driving information of the target vehicle in the current time period is obtained. Among them, the current time period may include several sampling time points, and the equivalent working information in the current time period is predicted, and then the engine sound data of the target vehicle in the current time period is synthesized based on the working information in the current time period, so as to synthesize the engine sound of the vehicle during driving to provide the user with an engine sound feedback equivalent to driving a traditional fuel vehicle.
[0122] It should be noted that the method for synthesizing the engine sound may include, but is not limited to, an order synthesis scheme, a wavetable synthesis scheme, and a particle synthesis scheme.
[0123] Taking the synthesis of the engine sound by using the particle synthesis scheme as an example, synthesizing the engine sound data of the target vehicle based on the working information may include: selecting matching audio particles from the audio particle set as the target audio particles based on the working information, and then synthesizing the engine sound data of the target vehicle based on the target audio particles that have been cached currently. It should be noted that the audio particle set may include audio particle subsets corresponding to different motion states. On this basis, specifically, the corresponding audio particle subset can be selected based on the motion state, the target audio particles can be selected from the audio particle subset based on the gear state and the predicted rotational speed data, and the engine sound data of the target vehicle can be synthesized based on the target audio particles that have been cached currently.
[0124] By improving the accuracy of engine simulation, accurately characterize the working state of the virtual engine, so as to be able to distinguish different motion states of acceleration, deceleration and idling, as well as different gear states, and combine with the vehicle speed to accurately synthesize the engine sound, improving the accuracy and authenticity of the engine sound simulation and enhancing the effect of the engine sound simulation.
[0125] Please refer to Figure 6 , Figure 6 which is a schematic framework diagram of an embodiment of the engine simulation device of the present application.
[0126] In this embodiment, the engine simulation device 60 includes an acquisition module 61 and a prediction module 62. The acquisition module 61 is used to acquire the driving information of the target vehicle. Among them, the driving information includes the driving speed data, the first depth data of the accelerator pedal, and the second depth data of the brake pedal collected at the same time point, and the driving speed data, the first depth data, and the second depth data have the same target length. The prediction module 62 is used to make a prediction based on the driving information of the target vehicle to obtain the equivalent working information of the virtual engine of the target vehicle at the time point. Among them, the working information includes at least one of the motion state and the gear state and the predicted rotational speed data with the target length.
[0127] Among them, the prediction module 62 includes an extraction sub-module and a prediction sub-module. The extraction sub-module is used to extract a first feature representation containing power characteristic information based on the first depth data, extract a second feature representation containing braking characteristic information based on the second depth data, and extract a third feature representation containing vehicle speed characteristic information based on the driving speed data. The prediction sub-module is used to make a prediction based on the first feature representation, the second feature representation, and the third feature representation to obtain the working information.
[0128] Among them, the prediction sub-module includes a fusion unit, a modeling unit, a splitting unit, and a mapping unit. The fusion unit is used to fuse the first feature representation, the second feature representation, and the third feature representation to obtain a first fusion feature. The modeling unit is used to perform time series modeling based on the first fusion feature to obtain a second fusion feature. Among them, the time series modeling includes at least fine-grained modeling and coarse-grained modeling. The splitting unit is used to split the second fusion feature to obtain a splitting result, where the splitting result includes at least one of a first splitting feature and a second splitting feature and a third splitting feature, and the first splitting feature contains feature information related to the motion state, the second splitting feature contains feature information related to the gear state, and the third splitting feature contains feature information related to the rotational speed. The mapping unit is used to map each splitting feature in the splitting result respectively to obtain the working information.
[0129] Among them, the modeling unit includes a fine-grained extraction subunit, an attention subunit, a fusion subunit, and a coarse-grained extraction subunit. The fine-grained extraction subunit is used to perform fine-grained feature extraction on the first fusion feature to obtain a short-term fusion feature. The attention subunit is used to process the first fusion feature based on the attention mechanism to obtain the attention weights of each element in the short-term fusion feature. Among them, the dimension of the short-term fusion feature is greater than that of the first fusion feature, and the attention weight of the element represents the importance of the element to the prediction work information. The fusion subunit is used to multiply each element in the short-term fusion feature by the attention weight of the element to obtain a third fusion feature. The coarse-grained extraction subunit is used to perform coarse-grained feature extraction on the third fusion feature to obtain a second fusion feature.
[0130] Among them, the fusion unit is used to fuse the first feature representation, the second feature representation, and the third feature representation to obtain a first fusion feature, specifically including: splicing the first feature representation, the second feature representation, and the third feature representation in the feature dimension to obtain a first fusion feature.
[0131] Among them, the splitting unit is used to split the second fusion feature to obtain a splitting result, specifically including: splitting the second fusion feature in the feature dimension to obtain a splitting result.
[0132] Among them, the acquisition module 61 includes an acquisition sub-module and an alignment sub-module. The acquisition sub-module is used to respectively acquire the driving speed, the first depth of the accelerator pedal, and the second depth of the brake pedal from the bus information of the target vehicle at the same time point. The alignment sub-module is used to perform time alignment on the acquired driving speed, first depth, and second depth to obtain driving information.
[0133] Among them, the work information is obtained by the engine simulation model predicting the driving information. The engine simulation model is trained based on the sample driving information of the sample vehicle. The sample driving information includes the sample driving speed data, the first sample depth data of the throttle pedal, and the second sample depth data of the brake pedal when the sample vehicle is actually driving, and the sample driving information is labeled with sample work information. The sample work information includes at least one of the sample motion state of the real engine, the sample gear state, and the sample rotation speed data when the sample vehicle is actually driving.
[0134] Among them, the engine simulation device further includes a training module, which is used to predict based on the sample driving information of the sample vehicle to obtain the sample prediction work information equivalent to the virtual engine of the sample vehicle. Among them, the sample prediction work information includes at least one of the sample prediction motion state and the sample prediction gear state and the sample prediction rotation speed data; according to the difference between the sample prediction work information and the sample work information, adjust the model parameters of the engine simulation model.
[0135] Please refer to Figure 7 , Figure 7It is a schematic framework diagram of an embodiment of the sound wave synthesis device of the present application.
[0136] In this embodiment, the sound wave synthesis device 70 includes a prediction module 71 and a synthesis module 72. The prediction module 71 is used to predict the equivalent working information of the virtual engine of the target vehicle under the current driving information; wherein, the working information is predicted based on the engine simulation method in any one of the above; the synthesis module 72 is used to synthesize the sound wave data of the target vehicle based on the working information.
[0137] Please refer to Figure 8 , Figure 8 It is a schematic framework diagram of an embodiment of the electronic device of the present application.
[0138] In this embodiment, the electronic device 80 includes a memory 81 and a processor 82, wherein the memory 81 is coupled to the processor 82. Specifically, the various components of the electronic device 80 can be coupled together through a bus, or the processor 82 of the electronic device 80 is respectively connected to other components one by one. The electronic device 80 can be any device with processing capabilities, such as a computer, a tablet computer, a mobile phone, etc.
[0139] The memory 81 is used to store the program data executed by the processor 82 and the data during the processing of the processor 82, etc. For example, vehicle speed data, motion state, etc. Among them, the memory 81 includes a non-volatile storage part for storing the above program data.
[0140] The processor 82 controls the operation of the electronic device 80. The processor 82 can also be called a CPU (Central Processing Unit, central processing unit). The processor 82 may be an integrated circuit chip with signal processing capabilities. The processor 82 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In addition, the processor 82 can be implemented jointly by multiple integrated circuit chips.
[0141] The processor 82 is used to execute instructions to implement any one of the above engine simulation methods or sound wave synthesis methods by calling the program data stored in the memory 81.
[0142] Please refer to Figure 9 , Figure 9 It is a schematic framework diagram of an embodiment of the computer-readable storage medium of the present application.
[0143] In this embodiment, the computer-readable storage medium 90 stores program data 91 that can be run by a processor. The program data 91 can be executed to implement any of the above engine simulation methods or sound synthesis methods.
[0144] Specifically, the computer-readable storage medium 90 can be a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc., which can store program data. Or it can also be a server that stores the program data 91. The server can send the stored program data to other devices for running, or it can also run the stored program data by itself.
[0145] In some embodiments, the computer-readable storage medium 90 can also be a memory as Figure 8 shown.
[0146] Please refer to Figure 10 , Figure 10 which is a schematic framework diagram of an embodiment of a vehicle in this application.
[0147] In this embodiment, the vehicle 100 includes a vehicle body 101 and an intelligent device 102 carried on the vehicle body. Among them, the intelligent device can be Figure 8 the electronic device shown, which can implement any of the above engine simulation methods or sound synthesis methods. The relevant descriptions can refer to the relevant content of the foregoing embodiments and will not be elaborated here.
[0148] If the technical solution of this application involves personal information, before the product applying the technical solution of this application processes personal information, it has clearly informed the personal information processing rules and obtained the personal's autonomous consent. If the technical solution of this application involves sensitive personal information, before the product applying the technical solution of this application processes sensitive personal information, it has obtained the personal's separate consent and at the same time meets the requirement of "express consent". For example, at a personal information collection device such as a camera, a clear and prominent sign is set to inform that the personal information collection range has been entered and personal information will be collected. If the personal voluntarily enters the collection range, it is regarded as consenting to the collection of their personal information; or on the personal information processing device, when the personal information processing rules are informed by obvious signs / information, personal authorization is obtained through pop-up information or asking the personal to upload their personal information by themselves, etc.; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.
[0149] The above are only the embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present application.
Claims
1. An engine simulation method, characterized in that, The method includes: Obtaining the driving information of the target vehicle; wherein, the driving information includes driving speed data, first depth data of the accelerator pedal, and second depth data of the brake pedal collected at the same time point within a time period, and the driving speed data, the first depth data, and the second depth data have the same target length, and the time period is a time period corresponding to an audio frame; Using an engine simulation model to extract a first feature representation containing power feature information based on the first depth data, and extract a second feature representation containing braking feature information based on the second depth data, and extract a third feature representation containing vehicle speed feature information based on the driving speed data; the dimension of the third feature representation and the dimension of the first feature representation are both greater than the dimension of the second feature representation; Fusing the first feature representation, the second feature representation, and the third feature representation to obtain a first fused feature; Performing temporal modeling based on the first fused feature to obtain a second fused feature; wherein, the temporal modeling includes at least fine-grained modeling and coarse-grained modeling; Splitting the second fused feature to obtain a splitting result, wherein the splitting result includes a first split feature, a second split feature, and a third split feature, and the first split feature contains feature information related to the motion state, the second split feature contains feature information related to the gear state, and the third split feature contains feature information related to the rotational speed; Mapping each split feature in the splitting result respectively to obtain the equivalent working information of the virtual engine of the target vehicle at the time point; wherein, the working information includes the motion state, the gear state, and predicted rotational speed data with the target length, and the motion state includes three types: acceleration, deceleration, and idling.
2. The method according to claim 1, characterized in that, The performing temporal modeling based on the first fused feature to obtain a second fused feature includes: Performing fine-grained feature extraction on the first fused feature to obtain a short-term fused feature, and processing the first fused feature based on an attention mechanism to obtain the attention weights of each element in the short-term fused feature; wherein, the dimension of the short-term fused feature is greater than the dimension of the first fused feature, and the attention weight of the element represents the importance of the element for predicting the working information; Multiplying each element in the short-term fused feature by the attention weight of the element to obtain a third fused feature; Performing coarse-grained feature extraction on the third fused feature to obtain the second fused feature.
3. The method according to claim 1, characterized in that, The fusing the first feature representation, the second feature representation, and the third feature representation to obtain a first fused feature includes: Concatenating the first feature representation, the second feature representation, and the third feature representation in the feature dimension to obtain the first fused feature; And / or, the splitting the second fused feature to obtain a splitting result includes: Splitting the second fused feature in the feature dimension to obtain the splitting result.
4. The method according to claim 1, characterized in that, The obtaining the driving information of the target vehicle includes: Collect the driving speed, the first depth of the accelerator pedal, and the second depth of the brake pedal from the bus information of the target vehicle at the same time point respectively. Perform time alignment on the collected driving speed, first depth, and second depth to obtain the driving information.
5. The method according to claim 1, characterized in that, The engine simulation model is trained based on the sample driving information of the sample vehicle. The sample driving information includes the sample driving speed data, the first sample depth data of the accelerator pedal, and the second sample depth data of the brake pedal when the sample vehicle is actually driving. And the sample driving information is labeled with sample working information, and the sample working information includes at least one of the sample motion state of the real engine, the sample gear state, and the sample rotation speed data when the sample vehicle is actually driving.
6. The method according to claim 5, characterized in that, The training steps of the engine simulation model include: Perform prediction based on the sample driving information of the sample vehicle to obtain the sample prediction working information equivalent to the virtual engine of the sample vehicle. Wherein, the sample prediction working information includes at least one of the sample prediction motion state, the sample prediction gear state, and the sample prediction rotation speed data. Adjust the model parameters of the engine simulation model according to the difference between the sample prediction working information and the sample working information.
7. A sound wave synthesis method, characterized in that, Include: Predict the working information equivalent to the virtual engine of the target vehicle under the current driving information; wherein, the working information is predicted based on the engine simulation method according to any one of claims 1 to 6. Based on the working information, synthesize the sound wave data of the target vehicle.
8. An electronic device, characterized in that, Includes a memory and a processor coupled to each other. The processor is configured to execute the program data stored in the memory to implement the engine simulation method according to any one of claims 1 to 6 or the sound wave synthesis method according to claim 7.
9. A vehicle, characterized in that, The vehicle at least includes: A vehicle body; An intelligent device carried on the vehicle body, and the intelligent device is the electronic device according to claim 8.
10. A computer-readable storage medium, on which program data is stored, characterized in that, When the program data is executed by the processor, it implements the engine simulation method according to any one of claims 1 to 6 or the sound wave synthesis method according to claim 7.
Citation Information
Patent Citations
Method for providing information relating to the operational state of a motor vehicle to a driver and motor vehicle having a control unit for carrying out the method
CN103717452A
Method and system for simulating engine sounding during automobile acceleration and deceleration
CN111731185A
Electric vehicle active sounding method and system based on shifting strategy migration
CN112298031A
Automobile engine sound real-time synthesis system and method based on deep learning
CN112652315A
Audio particle extraction method and device, sound synthesis method and device, equipment and medium
CN115662470A