Vehicle control method and device, computer equipment, storage medium and program product
By collecting voice signals in the vehicle and using local and cloud-based emotion recognition models for analysis, the problem of insufficient vehicle control accuracy is solved, and higher vehicle control accuracy and user experience is achieved.
Patent Information
- Application Number
- CN202510231398.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-05-27
AI Technical Summary
In the prior art, the intelligent interaction between the vehicle and the driver is processed through voice commands, and the accuracy of the recognition results is low, resulting in insufficient accuracy of vehicle control.
Voice signals are collected through the vehicle's microphone array and input them into the emotional recognition model of the local storage system and cloud server for analysis. Using local models to quickly respond and powerful computing resources and data sets of cloud models, we obtain target emotion recognition results through an arbitration mechanism, and thus perform corresponding vehicle control operations.
The accuracy of vehicle control is improved, so that the vehicle can be closer to the user's real emotional state, meet the user's current emotional needs, and enhance the safety and user experience of vehicle control.
Smart Images

Figure CN120048289A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of AI (Artificial Intelligence), and particularly to a vehicle control method, device, computer device, storage medium, and program product. Background Art
[0002] With the continuous development of artificial intelligence technology, the application of machine learning models has become more and more extensive. Especially in the intelligent interaction of new energy vehicles, drivers can achieve intelligent interaction with the vehicle through machine learning models.
[0003] In the related art, a driver can issue a voice command to the vehicle. After receiving the voice command, the vehicle processes the voice command through a machine learning model in the vehicle's whole vehicle control system, and makes a response to the driver according to the meaning of the voice command, or controls the vehicle to perform a specified operation.
[0004] However, in the solution shown in the above related art, the intelligent interaction based on the voice meaning is too single, and the accuracy of the vehicle's recognition result is relatively low, reducing the accuracy of vehicle control. Summary of the Invention
[0005] The embodiments of the present application provide a vehicle control method, device, computer device, storage medium, and program product, which can improve the accuracy of vehicle control. The technical solution is as follows:
[0006] On the one hand, a vehicle control method is provided. The method is executed by the vehicle's whole vehicle control system and includes:
[0007] Collect voice signals inside the vehicle through the vehicle's microphone array;
[0008] Input the voice signal into the first emotion recognition model of the vehicle's local storage system to obtain the first emotion recognition result output by the first emotion recognition model;
[0009] Upload the voice signal to the second emotion recognition model of the cloud server to obtain the second emotion recognition result output by the second emotion recognition model;
[0010] Perform an arbitration operation on the first emotion recognition result and the second emotion recognition result to obtain a target emotion recognition result; the target emotion recognition result is one of the first emotion recognition result and the second emotion recognition result;
[0011] According to the target emotion recognition result, perform a vehicle control operation corresponding to the target emotion recognition result.
[0012] On the other hand, a vehicle control device is provided, and the device includes:
[0013] A voice signal acquisition module, configured to acquire voice signals inside the vehicle through a microphone array of the vehicle;
[0014] A first emotion recognition result acquisition module, configured to input the voice signals into a first emotion recognition model of a local storage system of the vehicle, and acquire a first emotion recognition result output by the first emotion recognition model;
[0015] A second emotion recognition result acquisition module, configured to upload the voice signals to a second emotion recognition model of a cloud server, and acquire a second emotion recognition result output by the second emotion recognition model;
[0016] An arbitration operation execution module, configured to perform an arbitration operation on the first emotion recognition result and the second emotion recognition result, and acquire a target emotion recognition result; the target emotion recognition result is one of the first emotion recognition result and the second emotion recognition result;
[0017] A control operation execution module, configured to execute a vehicle control operation corresponding to the target emotion recognition result according to the target emotion recognition result.
[0018] In some embodiments, the arbitration operation execution module is configured to, when the first emotion recognition result is the same as the second emotion recognition result, acquire the first emotion recognition result or the second emotion recognition result as the target emotion recognition result;
[0019] When the first emotion recognition result is different from the second emotion recognition result, select the emotion recognition result with a higher confidence as the target emotion recognition result; the confidence of the first emotion recognition result is the classification probability output by the first emotion recognition model, and the confidence of the second emotion recognition result is the classification probability output by the second emotion recognition model.
[0020] In some embodiments, the arbitration operation execution module is configured to acquire, as the target emotion recognition result, the emotion recognition result that is generated first among the first emotion recognition result and the second emotion recognition result.
[0021] In some embodiments, the control operation execution module is configured to query a correspondence between the target emotion recognition result and a vehicle control operation according to the target emotion recognition result, and acquire a vehicle control operation corresponding to the target emotion recognition result.
[0022] In some embodiments, the vehicle control operation is a control operation performed on at least one of the ambient lights, fragrance, seats, speakers, and air conditioner of the vehicle.
[0023] In some embodiments, the voice signal acquisition module is configured to obtain at least two voice sub-signals inside the vehicle through the microphone array of the vehicle; the microphone array includes at least two microphones;
[0024] synthesize the at least two voice sub-signals to obtain a fused voice signal;
[0025] perform noise reduction processing on the fused voice signal;
[0026] perform echo cancellation on the fused voice signal after the noise reduction processing;
[0027] perform post-processing enhancement on the fused voice signal after the echo cancellation;
[0028] obtain the fused voice signal after the post-processing enhancement as the voice signal.
[0029] In another aspect, a computer device is provided. The computer device includes a processor and a memory. At least one instruction, at least one program, a code set, or an instruction set is stored in the memory. The at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the vehicle control method as described above.
[0030] In another aspect, a computer-readable storage medium is provided. At least one instruction, at least one program, a code set, or an instruction set is stored in the storage medium. The at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the vehicle control method as described above.
[0031] In still another aspect, a computer program product is provided. The computer program product includes a computer program, and the computer program is stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the vehicle control method provided in the above various optional implementation manners.
[0032] The technical solution provided by this application may include the following beneficial effects:
[0033] The vehicle control system collects the voice signals inside the vehicle through the microphone array deployed inside the vehicle, enabling the vehicle to obtain high-quality voice signals even in a noisy environment. The voice signals are respectively input into the emotion recognition models of the local storage system and the cloud server for analysis. Since the first emotion recognition model is deployed locally, it can respond quickly and output the first emotion recognition result quickly without relying on network connection. Since the second emotion recognition model is in the cloud, it has more powerful computing resources and updated data sets, and can output a more accurate second emotion recognition result. Based on the first emotion recognition result and the second emotion recognition result, the target emotion recognition result is obtained through an arbitration mechanism, ensuring the accuracy of the final target emotion recognition result. The target emotion recognition result can be closer to the user's true emotional state, and the subsequent vehicle control based on the target emotion recognition result can better meet the user's current emotional needs, effectively improving the accuracy of vehicle control.
[0034] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this application. Brief Description of the Drawings
[0035] The drawings herein are incorporated into the specification and constitute a part of this specification, showing the embodiments consistent with this application, and are used together with the specification to explain the principles of this application.
[0036] Figure 1 It is a system composition diagram of the vehicle control method according to an embodiment of this application;
[0037] Figure 2 It is a flowchart of a vehicle control method provided by an embodiment of this application;
[0038] Figure 3 It is a flowchart of a vehicle control method provided by an embodiment of this application;
[0039] Figure 4 It is a flowchart of a vehicle control method provided by an embodiment of this application;
[0040] Figure 5 It is a hardware structure diagram of an in-vehicle intelligent interaction system of this application;
[0041] Figure 6 It is a software framework diagram of an in-vehicle intelligent interaction system of this application;
[0042] Figure 7 It is a flowchart of an in-vehicle intelligent interaction method of this application;
[0043] Figure 8 It is a block diagram of an in-vehicle intelligent interaction device of this application;
[0044] Figure 9 is a block diagram of a vehicle control device provided by an exemplary embodiment of the present application;
[0045] Figure 10 is a schematic structural diagram of a computer device provided by an exemplary embodiment of the present application. Detailed Description of the Invention
[0046] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application.
[0047] On the contrary, they are only examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0048] Figure 1 is a system configuration diagram of a vehicle control method according to an embodiment of the present application. As Figure 1 shown, Figure 1 it includes a vehicle 100 and a cloud server 101 corresponding to the vehicle 100. The vehicle 100 includes a vehicle control system 100a, a microphone array 10, and a first emotion recognition model 11 deployed locally. The microphone array 10 is composed of multiple microphones; the cloud server 101 includes a second emotion recognition model 12.
[0049] In the embodiment of the present application, the vehicle control system 100a collects voice signals inside the vehicle 100 through the microphone array 10 of the vehicle 100; inputs the collected voice signals into the first emotion recognition model 11 of the local storage system of the vehicle 100 and the second emotion recognition model 12 in the cloud server 101 respectively to obtain a first emotion recognition result output by the first emotion recognition model 11 and a second emotion recognition result output by the second emotion recognition model 12; performs arbitration operations on the first emotion recognition result and the second emotion recognition result respectively to obtain a target emotion recognition result; the target emotion recognition result is one of the first emotion recognition result and the second emotion recognition result; according to the target emotion recognition result, a vehicle 100 control operation corresponding to the target emotion recognition result is performed.
[0050] Exemplarily, please refer to Figure 2 , Figure 2 is a flowchart of a vehicle control method provided by an exemplary embodiment of the present application. Figure 2 The vehicle control method shown can be executed by the vehicle control system of the vehicle. For example, the vehicle can be the vehicle 100 shown above Figure 1 , and the vehicle control system can be the vehicle control system 100a shown Figure 1 above.
[0051] As Figure 2 shown, the above vehicle control method may include step 210, step 220, step 230, step 240, and step 250, and the specific implementation is as follows.
[0052] Step 210: Collect voice signals inside the vehicle through the microphone array of the vehicle.
[0053] Among them, the above microphone array is a system composed of multiple microphones.
[0054] In the embodiment of the present application, the vehicle may be provided with multiple microphone arrays. One microphone array includes multiple microphones, and different microphone arrays and different microphones may be respectively arranged at different positions of the vehicle.
[0055] For example, a microphone array may be respectively arranged at the driver's seat, the passenger seat, and the rear seat of the vehicle, and the microphones in the microphone array are respectively arranged in different areas of the seat where the microphone array is located.
[0056] In the embodiment of the present application, the microphone array of the vehicle may collect voice signals in the vehicle environment in real time. The vehicle control system may obtain the voice signals collected by multiple microphones in the microphone array and fuse the voice signals collected by the multiple microphones into one voice signal.
[0057] Step 220: Input the voice signal into the first emotion recognition model of the local storage system of the vehicle to obtain the first emotion recognition result output by the first emotion recognition model.
[0058] Among them, the above first emotion recognition model is a machine learning model capable of generating a first emotion recognition result according to the voice signal.
[0059] Among them, the above first emotion recognition result may be a directly output emotion type or the confidence levels corresponding to multiple emotion types.
[0060] In the embodiment of the present application, the first emotion recognition model is set in the vehicle control system. In the case where the voice signal inside the vehicle is collected through the microphone array of the vehicle, the first emotion recognition result is directly obtained through the first emotion recognition model.
[0061] Step 230: Upload the voice signal to the second emotion recognition model of the cloud server to obtain the second emotion recognition result output by the second emotion recognition model.
[0062] Among them, the above cloud server is a high-performance computing resource located in the Internet data center, which can provide powerful computing capabilities and rich dataset support.
[0063] Among them, the above-mentioned second emotion recognition model is a machine learning model capable of obtaining a second emotion recognition result based on a voice signal.
[0064] Among them, the above-mentioned second emotion recognition result can be a directly output emotion type or the confidence levels corresponding to multiple emotion types.
[0065] In an embodiment of the present application, after obtaining a voice signal, the vehicle control system uploads the voice signal to the cloud server. The cloud server inputs the voice signal into the second emotion recognition model and sends the second emotion recognition result output by the second emotion recognition model to the vehicle control system.
[0066] Step 240: Perform an arbitration operation on the first emotion recognition result and the second emotion recognition result to obtain a target emotion recognition result; the target emotion recognition result is one of the first emotion recognition result and the second emotion recognition result.
[0067] Among them, the above-mentioned arbitration operation is a decision-making mechanism for obtaining a target emotion recognition result based on the first emotion recognition result and the second emotion recognition result.
[0068] In an embodiment of the present application, after obtaining the first emotion recognition result and the second emotion recognition result, the vehicle control system performs an arbitration operation on the first emotion recognition result and the second emotion recognition result through a preset arbitration strategy to obtain a target emotion recognition result.
[0069] Step 250: Perform a vehicle control operation corresponding to the target emotion recognition result according to the target emotion recognition result.
[0070] Based on the solutions shown in any one or more of the above-mentioned embodiments, in some embodiments, the vehicle control system collects the voice signal inside the vehicle and the sound source corresponding to the voice signal through the vehicle's microphone array; inputs the voice signal into the first emotion recognition model of the vehicle's local storage system to obtain the first emotion recognition result output by the first emotion recognition model; uploads the voice signal to the second emotion recognition model of the cloud server to obtain the second emotion recognition result output by the second emotion recognition model; performs an arbitration operation on the first emotion recognition result and the second emotion recognition result to obtain a target emotion recognition result; the target emotion recognition result is one of the first emotion recognition result and the second emotion recognition result; and performs a vehicle control operation on the seat where the sound source is located according to the target emotion recognition result and the position area of the sound source corresponding to the target emotion recognition result.
[0071] Among them, the above-mentioned position area of the sound source refers to the position area where the voice signal is emitted.
[0072] In the embodiments of the present application, different microphone arrays can be arranged in different areas of the vehicle. For example, a microphone array can be arranged on each seat.
[0073] Exemplarily, the vehicle control system collects a voice signal through the microphone array of the vehicle, identifies the sound source position corresponding to the voice signal as the co-pilot position, inputs the voice signal into the first emotion recognition model of the local storage system and the second emotion recognition model of the cloud server respectively, obtains that both the first emotion recognition result and the second emotion recognition result are "anxiety", and determines that the final target emotion recognition result is "anxiety". Based on this target emotion recognition result and the seat where the corresponding sound source is located, the vehicle control system adjusts the air conditioning temperature of the co-pilot to increase comfort, and activates the seat massage function to relieve the pressure of the passenger.
[0074] In the embodiments of the present application, the vehicle control system collects the voice signal inside the vehicle through the microphone array deployed inside the vehicle, so that the vehicle can obtain high-quality voice signals even in a noisy environment. The voice signals are respectively input into the emotion recognition models of the local storage system and the cloud server for analysis. Since the first emotion recognition model is deployed locally, it can respond quickly and output the first emotion recognition result quickly without relying on network connection. Since the second emotion recognition model is in the cloud, it has more powerful computing resources and updated data sets, and can output a more accurate second emotion recognition result. Based on the first emotion recognition result and the second emotion recognition result, the target emotion recognition result is obtained through an arbitration mechanism, which ensures the accuracy of the final target emotion recognition result. The target emotion recognition result can be closer to the user's true emotional state, and the subsequent vehicle control based on the target emotion recognition result can better meet the user's current emotional needs, effectively improving the accuracy of vehicle control.
[0075] Based on the solutions shown in any one or more of the above embodiments of the present application, please refer to Figure 3 , Figure 3 is a flowchart of a vehicle control method provided by an embodiment of the present application. The above step 240 can be implemented as step 240a and step 240b, specifically as follows.
[0076] Step 240a: When the first emotion recognition result is the same as the second emotion recognition result, obtain the first emotion recognition result or the second emotion recognition result as the target emotion recognition result.
[0077] Step 240b: When the first emotion recognition result is different from the second emotion recognition result, select the emotion recognition result with a higher confidence level as the target emotion recognition result; the confidence level of the first emotion recognition result is the classification probability output by the first emotion recognition model, and the confidence level of the second emotion recognition result is the classification probability output by the second emotion recognition model.
[0078] In an embodiment of the present application, when the first emotion recognition result is different from the second emotion recognition result, the vehicle control system respectively obtains the confidence level of the first emotion recognition result and the confidence level of the second emotion recognition result, compares these two confidence levels, and obtains the emotion recognition result corresponding to the higher confidence level as the target emotion recognition result.
[0079] In an embodiment of the present application, when any one of the confidence levels of the above-mentioned first emotion recognition result and the second emotion recognition result does not meet the confidence level threshold, the vehicle control system discards the first emotion recognition result and the second emotion recognition result to avoid inaccurate emotion recognition results caused by extremely low confidence levels in extreme cases. The confidence level threshold is a relatively low threshold, for example, 50%.
[0080] In some embodiments, when the first emotion recognition result and / or the second emotion recognition result indicate the classification probabilities corresponding to multiple emotion types, obtain the emotion recognition result corresponding to the highest classification probability among the first emotion recognition result and the second emotion recognition result as the target emotion recognition result.
[0081] In an embodiment of the present application, the above solution expands the arbitration mechanism. When the emotion recognition results (the first emotion recognition result and the second emotion recognition result) of the local and the cloud are the same, the vehicle control system directly uses any one of the results as the target emotion recognition result. When the first emotion recognition result is different from the second emotion recognition result, it is determined based on the confidence levels of the first emotion recognition result and the second emotion recognition result. Since the confidence level reflects the degree of certainty of each emotion recognition model for its own emotion recognition result, selecting the emotion recognition result with a higher confidence level as the target emotion recognition result can make full use of the respective advantages of the two emotion recognition models, improve the quality of the target emotion recognition result, and improve the accuracy of obtaining the target emotion recognition result.
[0082] Based on the solution shown in any one or more of the above embodiments of the present application, please refer to Figure 4 , Figure 4 is a flowchart of a vehicle control method provided by an embodiment of the present application. The above step 240 can be implemented as step 240c, which is specifically as follows.
[0083] Step 240c: Obtain the earliest generated emotion recognition result among the first emotion recognition result and the second emotion recognition result as the target emotion recognition result.
[0084] In the embodiment of the present application, after the vehicle control system sends the voice signal to the first emotion recognition model and the second emotion recognition model simultaneously, it obtains the first output emotion recognition result as the target emotion recognition result.
[0085] In some embodiments, when the first emotion recognition model outputs the first emotion recognition result, it simultaneously generates a timestamp indicating the generation time of the first emotion recognition result; when the second emotion recognition model outputs the second emotion recognition result, it simultaneously generates a timestamp indicating the generation time of the second emotion recognition result. The vehicle control system can compare these two timestamps and select the emotion recognition result with an earlier generation time as the target emotion recognition result. In addition to immediate generation, it is also possible to compare the two generation times after both emotion recognition results are generated and obtain the earlier generated one.
[0086] For example, if the first emotion recognition result and the second emotion recognition result are generated on the same day, the generation time of the first emotion recognition result is 10:05:03; the generation time of the second emotion recognition result is 10:06:00, the vehicle control system obtains the first emotion recognition result as the target emotion recognition result.
[0087] In some embodiments, the vehicle control system obtains the earliest generated emotion recognition result among the first emotion recognition result and the second emotion recognition result, and whose confidence level meets the specified condition as the target emotion recognition result.
[0088] Among them, the above-mentioned specified condition can be a relatively high confidence level, such as 90%.
[0089] In the embodiment of the present application, after the vehicle control system obtains the earliest generated emotion recognition result, it judges the confidence level corresponding to this emotion recognition result. When the confidence level does not meet the specified condition, it continues to wait for the second emotion recognition result to be generated to avoid inaccurate emotion recognition results caused by extreme situations.
[0090] In the embodiment of the present application, the above solution extends an arbitration strategy, that is, preferentially obtaining the earliest generated emotion recognition result as the target emotion recognition result, effectively reducing the delay in obtaining the target emotion recognition result, accelerating the response speed, reducing the computational complexity of the vehicle control system through simplified decision logic, being beneficial to reducing the power consumption of the vehicle control system, and improving the operation efficiency.
[0091] Based on the solutions shown in any one or more of the above embodiments, in some embodiments, step 250 above may be implemented as: according to the target emotion recognition result, query the corresponding relationship between the target emotion recognition result and the vehicle control operation, and obtain the vehicle control operation corresponding to the target emotion recognition result.
[0092] In the embodiments of the present application, a mapping table may be set in the vehicle control system. The mapping table records the corresponding relationship between the target emotion recognition result and the vehicle control operation. One target emotion recognition result may correspond to at least one vehicle control operation, and one vehicle control operation may correspond to at least one target emotion recognition result; after the vehicle control system obtains the target emotion recognition result, query the mapping table according to the target emotion recognition result, and obtain the vehicle control operation corresponding to the target emotion recognition result in the mapping table.
[0093] For example, the target emotion recognition result is "happy", and the vehicle control operations corresponding to this target emotion recognition result in the mapping table are to set the atmosphere light to yellow and play background music; after the vehicle control system obtains the target emotion recognition result, query and obtain the vehicle control operations corresponding to the target emotion recognition result in the mapping table, and execute the specified operations according to the vehicle control operations.
[0094] Further, in the embodiments of the present application, the vehicle control system may introduce a dynamic learning mechanism to regularly collect the user's voice signal and the corresponding target emotion recognition result of the voice signal to update the corresponding relationship between the target emotion recognition result and the vehicle control operation in the mapping table.
[0095] In the embodiments of the present application, the vehicle control system realizes the control operation of the vehicle through a pre-set mapping relationship table between the target emotion recognition result and the specific vehicle control operation, and dynamically adjusts the operation based on the emotion feedback, so that the vehicle control system can take measures in time according to the emotional state of the passengers, ensuring the accuracy of vehicle control, avoiding potential risks, and effectively improving the safety of vehicle control.
[0096] Based on the solutions shown in any one or more of the above embodiments, in some embodiments, the vehicle control operation is a control operation performed on at least one of the vehicle's atmosphere light, fragrance, seat, speaker, and air conditioner.
[0097] Optionally, the vehicle control system may adjust the atmosphere light in the vehicle according to the target emotion recognition result. For example, adjust the color and brightness.
[0098] For example, when the target emotion recognition result indicates that the user is relaxed, the vehicle control operation may be to adjust the ambient light to a soft blue or green tone; when the target emotion recognition result indicates that the user is excited, the vehicle control operation may be to adjust the ambient light to a bright orange or yellow.
[0099] Optionally, the vehicle control system can adjust the fragrance system in the vehicle based on the target emotion recognition results.
[0100] For example, where the target emotion recognition result indicates that the user is in a state of tension, the vehicle control operation may be to release a lavender-scented fragrance that has a calming effect.
[0101] Optionally, the vehicle control system can adjust the vehicle's seats based on the target emotion recognition results.
[0102] For example, when the target emotion recognition result indicates that the user is in a tired state, the vehicle control operation may be to start the massage function of the seat and appropriately adjust the backrest angle to increase support; when the target emotion recognition result indicates that the user is in an anxious state, the vehicle control operation may be to adjust the seat to a more relaxing position.
[0103] Optionally, the vehicle control system can adjust the speakers based on the target emotion recognition results.
[0104] For example, when the target emotion recognition result indicates that the user is in a happy state, the vehicle control operation may be to play cheerful music and adjust the equalizer settings to enhance the low-frequency effect; when the target emotion recognition result indicates that the user is in a tired state, the vehicle control operation may be to play soothing background music and lower the overall volume.
[0105] Optionally, the vehicle control system can adjust the vehicle's air conditioning based on the target emotion recognition results.
[0106] For example, in hot weather, the target emotion recognition result indicates that the user is in an irritable state, and the vehicle control operation may be to lower the air conditioning temperature and increase the wind speed.
[0107] In the embodiment of the present application, the above scheme expands the specific types of vehicle control operations that can be performed. The vehicle control system dynamically adjusts various devices inside the vehicle according to different target emotion recognition results, which can effectively improve the passengers' vehicle experience and comfort.
[0108] Based on the solutions shown in any one or more of the above embodiments, in some embodiments, the vehicle control system obtains at least two voice sub-signals inside the vehicle through a microphone array of the vehicle; the microphone array includes at least two microphones; synthesizes the at least two voice sub-signals to obtain a fused voice signal; performs noise reduction processing on the fused voice signal; performs echo cancellation operation on the fused voice signal after the noise reduction processing; performs post-processing enhancement on the fused voice signal after the echo cancellation operation; obtains the fused voice signal after the post-processing enhancement as the voice signal.
[0109] In the embodiments of the present application, after the vehicle control system obtains at least two voice sub-signals inside the vehicle, it divides the at least two voice confidence signals into at least two short-time frames, and performs beamforming processing on the at least two short-time frames to obtain a fused voice signal.
[0110] The vehicle control system removes engine noise in the fused voice signal through an adaptive algorithm; inputs the fused voice signal after removing engine noise into a convolutional neural network to obtain the noise-reduced fused voice signal output by the convolutional neural network, and then performs an echo cancellation operation on the fused voice signal after the noise reduction processing. After the echo cancellation operation is completed, post-processing enhancement is performed on the fused voice signal to further optimize the quality and intelligibility of the fused voice signal, obtain the final fused voice signal, and use it as the voice signal.
[0111] In the embodiments of the present application, through steps such as noise reduction, echo cancellation, and post-processing enhancement, the purity of the voice signal can be effectively guaranteed, the quality of the collected voice signal is significantly improved, and the possibility of misjudgment can be reduced through high-quality voice input, improving the accuracy of emotion recognition.
[0112] Based on the above Figures 2 to 4 steps of the embodiments, the embodiments of the present application show a design of an in-vehicle intelligent interaction system based on voice emotion recognition.
[0113] The purpose of this embodiment is to provide an in-vehicle intelligent interaction system based on voice emotion recognition. The in-vehicle intelligent interaction system can identify the emotional state of the driver by analyzing the voice characteristics of the driver, and dynamically adjust the functions of the in-vehicle system based on the recognition result.
[0114] Exemplarily, please refer to Figure 5 , Figure 5 which is a hardware structure diagram of an in-vehicle intelligent interaction system of the present application. The in-vehicle intelligent interaction system 500 includes a signal acquisition module 501, a signal processing module 502, and a policy execution module 503.
[0115] 1) Signal acquisition module 501
[0116] In the embodiment of the present application, the signal acquisition module 501 can collect the sound signals inside the vehicle through a microphone array. Since the environment inside the vehicle is complex and there are various noise sources, such as engine noise, wind noise, tire noise, and conversations among passengers, etc., the microphone array can effectively suppress noise by fusing the sound signals of multiple microphones, improve the signal-to-noise ratio of the voice signal, and thus enhance the purity of the sound data.
[0117] 2) Signal processing module 502
[0118] In the embodiment of the present application, the in-vehicle intelligent interaction system 500 can run the in-vehicle operating system and the trained voice recognition model on a board-level chip. The signal processing module 502 is responsible for performing emotion recognition on the sound signal and sending the emotion recognition result to the policy execution module.
[0119] 3) Policy execution module 503
[0120] In the embodiment of the present application, the policy execution module 503 can adjust the in-vehicle environment of the vehicle to the corresponding state according to the emotion recognition result sent by the signal processing module 502. For example, it can adjust the hardware devices such as the atmosphere light strip, fragrance, air-conditioned seat, and speaker.
[0121] In the embodiment of the present application, each hardware in the in-vehicle intelligent interaction device communicates and integrates through a bus and an SOC (System on Chip).
[0122] Exemplarily, please refer to Figure 6 , Figure 6 which is a software framework diagram of an in-vehicle intelligent interaction system of the present application. As Figure 6 shown, the in-vehicle intelligent interaction system 500 includes a voice engine 601 and an application program 602.
[0123] In the embodiment of the present application, the voice engine 601 includes a sound data module 60, a data preprocessing module 61, a local engine 62, a cloud engine 63, an arbiter 64, and an intent distributor 65.
[0124] Among them, the above-mentioned sound data module 60 is responsible for collecting sound data and converting the analog signal obtained from the microphone into a digital signal.
[0125] Among them, the above-mentioned data preprocessing module 61 performs preprocessing on the sound data, that is, performs preprocessing such as noise reduction and echo cancellation on the sound data to ensure the purity of the sound data.
[0126] Among them, the above-mentioned local engine 62 can perform emotion recognition on sounds. The local engine 62 can deploy the emotion recognition model locally and use deep learning algorithms to extract and analyze the features of sound data, solving the problem that sound emotions cannot be recognized in the absence of a network.
[0127] Among them, the above-mentioned cloud engine 63 can perform emotion recognition on sounds. By uploading the local sound data to the cloud and utilizing the more powerful computing power of the cloud, it can more accurately and quickly identify the emotion results.
[0128] Among them, the above-mentioned arbiter 64 arbitrates the emotion recognition results of the local engine 62 and the cloud engine 63, and selects the optimal result.
[0129] Among them, the above-mentioned intent distributor 65 can distribute the results of emotion recognition to each corresponding application of the vehicle for processing.
[0130] In the embodiment of the present application, the application program 602 includes an ambient light application program 66, an aroma application program 67, a seat application program 68, a music application program 69, an air conditioner application program 70, and a voice application program 71.
[0131] Among them, the above-mentioned ambient light application program 66 can adjust the ambient light of the vehicle to the corresponding color according to the emotion recognition result.
[0132] Among them, the above-mentioned aroma application program 67 can adjust the aroma in the vehicle according to the emotion recognition result.
[0133] Among them, the above-mentioned seat application program 68 can adjust the angle, height of the vehicle seat and turn on the seat massage according to the emotion recognition result.
[0134] Among them, the above-mentioned music application program 69 can adjust the music in the vehicle according to the emotion recognition result.
[0135] Among them, the above-mentioned air conditioner application program 70 can adjust the air conditioner temperature and air volume in the vehicle according to the emotion recognition result.
[0136] Among them, the above-mentioned voice application program 71 can play the corresponding greeting according to the emotion recognition result.
[0137] Exemplarily, please refer to Figure 7 , Figure 7 which is a flowchart of an in-vehicle intelligent interaction method of the present application. This method can be executed by the Figure 5 shown in-vehicle intelligent interaction system 500. As Figure 7 shown, the above-mentioned in-vehicle intelligent interaction method includes steps 710, 720, 730, and 740, which are specifically as follows.
[0138] Step 710: Sound collection
[0139] In an embodiment of the present application, the in-vehicle intelligent interaction system can collect the voice signals of the driver during vehicle driving through a microphone array installed in the vehicle.
[0140] Step 720: Data preprocessing
[0141] In an embodiment of the present application, the in-vehicle intelligent interaction system can preprocess the collected voice signals, such as noise reduction, normalization, etc., and then extract the emotional features in the voice signals through a deep learning model.
[0142] Step 730: Emotion recognition
[0143] In an embodiment of the present application, the in-vehicle intelligent interaction system can use a deep learning algorithm to train an emotion recognition model, and use the trained emotion recognition model to analyze the extracted emotional features to identify the emotional state of the driver (such as anger, tension, fatigue).
[0144] In an embodiment of the present application, based on the emotion recognition result, the in-vehicle intelligent interaction system can generate corresponding interaction strategies. For example, when the driver is emotional, it is recommended to relax or adjust the driving mode and adjust the played music.
[0145] Step 740: Execute corresponding strategies according to the emotion recognition result
[0146] In an embodiment of the present application, the in-vehicle intelligent interaction system can execute corresponding interaction strategies through the in-vehicle control unit. After the in-vehicle intelligent interaction system recognizes the emotional state, it intelligently adjusts various functions of the vehicle. For example, when it recognizes that the driver is emotional, the in-vehicle intelligent interaction system can automatically play soothing music or prompt the driver to take a deep breath to relax.
[0147] Exemplarily, please refer to Figure 8 , Figure 8 is a block diagram of an in-vehicle intelligent interaction device of the present application. As Figure 8 shown, the in-vehicle intelligent interaction device includes the following modules.
[0148] The sound data collection module 810 needs to install a microphone array (such as a dual microphone, a four-microphone) in the vehicle.
[0149] The emotion recognition module 820 requires a chip with a certain computing power to be deployed inside the vehicle, requires a large amount of sound data, manually marks each piece of sound data, and uses a deep learning algorithm to train the sound data after marking. During the training process, the model parameters are continuously tuned until a relatively satisfactory model is trained, and a high-performance sound emotion recognition model is trained and deployed in the cloud.
[0150] The emotional response module 830 requires the deployment of corresponding intelligent hardware, such as speakers, air conditioners, seats, and ambient lights. Each intelligent hardware needs to expose relevant function interfaces to the vehicle operating system for various application programs to use.
[0151] In the embodiments of the present application, through the above solutions, the intelligence level of the vehicle system can be significantly improved, enabling it not only to interact with the driver through voice, but also to understand the driver's emotional state, thereby providing more personalized services and enhancing driving safety and user experience.
[0152] Please refer to Figure 9 , which shows a block diagram of a vehicle control device provided by an exemplary embodiment of the present application. The vehicle control device can be implemented as all or part of a computer device in a hardware or software-hardware combination manner to implement all or part of the steps in the embodiments shown as above Figures 2 to 4 shown. As Figure 9 shown, the vehicle control device includes:
[0153] A voice signal acquisition module 901, configured to acquire voice signals inside the vehicle through a microphone array of the vehicle;
[0154] A first emotional recognition result acquisition module 902, configured to input the voice signal into a first emotional recognition model of the local storage system of the vehicle and obtain a first emotional recognition result output by the first emotional recognition model;
[0155] A second emotional recognition result acquisition module 903, configured to upload the voice signal to a second emotional recognition model of the cloud server and obtain a second emotional recognition result output by the second emotional recognition model;
[0156] An arbitration operation execution module 904, configured to perform an arbitration operation on the first emotional recognition result and the second emotional recognition result to obtain a target emotional recognition result; the target emotional recognition result is one of the first emotional recognition result and the second emotional recognition result;
[0157] A control operation execution module 905, configured to perform a vehicle control operation corresponding to the target emotional recognition result according to the target emotional recognition result.
[0158] In some embodiments, the arbitration operation execution module 904 is configured to obtain the first emotional recognition result or the second emotional recognition result as the target emotional recognition result when the first emotional recognition result is the same as the second emotional recognition result;
[0159] In the case where the first emotion recognition result is different from the second emotion recognition result, select the emotion recognition result with a higher confidence as the target emotion recognition result; the confidence of the first emotion recognition result is the classification probability output by the first emotion recognition model, and the confidence of the second emotion recognition result is the classification probability output by the second emotion recognition model.
[0160] In some embodiments, the arbitration operation execution module 904 is configured to obtain the emotion recognition result that is generated first among the first emotion recognition result and the second emotion recognition result as the target emotion recognition result.
[0161] In some embodiments, the control operation execution module 905 is configured to query the correspondence between the target emotion recognition result and the vehicle control operation according to the target emotion recognition result, and obtain the vehicle control operation corresponding to the target emotion recognition result.
[0162] In some embodiments, the vehicle control operation is a control operation performed on at least one of the vehicle's ambient lights, aromatherapy, seats, speakers, and air conditioner.
[0163] In some embodiments, the voice signal acquisition module 901 is configured to obtain at least two voice sub-signals inside the vehicle through the vehicle's microphone array; the microphone array includes at least two microphones;
[0164] Synthesize at least two voice sub-signals to obtain a fused voice signal;
[0165] Perform noise reduction processing on the fused voice signal;
[0166] Perform an echo cancellation operation on the fused voice signal after the noise reduction processing;
[0167] Perform post-processing enhancement on the fused voice signal after the echo cancellation operation;
[0168] Obtain the fused voice signal after the post-processing enhancement as the voice signal.
[0169] Please refer to Figure 10 , Figure 10It is a schematic structural diagram of a computer device provided by an exemplary embodiment of the present application. The computer device 1000 includes a Central Processing Unit (CPU) 1001, a system memory 1004 including a Random Access Memory (RAM) 1002 and a Read-Only Memory (ROM) 1003, and a system bus 1005 connecting the system memory 1004 and the central processing unit 1001. The computer device 1000 further includes a Basic Input / Output System (InputOutput System, I / O system) 1006 for facilitating information transmission between various components within the computer, and a mass storage device 1007 for storing an operating system 1013, application programs 1014, and other program modules 1015.
[0170] The basic input / output system 1006 includes a display 1008 for displaying information and input devices 1009 such as a mouse, keyboard, etc. for user input of information. Both the display 1008 and the input devices 1009 are connected to the central processing unit 1001 through an input / output controller 1010 connected to the system bus 1005. The basic input / output system 1006 may further include an input / output controller 1010 for receiving and processing inputs from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 1010 also provides output to a display screen, printer, or other types of output devices.
[0171] The mass storage device 1007 is connected to the central processing unit 1001 through a mass storage controller (not shown) connected to the system bus 1005. The mass storage device 1007 and its associated computer-readable medium provide non-volatile storage for the computer device 1000. That is to say, the mass storage device 1007 may include computer-readable media (not shown) such as a hard disk or a CD-ROM (Compact Disc Read-Only Memory) drive.
[0172] Without loss of generality, computer-readable media can include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes RAM (Random Access Memory), ROM (Read-Only Memory), EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), flash memory or other solid-state storage technologies, CD-ROM, DVD (Digital Video Disc), or other optical storage, magnetic tape cartridges, tapes, disk storage, or other magnetic storage devices. Of course, those skilled in the art will understand that computer storage media is not limited to the above several types. The above-mentioned system memory 1004 and mass storage device 1007 can be collectively referred to as memory.
[0173] The computer device 1000 can be connected to the Internet or other network devices through a network interface unit 1011 connected to the system bus 1005.
[0174] The memory further includes one or more programs. The one or more programs are stored in the memory, and the central processing unit 1001 implements Figures 2 to 4 all or part of the steps in the method shown.
[0175] In an exemplary embodiment, a chip is further provided. The chip includes programmable logic circuits and / or program instructions, and when the chip runs on a computer device, it is used to implement all or part of the steps of the methods shown in the above various embodiments of the present application.
[0176] In an exemplary embodiment, a computer program product is further provided. The computer program product includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor reads and executes the computer instructions to implement all or part of the steps of the methods shown in the above various embodiments of the present application.
[0177] In an exemplary embodiment, a computer-readable storage medium is further provided. A computer program is stored in the computer-readable storage medium and is loaded and executed by a processor to implement all or part of the steps of the methods shown in the foregoing various embodiments of the present application.
[0178] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The above program can be stored in a computer-readable storage medium, and the storage medium mentioned above can be a read-only memory, a magnetic disk, an optical disk, or the like.
[0179] Those skilled in the art should be able to realize that in the above one or more examples, the functions described in the embodiments of the present application can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. The computer-readable medium includes computer storage media and communication media, where the communication media includes any medium that facilitates the transfer of a computer program from one place to another. The storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0180] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A vehicle control method, characterized in that: The method is executed by a vehicle control system of the vehicle, and includes: collecting voice signals in the vehicle through a microphone array of the vehicle; Inputting the voice signal into a first emotion recognition model of a local storage system of the vehicle, and obtaining a first emotion recognition result output by the first emotion recognition model; Uploading the voice signal to a second emotion recognition model of a cloud server, and obtaining a second emotion recognition result output by the second emotion recognition model; Perform an arbitration operation on the first emotion recognition result and the second emotion recognition result to obtain a target emotion recognition result; the target emotion recognition result is one of the first emotion recognition result and the second emotion recognition result; According to the target emotion recognition result, a vehicle control operation corresponding to the target emotion recognition result is performed.
2. The method according to claim 1, characterized in that The performing an arbitration operation on the first emotion recognition result and the second emotion recognition result to obtain a target emotion recognition result includes: When the first emotion recognition result is the same as the second emotion recognition result, obtaining the first emotion recognition result or the second emotion recognition result as a target emotion recognition result; When the first emotion recognition result is different from the second emotion recognition result, the emotion recognition result with higher confidence is selected as the target emotion recognition result; the confidence of the first emotion recognition result is the classification probability of the first emotion recognition result output by the first emotion recognition model; the confidence of the second emotion recognition result is the classification probability of the second emotion recognition result output by the second emotion recognition model.
3. The method according to claim 1, characterized in that The first emotion recognition result and the second emotion recognition result perform an arbitration operation to obtain a target emotion recognition result, including: The emotion recognition result generated first between the first emotion recognition result and the second emotion recognition result is obtained as the target emotion recognition result.
4. The method according to claim 1, characterized in that: The performing a vehicle control operation corresponding to the target emotion recognition result according to the target emotion recognition result includes: According to the target emotion recognition result, the correspondence between the target emotion recognition result and the vehicle control operation is queried to obtain the vehicle control operation corresponding to the target emotion recognition result.
5. The method according to claim 4, characterized in that The vehicle control operation is a control operation performed on at least one of an ambient light, a fragrance, a seat, a speaker, and an air conditioner of the vehicle.
6. The method according to claim 1, characterized in that The collecting of the voice signal in the vehicle by means of the microphone array of the vehicle comprises: Acquiring at least two speech sub-signals in the vehicle through a microphone array of the vehicle; the microphone array comprises at least two microphones; synthesizing the at least two speech sub-signals to obtain a fused speech signal; Performing noise reduction processing on the fused speech signal; Performing an echo cancellation operation on the fused speech signal after the noise reduction process; Performing post-processing enhancement on the fused speech signal after the echo cancellation operation; The fused speech signal after the post-processing enhancement is obtained as a speech signal.
7. A vehicle control device, characterized in that: The device comprises: A voice signal acquisition module, used to collect voice signals in the vehicle through a microphone array of the vehicle; A first emotion recognition result acquisition module, used for inputting the voice signal into a first emotion recognition model of the local storage system of the vehicle, and acquiring a first emotion recognition result output by the first emotion recognition model; A second emotion recognition result acquisition module, used to upload the voice signal to a second emotion recognition model of a cloud server, and obtain a second emotion recognition result output by the second emotion recognition model; An arbitration operation execution module, used to perform an arbitration operation on the first emotion recognition result and the second emotion recognition result to obtain a target emotion recognition result; the target emotion recognition result is one of the first emotion recognition result and the second emotion recognition result; A control operation execution module is used to execute a vehicle control operation corresponding to the target emotion recognition result according to the target emotion recognition result.
8. A computer device, characterized in that: The computer device includes a processor and a memory, wherein instructions are stored in the memory, and the instructions are executed by the processor to implement the vehicle control method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The storage medium stores instructions, and the instructions are executed by a processor of a computer device to implement the vehicle control method as described in any one of claims 1 to 6.
10. A computer program product, characterized in that The computer program product includes computer instructions, which are stored in a computer-readable storage medium; the computer instructions are read and executed by a processor of a computer device to implement the vehicle control method as described in any one of claims 1 to 6.