Intelligent cockpit voice assistant adaptive voice enhancement method, device and system, computer equipment and medium
By extracting multiple environments and user feedback features in the intelligent voice assistant system, and adjusting filter parameters and speech enhancement models using deep reinforcement learning technology, the problem of low recognition accuracy in complex environments is solved, and efficient and personalized speech recognition processing is achieved.
Patent Information
- Application Number
- CN202510162134.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-14
Smart Images

Figure CN119993182A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech recognition technology, and in particular to a method, device, system, computer equipment and medium for adaptive speech enhancement of an intelligent cockpit voice assistant. Background Art
[0002] Traditional voice assistant systems typically use fixed-parameter filters and speech enhancement models to process in-car voice signals. These systems perform speech enhancement and noise suppression based on a preset ambient noise model. However, due to the complexity and variability of the in-car environment, traditional fixed-parameter methods often fail to cope with diverse noise scenarios, resulting in low speech recognition accuracy, especially in noisy environments, and a poor user experience.
[0003] To address this, some intelligent voice assistant systems have begun incorporating adaptive filtering technology. These technologies utilize ambient noise sensors to detect in-vehicle noise in real time and adjust filter parameters to enhance the voice signal. While adaptive filtering can improve voice signal quality to a certain extent, existing adaptive technologies often rely solely on the amplitude characteristics of the noise signal and lack consideration of temporal characteristics and user feedback, resulting in insufficient system adaptability and personalized services.
[0004] In short, the existing technology lacks effective adaptive and self-learning mechanisms and cannot optimize system performance in real time based on environmental changes and user feedback. Summary of the Invention
[0005] The embodiments of the present application provide a method, apparatus, system, computer equipment, and medium for adaptive speech enhancement of an intelligent cockpit voice assistant. These methods not only take into account the limitations of traditional adaptive filters, but also achieve more intelligent and personalized speech recognition processing by introducing user feedback and noise timing characteristics.
[0006] In a first aspect, the present application provides a method for adaptive voice enhancement of an intelligent cockpit voice assistant, comprising:
[0007] Extracting a current state from the current environment, wherein the current state includes: noise characteristics, current speech characteristics, filter state, and user feedback;
[0008] Select actions based on the current state and strategy to adjust filter parameters or speech enhancement model weights;
[0009] Use the adjusted speech enhancement model to process the current speech signal to obtain a recognition result. Calculate the reward based on the recognition result of the speech enhancement model and user feedback.
[0010] The current state, action, reward, and next-moment state are used as samples to train the neural network to update the strategy.
[0011] In some examples, extracting the current state from the current environment includes:
[0012] Extract the time series characteristics of the in-car noise as the noise state vector, extract the characteristics of the current speech signal, generate the speech state vector, obtain the parameter configuration of the current adaptive filter, and obtain user feedback;
[0013] The current state is composed of the noise state vector, the speech state vector, the parameter configuration of the current adaptive filter and the user feedback.
[0014] In some instances, an LSTM network is used to extract the temporal features of the in-car noise as a noise state vector; and the MFCC features of the current speech signal are extracted as a speech state vector.
[0015] In some instances, R(t)=λ1·recognition result+λ2·user feedback, where λ1 and λ2 are weight parameters.
[0016] In some examples, the present invention trains a neural network using the current state, action, reward, and next state as samples to update the strategy, including:
[0017] The current state, action, reward, and next-moment state are used as samples to train the Q network to continuously update the Q-value function, thereby adopting the best filtering parameters and speech enhancement strategies in different noise environments, so as to select the optimal action in a given current state to maximize the cumulative reward.
[0018] In some instances, Update the Q value function, where S(t) is the current state, A(t) is the action, α is the learning rate, γ is the discount factor, and A ′ is the possible action for the next step, and S(t+1) is the state at the next moment.
[0019] In a second aspect, the present application provides an adaptive voice enhancement device for an intelligent cockpit voice assistant, comprising:
[0020] A state design module is used to extract the current state from the current environment, wherein the current state includes: noise characteristics, current speech characteristics, filter state and user feedback;
[0021] An action design module, which selects actions based on the current state and strategy to adjust filter parameters or speech enhancement model weights;
[0022] The reward design module is used to process the current speech signal using the adjusted speech enhancement model to obtain a recognition result, and calculate the reward based on the recognition result of the speech enhancement model and user feedback;
[0023] The strategy optimization module is used to train the neural network using the current state, action, reward, and next-moment state as samples to update the strategy.
[0024] In a third aspect, the present application provides an intelligent cockpit system including the above-mentioned intelligent cockpit voice assistant adaptive voice enhancement device.
[0025] In a fourth aspect, the present application provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any of the above methods when executing the computer program.
[0026] In a fifth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of any of the above methods when executed by a processor.
[0027] In general, the above technical solutions conceived by the present invention can achieve the following beneficial effects compared with the prior art:
[0028] (1) Enhanced dynamic adaptability: Ability to adjust system parameters in real time to cope with complex and changing in-vehicle environments, significantly improving the accuracy of speech recognition.
[0029] (2) User-driven self-learning mechanism: Optimize the system through user feedback, improve personalized service levels, and enhance user experience.
[0030] (3) Lightweight design: Optimize algorithms to ensure real-time performance and high efficiency in embedded environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0032] Figure 1 Schematic diagram of the adaptive voice enhancement method for a smart cockpit voice assistant provided in an embodiment of the present application;
[0033] Figure 2 Schematic diagram of an adaptive voice enhancement device for a smart cockpit voice assistant provided in an embodiment of the present application;
[0034] Figure 3 This is a schematic diagram of the structure of a computer device according to an embodiment of the present application. DETAILED DESCRIPTION
[0035] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.
[0036] In the following description, the specific embodiments of the present application will be described with reference to steps and symbols performed by one or more computers, unless otherwise stated. Therefore, these steps and operations will be mentioned several times as being performed by a computer, and the computer execution referred to herein includes the operation of a computer processing unit by an electronic signal representing data in a structured form. This operation converts the data or maintains it at a location in the computer's memory system, which can be reconfigured or otherwise change the operation of the computer in a manner familiar to testers in the field. The data structure in which the data is maintained is a physical location in the memory, which has specific characteristics defined by the data format. However, the principles of the present application are described in the above text, which does not represent a limitation, and testers in the field will understand that the following various steps and operations can also be implemented in hardware.
[0037] As used herein, the terms "module" or "unit" may be considered software objects executed on the computing system. The various components, modules, engines, and services herein may be considered implementation objects on the computing system. While the apparatus and methods herein are preferably implemented in software, they may also be implemented in hardware and remain within the scope of protection of this application.
[0038] It will be understood by those skilled in the art that, unless expressly stated otherwise, the singular forms "a", "an", "" and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of this application refers to the presence of features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connections or wireless couplings. The term "and / or" as used herein includes all or any units and all combinations of one or more associated listed items.
[0039] In the first aspect, the present application provides an adaptive voice enhancement method for an intelligent cockpit voice assistant, aiming to provide an intelligent cockpit voice assistant with a voice feature enhancement technology that can adapt to complex environments in real time, such as Figure 1 As shown, the following steps are included:
[0040] S101: extracting a current state from the current environment, where the current state includes: noise features, current speech features, filter state, and user feedback;
[0041] S102: Selecting an action based on the current state and strategy to adjust filter parameters or speech enhancement model weights;
[0042] Among them, a deep Q-network can be used to select an action, that is, to adjust the filter parameters or model weights.
[0043] S103: Using the adjusted speech enhancement model to process the current speech signal to obtain a recognition result, and calculating a reward based on the recognition result of the speech enhancement model and user feedback;
[0044] S104: Use the current state, action, reward, and next moment state as samples to train the neural network to update the strategy.
[0045] Furthermore, after the current execution is completed, the next time step "t+1" is entered, and steps S101 to S104 are repeated until convergence or the specified training round is reached.
[0046] In another specific example, the state space is the basis of the algorithm, which defines the system's environmental perception capabilities at each point in time. In the embodiment of the present application, the state space consists of the following parts:
[0047] Noise characteristics: Use a long short-term memory (LSTM) network to extract the temporal characteristics of the noise inside the vehicle and generate the noise state vector N(t);
[0048] Speech features: Extract the Mel Frequency Cepstrum Coefficient (MFCC) features of the current speech signal to generate the speech state vector V(t);
[0049] Filter state: the current adaptive filter parameter configuration F(t);
[0050] User feedback: Historical feedback data U(t-1), including user feedback information at the previous moment.
[0051] Therefore, the complete state space is defined as: S(t) = {N(t), V(t), F(t), U(t-1)}.
[0052] In another specific example, the action space defines the adjustment strategies that the system can adopt at each time point. The action space of the embodiment of the present application includes:
[0053] Adjust filter parameters: such as changing filter coefficients, increasing or decreasing certain parameters to adapt to current noise characteristics;
[0054] Update the speech enhancement model: Enhance the speech feature processing capability by adjusting the weights of the LSTM network or adjusting the parameters of the feedforward network.
[0055] Therefore, the action space is expressed as: A(t) = {adjust filter parameters, update model weights}.
[0056] In another specific example, the design of the reward function directly affects the effectiveness of the reinforcement learning model. The reward function in the embodiment of the present application combines user feedback and recognition accuracy and is defined as follows: R(t) = λ1·Accuracy Improvement+λ2·User Feedback Score, where λ1 and λ2 are weight parameters used to balance the influence of the two parts.
[0057] In another specific example, in this application, a Deep Q-Network (DQN) is used to optimize the strategy, ensuring that the optimal action is selected in a given state to maximize the cumulative reward. By continuously updating the Q-value function, the system gradually learns to adopt the best filter parameter adjustment and speech enhancement strategy in different noise environments.
[0058] In another specific example, the update formula for the value is:
[0059]
[0060] Among them, α is the learning rate, γ is the discount factor, and A ′ Possible next steps.
[0061] In another specific example, at the beginning, parameters need to be initialized, including initial filter parameters, weights of the deep learning model, and the Q network of reinforcement learning.
[0062] Secondly, to facilitate better implementation of the method provided in the embodiment of the present application, the embodiment of the present application also provides a device based on the above method. The meanings of the terms are the same as in the above method, and the specific implementation details can be referred to the description in the method embodiment.
[0063] See also Figure 2 , Figure 2 This is a schematic diagram of the structure of the device provided in an embodiment of the present application, wherein the device 200 may include a state design module 201, an action design module 202, a reward design module 203 and a strategy optimization module 204, wherein:
[0064] The state design module 201 is used to extract the current state from the current environment, wherein the current state includes: noise characteristics, current speech characteristics, filter state and user feedback;
[0065] An action design module 202 for selecting actions based on the current state and strategy to adjust filter parameters or speech enhancement model weights;
[0066] A reward design module 203 is configured to process the current speech signal using the adjusted speech enhancement model to obtain a recognition result, and calculate a reward based on the recognition result of the speech enhancement model and user feedback;
[0067] The strategy optimization module 204 is used to train the neural network using the current state, action, reward, and next moment state as samples to update the strategy.
[0068] In another specific example, the state design module 201 is specifically configured to:
[0069] Extract the time series characteristics of the in-car noise as the noise state vector, extract the characteristics of the current speech signal, generate the speech state vector, obtain the parameter configuration of the current adaptive filter, and obtain user feedback;
[0070] The current state is composed of the noise state vector, the speech state vector, the parameter configuration of the current adaptive filter and the user feedback.
[0071] Specifically, the LSTM network is used to extract the temporal features of the in-car noise as the noise state vector; and the MFCC features of the current speech signal are extracted as the speech state vector.
[0072] In another specific example, R(t)=λ1·recognition result+λ2·user feedback, where λ1 and λ2 are weight parameters.
[0073] In some instances, the strategy optimization module 204 is specifically used to train the Q network using the current state, action, reward, and next-moment state as samples to continuously update the Q-value function, thereby adopting the best filtering parameters and speech enhancement strategies in different noise environments, so as to select the optimal action in a given current state to maximize the cumulative reward.
[0074] In some instances, Update the Q value function, where S(t) is the current state, A(t) is the action, α is the learning rate, γ is the discount factor, and A ′ is the possible action for the next step, and S(t+1) is the state at the next moment.
[0075] Thirdly, this application also provides a smart cockpit system including the aforementioned smart cockpit voice assistant adaptive speech enhancement device 200. The primary hardware requirements include a microphone array and an onboard processor. During the manufacturing process, the microphone array layout must be optimized to ensure effective separation of noise and speech signals, while also optimizing the computing power allocation of the onboard processor to support real-time adaptive speech feature enhancement processing.
[0076] In the embodiment of the present application, the vehicle-mounted processor may be a terminal device or a server.
[0077] In the embodiment of the present application, when the on-board processor is a server, the server can be an independent server or a server network or server cluster composed of servers. For example, the server described in the embodiment of the present application includes but is not limited to a computer, a network host, a single network server, a plurality of network server sets or a cloud server composed of multiple servers. Among them, the cloud server is composed of a large number of computers or network servers based on cloud computing. In the embodiment of the present application, communication between the server and the client can be achieved through any communication method, including but not limited to mobile communications based on the 3rd Generation Partnership Project (3GPP), Long Term Evolution (LTE), Worldwide Interoperability for Microwave Access (WiMAX), or computer network communications based on the TCP / IP Protocol Suite (TCP / IP) and User Datagram Protocol (UDP).
[0078] It is understood that when the vehicle-mounted processor used in the embodiments of the present application is a terminal device, the terminal device can be a device that includes both receiving hardware and transmitting hardware, that is, a device having receiving and transmitting hardware capable of performing two-way communication on a two-way communication link. Such terminal devices may include: cellular or other communication devices that have a single-line display or a multi-line display or a cellular or other communication device without a multi-line display. The specific vehicle-mounted processor can be a desktop terminal or a mobile terminal, and the vehicle-mounted processor can be one of a mobile phone, a tablet computer, a laptop computer, etc.
[0079] In a fourth aspect, the present application also provides a computer device, such as Figure 3 , which shows a schematic diagram of the structure of the computer device involved in the embodiment of the present application, specifically:
[0080] The computer device may include one or more processing core processors 301, one or more computer readable storage media memories 302, a power supply 303, an input unit 304 and other components. Those skilled in the art will understand that Figure 3 The computer device structure shown in the figure does not constitute a limitation on the computer device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange components differently.
[0081] Processor 301 is the control center of the computer device. It connects the various components of the entire computer device using various interfaces and lines. By running or executing software programs and / or modules stored in memory 302 and accessing data stored in memory 302, it performs various functions of the computer device and processes data, thereby providing overall monitoring of the computer device. Optionally, processor 301 may include one or more processing cores; preferably, processor 301 may integrate an application processor and a modem processor, wherein the application processor primarily handles operating storage media, user interfaces, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 301.
[0082] The memory 302 can be used to store software programs and modules. The processor 301 executes various functional applications and data processing by running the software programs and modules stored in the memory 302. The memory 302 may mainly include a program storage area and a data storage area, wherein the program storage area may store operating storage media, applications required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 302 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. Accordingly, the memory 302 may also include a controller to provide the processor 301 with access to the memory 302.
[0083] The computer device also includes a power supply 303 for supplying power to various components. Preferably, the power supply 303 can be logically connected to the processor 301 via a power management storage medium, thereby implementing functions such as charging, discharging, and power consumption management through the power management storage medium. The power supply 303 can also include one or more DC or AC power supplies, a recharge storage medium, a power failure detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components.
[0084] The computer device may further include an input unit 304 , which may be configured to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0085] Although not shown, the computer device may further include a display unit, etc., which will not be described in detail herein. Specifically, in this embodiment, the processor 301 in the computer device loads the executable files corresponding to one or more application processes into the memory 302 according to the following instructions, and the processor 301 runs the application stored in the memory 302, thereby implementing the steps in the above method embodiment.
[0086] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be accomplished by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.
[0087] To this end, an embodiment of the present application provides a computer-readable storage medium on which a computer program is stored. The computer program is loaded by a processor to execute the steps of any method provided in the embodiment of the present application.
[0088] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.
[0089] The computer-readable storage medium may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0090] Since the computer program stored in the computer-readable storage medium can execute the steps of any method provided in the embodiments of the present application, the beneficial effects that can be achieved by any method provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.
[0091] The above is a detailed introduction to the adaptive voice enhancement method, device, system, computer equipment and medium for an intelligent cockpit voice assistant provided in the embodiments of the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for technical personnel in this field, based on the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. An adaptive voice enhancement method for an intelligent cockpit voice assistant, characterized in that: include: Extracting a current state from the current moment environment, wherein the current state includes: noise characteristics, current speech characteristics, filter state, and user feedback; Select actions based on the current state and strategy to adjust filter parameters or speech enhancement model weights; Use the adjusted speech enhancement model to process the current speech signal to obtain the recognition result, and calculate the reward based on the recognition result of the speech enhancement model and user feedback; The current state, action, reward, and next state are used as samples to train the neural network to update the strategy.
2. The method according to claim 1, characterized in that The extracting the current state from the current environment includes: Extract the time series characteristics of the in-car noise as the noise state vector, extract the characteristics of the current speech signal, generate the speech state vector, obtain the parameter configuration of the current adaptive filter, and obtain user feedback; The noise state vector, the speech state vector, the parameter configuration of the current adaptive filter and the user feedback constitute the current state.
3. The method according to claim 2, characterized in that The LSTM network is used to extract the time series features of the in-car noise as the noise state vector; the MFCC features of the current speech signal are extracted as the speech state vector.
4. The method according to any one of claims 1 to 3, characterized in that: R(t)=λ1·recognition result+λ2·user feedback, where λ1 and λ2 are weight parameters.
5. The method according to claim 4, characterized in that The current state, action, reward, and next moment state are used as samples to train the neural network to update the strategy, including: The current state, action, reward and next moment state are used as samples to train the Q network to continuously update the Q value function, so as to adopt the best filtering parameters and speech enhancement strategies in different noise environments, so as to select the optimal action under the given current state to maximize the cumulative reward.
6. The method according to claim 5, characterized in that Depend on Update the Q value function, where S(t) is the current state, A(t) is the action, α is the learning rate, γ is the discount factor, and A ′ is the possible action for the next step, and S(t+1) is the state at the next moment.
7. An adaptive voice enhancement device for a smart cockpit voice assistant, characterized in that: include: A state design module, used to extract the current state from the current environment, wherein the current state includes: noise characteristics, current speech characteristics, filter state and user feedback; An action design module for selecting actions based on the current state and strategy to adjust filter parameters or speech enhancement model weights; A reward design module is used to process the current speech signal using the adjusted speech enhancement model to obtain a recognition result, and calculate a reward based on the recognition result of the speech enhancement model and user feedback; The strategy optimization module is used to train the neural network using the current state, action, reward, and next-moment state as samples to update the strategy.
8. An intelligent cockpit system comprising the intelligent cockpit voice assistant adaptive voice enhancement device according to claim 7.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method according to any one of claims 1 to 6 when executing the computer program.
10. A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Voice message processing method and device
CN115662398A
Dynamic sound filtering and classifying system and method in real-time microphone connection environment and storage medium
CN117636890A
Method and system for enhancing speech based on reinforcement learning
KR102017173B1
Condition monitoring system and method for water treatment equipment
KR1020210026332A
Real-time contextually aware artificial intelligence (AI) assistant system and a method for providing a contextualized response to a user using ai
US20240412720A1