A vehicle-mounted speech enhancement method, device, storage medium and equipment
By acquiring in-vehicle auxiliary information and using a voice enhancement model to process in-vehicle voice, the problems of low wake-up rate and positioning errors under conditions of high-speed vehicle driving or interference from multiple people are solved, thus improving the effect of in-vehicle voice interaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- IFLYTEK CO LTD
- Filing Date
- 2022-10-13
- Publication Date
- 2026-05-05
AI Technical Summary
In low signal-to-noise ratio scenarios such as high-speed vehicle travel or interference from multiple speakers, the wake-up rate of in-vehicle voice interaction systems is low and the positioning is incorrect, resulting in a poor user voice interaction experience.
By acquiring in-vehicle auxiliary information such as seat information, vehicle speed information, and window status information, Fourier transform and speech enhancement models are used to enhance the target speech information, predict the speech information weights of each voice region, and perform multiplication calculations and inverse Fourier transforms to improve the speech signal-to-noise ratio.
It improves the wake-up, positioning, and recognition effects while the vehicle is in motion, enhancing the user's voice interaction experience.
Smart Images

Figure CN115641861B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech processing technology, and in particular to a method, apparatus, storage medium and device for in-vehicle speech enhancement. Background Technology
[0002] With the improvement of people's living standards and the rapid development of the social economy, the usage rate of cars has gradually increased, and more and more cars have entered people's lives, bringing great convenience to all aspects of life. Among them, voice interaction systems have also become widespread in smart cars.
[0003] Currently, vehicles are typically divided into multiple audio zones based on seat positions. For example, a four-seater vehicle, including the driver's seat, front passenger seat, rear passenger seat, and rear front passenger seat, is divided into four audio zones. Then, directional beamforming or speech separation models are used to obtain the speaker audio corresponding to each audio zone and send it to the backend for wake-up. The wake-up results are then compared to determine the location of the wake-up person, thus enabling subsequent recognition of the target speaker. However, in scenarios with low signal-to-noise ratios, such as high-speed driving with open windows or interference from multiple speakers, wake-up rates become low, and even after wake-up, location errors may occur, resulting in a poor voice interaction experience for in-vehicle users. Summary of the Invention
[0004] The main objective of this application is to provide an in-vehicle voice enhancement method, apparatus, storage medium, and device that can enhance the speaker's voice based on in-vehicle auxiliary information, thereby improving wake-up, positioning, and recognition effects, and ultimately enhancing the user's voice interaction experience while driving.
[0005] This application provides an in-vehicle voice enhancement method, including:
[0006] Acquire the vehicle assistance information of the target vehicle, and acquire the target voice information of the vehicle users in each audio zone of the target vehicle;
[0007] The target speech information is enhanced using the vehicle-mounted auxiliary information to obtain enhanced target speech information.
[0008] Based on the enhanced target voice information, preset operation processing is performed on the in-vehicle user and / or target vehicle to obtain the processing result.
[0009] In one possible implementation, the vehicle assistance information of the target vehicle includes the seat information of the target vehicle; the step of using the vehicle assistance information to enhance the target voice information to obtain enhanced target voice information includes:
[0010] The target speech information is subjected to Fourier transform to obtain the transformed target speech information;
[0011] A combined vector is constructed using the seat information of the target vehicle and the converted target speech information, and the combined vector is input into a pre-constructed speech enhancement model to predict the weights of the target speech information in each voice region of the target vehicle.
[0012] The weights are multiplied by the corresponding target speech information, and the results are subjected to inverse Fourier transform to obtain the enhanced target speech information.
[0013] In one possible implementation, constructing a combined vector using the seat information of the target vehicle and the converted target speech information includes:
[0014] The vector corresponding to the seat information of the target vehicle and the vector corresponding to the converted target speech information are concatenated to obtain a combined vector; or, a combined vector is constructed using the seat information of the target vehicle and the converted target speech information through gating.
[0015] In one possible implementation, the vehicle assistance information of the target vehicle includes the vehicle speed information and window status information; the step of using the vehicle assistance information to enhance the target voice information to obtain enhanced target voice information includes:
[0016] The target speech information is subjected to Fourier transform to obtain the transformed target speech information;
[0017] A combined vector is constructed using the vehicle speed information and window status information of the target vehicle and the converted target speech information. The combined vector is then input into a pre-built speech enhancement model to predict the weights of the target speech information in each vocal range of the target vehicle.
[0018] The weights are multiplied by the corresponding target speech information, and the results are subjected to inverse Fourier transform to obtain the enhanced target speech information.
[0019] In one possible implementation, constructing a combined vector using the target vehicle's speed information, window status information, and the converted target speech information includes:
[0020] The vector corresponding to the vehicle speed information of the target vehicle, the vector corresponding to the window status information, and the vector corresponding to the converted target voice information are concatenated to obtain a combined vector; or, a combined vector is constructed using the vehicle speed information, window status information, and the converted target voice information through gating.
[0021] In one possible implementation, the speech enhancement model includes at least one of a convolutional neural network (CNN), a recurrent neural network (RNN), or a real or complex network.
[0022] In one possible implementation, inputting the combined vector into a pre-built speech enhancement model to predict the weights of the target speech information for each register on the target vehicle includes:
[0023] The combined vector is input into a pre-built speech enhancement model, and the ratio between the frequency domain signal of the target speech information collected in each voice region and the frequency domain signal with noise frequency collected in the preset voice region is calculated as the weight of the target speech information in each voice region.
[0024] In one possible implementation, the step of performing preset operation processing on the in-vehicle user and / or target vehicle based on the enhanced target voice information to obtain a processing result includes:
[0025] Based on the enhanced target voice information, the preset device of the target vehicle is activated, and the vehicle user who issued the activation voice is located and identified to obtain the processing result.
[0026] This application also provides an in-vehicle voice enhancement device, including:
[0027] The acquisition unit is used to acquire the vehicle assistance information of the target vehicle and the target voice information of the vehicle users in each audio zone of the target vehicle.
[0028] The enhancement unit is used to enhance the target speech information using the vehicle-mounted auxiliary information to obtain enhanced target speech information;
[0029] The processing unit is used to perform preset operation processing on the in-vehicle user and / or target vehicle based on the enhanced target voice information, and obtain the processing result.
[0030] In one possible implementation, the onboard assistance information of the target vehicle includes the seat information of the target vehicle; the enhancement unit includes:
[0031] The first transformation subunit is used to perform a Fourier transform on the target speech information to obtain the transformed target speech information;
[0032] The first prediction subunit is used to construct a combined vector using the seat information of the target vehicle and the converted target speech information, and input the combined vector into a pre-built speech enhancement model to predict the weights of the target speech information in each voice region of the target vehicle.
[0033] The first calculation subunit is used to multiply the weights with the corresponding target speech information respectively, and perform inverse Fourier transform on the calculation results to obtain the enhanced target speech information.
[0034] In one possible implementation, the first prediction subunit is specifically used for:
[0035] The vector corresponding to the seat information of the target vehicle and the vector corresponding to the converted target speech information are concatenated to obtain a combined vector; or, a combined vector is constructed using the seat information of the target vehicle and the converted target speech information through gating.
[0036] In one possible implementation, the vehicle assistance information of the target vehicle includes the vehicle speed information and window status information; the enhancement unit includes:
[0037] The second transformation subunit is used to perform Fourier transform on the target speech information to obtain the transformed target speech information.
[0038] The second prediction subunit is used to construct a combined vector using the vehicle speed information and window status information of the target vehicle and the converted target speech information, and input the combined vector into a pre-built speech enhancement model to predict the weights of the target speech information in each voice region of the target vehicle.
[0039] The second calculation subunit is used to multiply the weights with the corresponding target speech information respectively, and perform inverse Fourier transform on the calculation results to obtain the enhanced target speech information.
[0040] In one possible implementation, the second prediction subunit is specifically used for:
[0041] The vector corresponding to the vehicle speed information of the target vehicle, the vector corresponding to the window status information, and the vector corresponding to the converted target voice information are concatenated to obtain a combined vector; or, a combined vector is constructed using the vehicle speed information, window status information, and the converted target voice information through gating.
[0042] In one possible implementation, the speech enhancement model includes at least one of a convolutional neural network (CNN), a recurrent neural network (RNN), or a real or complex network.
[0043] In one possible implementation, the first prediction subunit or the second prediction subunit is specifically used for:
[0044] The combined vector is input into a pre-built speech enhancement model, and the ratio between the frequency domain signal of the target speech information collected in each voice region and the frequency domain signal with noise frequency collected in the preset voice region is calculated as the weight of the target speech information in each voice region.
[0045] In one possible implementation, the processing unit is specifically used for:
[0046] Based on the enhanced target voice information, the preset device of the target vehicle is activated, and the vehicle user who issued the activation voice is located and identified to obtain the processing result.
[0047] This application also provides an in-vehicle voice enhancement device, including: a processor, a memory, and a system bus;
[0048] The processor and the memory are connected via the system bus;
[0049] The memory is used to store one or more programs, the one or more programs including instructions, which, when executed by the processor, cause the processor to perform any of the above-described implementations of the in-vehicle voice enhancement method.
[0050] This application also provides a computer-readable storage medium storing instructions that, when executed on a terminal device, cause the terminal device to perform any of the above-described implementations of the in-vehicle voice enhancement method.
[0051] This application also provides a computer program product, which, when run on a terminal device, causes the terminal device to execute any of the above-described implementations of the in-vehicle voice enhancement method.
[0052] This application provides a method, apparatus, storage medium, and device for enhancing in-vehicle voice. First, it acquires in-vehicle auxiliary information of the target vehicle and target voice information of in-vehicle users in each voice zone of the target vehicle. Then, it uses the in-vehicle auxiliary information to enhance the target voice information, obtaining enhanced target voice information. Next, based on the enhanced target voice information, it performs preset operation processing on the in-vehicle users and / or the target vehicle to obtain the processing result. It is evident that this application first enhances the voice of in-vehicle users in each voice zone of the vehicle based on the in-vehicle auxiliary information, and then uses the enhanced voice for subsequent preset operation processing such as vehicle wake-up, user location, and recognition. This improves the wake-up, location, and recognition effects, thereby enhancing the user's voice interaction experience while the target vehicle is in motion. Attached Figure Description
[0053] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0054] Figure 1 A schematic diagram illustrating the division of a four-seater vehicle, including a driver's seat, a front passenger seat, a rear passenger seat, and a rear passenger seat, into four frequency bands, provided for an embodiment of this application;
[0055] Figure 2 A flowchart illustrating an in-vehicle voice enhancement method provided in an embodiment of this application;
[0056] Figure 3 This is a schematic diagram illustrating the composition of an in-vehicle voice enhancement device provided in an embodiment of this application. Detailed Implementation
[0057] With the continuous development of technology, voice interaction systems have become widespread in smart cars. To achieve better information interaction between multiple speakers and voice devices within the vehicle, the vehicle is typically divided into multiple sound zones based on seating positions. For example, a four-seater car, including the driver's seat, front passenger seat, rear passenger seat, and rear front passenger seat, is divided into four sound zones. Figure 1 As shown, the system then uses directional beamforming or a speech separation model to obtain speaker audio from multiple audio regions for backend wake-up. By comparing the wake-up results, the location of the person waking up in the vehicle is determined, thus enabling subsequent identification of the target speaker. However, in scenarios with low signal-to-noise ratios, such as high-speed driving with open windows or interference from multiple speakers, low wake-up rates occur. Even when wake-up is successful, incorrect speaker location may occur, leading to a poor voice interaction experience for in-vehicle users.
[0058] To address the aforementioned shortcomings, this application provides an in-vehicle voice enhancement method. First, it acquires the in-vehicle auxiliary information of the target vehicle and the target voice information of in-vehicle users in each voice zone of the target vehicle. Then, it uses the in-vehicle auxiliary information to enhance the target voice information, obtaining enhanced target voice information. Next, based on the enhanced target voice information, it performs preset operation processing on the in-vehicle user and / or the target vehicle to obtain the processing result. As can be seen, this application first enhances the voice of in-vehicle users in each voice zone of the vehicle based on the in-vehicle auxiliary information, and then uses the enhanced voice to perform subsequent preset operation processing such as vehicle wake-up, user location, and recognition. This improves the wake-up, location, and recognition effects, thereby enhancing the user's voice interaction experience while the target vehicle is in motion.
[0059] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0060] First Embodiment
[0061] See Figure 2 This is a flowchart illustrating an in-vehicle voice enhancement method provided in this embodiment. The method includes the following steps:
[0062] S201: Obtain the vehicle assistance information of the target vehicle, and obtain the target voice information of the vehicle users in each voice zone of the target vehicle.
[0063] In this embodiment, any vehicle requiring in-vehicle voice enhancement is defined as the target vehicle, and the voice information emitted by in-vehicle users in each voice zone of the target vehicle is defined as the target voice information to be enhanced. It should be noted that this embodiment does not limit the language type of the target voice information; for example, the target voice can be composed of Chinese or English. Furthermore, this embodiment does not limit the length of the target voice; for example, the target voice can be a sentence or a paragraph.
[0064] In practical applications, in-vehicle voice interaction scenarios, compared to home interaction scenarios, are characterized by fixed user locations (such as the driver's seat, passenger seat, etc.) and the number of users being related to the vehicle type (such as four-seater, seven-seater, etc.). During vehicle operation, the main sources of noise are usually wind noise and tire noise generated while driving. Furthermore, different noise levels occur at different vehicle speeds and are related to the window status; for example, when the windows are closed, the vehicle achieves a certain degree of passive noise reduction.
[0065] Therefore, when enhancing the target voice information of the in-vehicle user, this application fully considers the in-vehicle assistance information of the target vehicle. This in-vehicle assistance information includes, but is not limited to, the target vehicle's seat information, speed information, and window status information. Furthermore, after obtaining the target vehicle's in-vehicle assistance information, existing or future vector conversion methods can be used to convert the target vehicle's in-vehicle assistance information into a corresponding in-vehicle assistance information representation vector. The specific format of this representation vector can be set according to the actual situation; this embodiment does not limit it. For example, the representation vector can be a 4-dimensional vector, etc., for executing the subsequent step S202.
[0066] It should be noted that this application does not limit the method of acquiring the target voice information of in-vehicle users in each audio zone of the target vehicle. For example, using... Figure 1 Taking a target vehicle with four audio zones as an example, the vehicle seats include the driver's seat, the front passenger seat, the rear passenger seat, and the rear passenger seat. A microphone can be pre-installed on the roof corresponding to each seat to collect the voice information emitted by the vehicle users in each audio zone, which will serve as the target voice information for the vehicle users in that audio zone.
[0067] It should also be noted that this application does not limit the specific content of the vehicle assistance information of the target vehicle or the method of obtaining it. The corresponding method of obtaining the information can be selected according to the content of the vehicle assistance information. For example, when the vehicle assistance information of the target vehicle includes the seat information of the target vehicle, the seat information of the target vehicle can be obtained by monitoring the seat belt wearing status, using a pressure sensor to detect the seat status, or obtaining human activity information at the corresponding position through a human infrared detection sensor or a vehicle camera.
[0068] S202: Enhance the target speech information using vehicle-mounted auxiliary information to obtain enhanced target speech information.
[0069] In this embodiment, after obtaining the vehicle assistance information of the target vehicle and the target voice information of the vehicle users in each voice zone of the target vehicle through step S201, in order to improve the voice interaction experience of the vehicle users, the vehicle assistance information can be used to enhance the target voice information to obtain the enhanced target voice information, which is then used to execute the subsequent step S203.
[0070] Specifically, one possible implementation is that the vehicle-mounted auxiliary information of the target vehicle may include the seat information of the target vehicle. In this case, the implementation process of step S202 may include: first, performing a Fourier transform on the target speech information to obtain the transformed target speech information; then, constructing a combined vector using the seat information of the target vehicle and the transformed target speech information, and inputting the combined vector into a pre-built speech enhancement model to predict the weights of the target speech information in each voice region of the target vehicle; next, multiplying each weight with the target speech information of the corresponding voice region, and performing an inverse Fourier transform on the calculated result to obtain the enhanced target speech information, which is then used to execute the subsequent step S203.
[0071] In this implementation, there are no restrictions on the method of acquiring user data and seat information within the target vehicle. Furthermore, after acquiring the seat information, it can be converted into a representation vector of the corresponding dimension based on the vehicle's audio register division. For example, the seat information of a four-seater vehicle can be converted into a 4-dimensional representation vector, or the seat information of a seven-seater vehicle can be converted into a 7-dimensional representation vector. The values of each dimension represent the seat information in each audio register (i.e., whether there are any in-vehicle users).
[0072] For example: Figure 1 Taking a target vehicle with four sound zones as an example, if the seat belt wearing status is monitored, the seat status is detected by pressure sensors, and human activity information at the corresponding position is obtained by human infrared detection sensors or vehicle cameras, and it is determined that the occupants of the target vehicle are located in the driver's seat and the front passenger seat, then the seat information of the target vehicle can be converted into the corresponding 4-dimensional representation vector: Information_position = [1,1,0,0]. Similarly, when it is determined that the occupants are located in the driver's seat and the rear passenger seat, the seat information of the target vehicle can be converted into the corresponding 4-dimensional representation vector: Information_position = [1,0,1,0].
[0073] Based on this, in order to improve the voice interaction experience of in-vehicle users, after obtaining the target voice information of in-vehicle users in each voice zone of the target vehicle, Fourier transform can be performed on it to obtain the transformed target voice information in the frequency domain.
[0074] For example: still using Figure 1 Taking a target vehicle with four audio zones as an example, assuming that a microphone is pre-installed on the roof corresponding to each of the four audio zones—driver's seat, front passenger seat, rear passenger seat, and rear passenger seat—we can define them as microphone 1, microphone 2, microphone 3, and microphone 4, respectively. Then, the noisy target voice information of the in-vehicle user in each audio zone can be further defined as follows:
[0075] yi=si1+si2+si3+si4+ni (1)
[0076] Where yi represents the target voice information of the vehicle user collected by the i-th microphone in the i-th sound zone; si1, si2, si3, and si4 represent the signals of the voices emitted by the four vehicle users in the four sound zones of the driver's seat, passenger seat, driver's rear seat, and passenger's rear seat, respectively, reaching the i-th microphone; and ni represents the signal of noise reaching the i-th microphone.
[0077] For example, y1 = s11 + s12 + s13 + s14 + n1 represents the target voice information of the in-vehicle user collected by the microphone in the driver's seat area. Similarly, y2 = s21 + s22 + s23 + s24 + n2, y3 = s31 + s32 + s33 + s34 + n3, and y4 = s41 + s42 + s43 + s44 + n4 represent the target voice information of the in-vehicle user collected by the microphones in the respective areas of the passenger seat, the driver's rear seat, and the passenger's rear seat.
[0078] The speech information represented by the above formula is then subjected to Fourier transform to obtain the transformed target speech information in the frequency domain, as shown in the following formula (2):
[0079] Yi = Si1 + Si2 + Si3 + Si4 + Ni (2)
[0080] Furthermore, existing or future vector combination methods can be utilized to fuse the 4D representation vector corresponding to the target vehicle's seat information with the transformed frequency domain target speech information (i.e., Yi) to construct a combined vector, which is defined as a feature. Specifically, the vector corresponding to the target vehicle's seat information and the vector corresponding to the transformed target speech information can be concatenated to obtain the combined vector as: feature = [Y1, Y2, Y3, Y4, Information_position]; or, gating can be applied to the input or intermediate layers of the speech enhancement model to guide the model to better separate speech by utilizing both the amplitude and phase differences between microphones and seat-assisted information, thereby achieving vector combination of the target vehicle's seat information and the transformed target speech information.
[0081] The obtained combined vector (feature) is then input into the pre-built speech enhancement model to calculate the ratio between the frequency domain signal of the target speech information collected in each voice region and the frequency domain signal with noise collected in the preset voice region. This ratio is used as the weight of the target speech information in each voice region, thereby predicting the weight of the target speech information in the four voice regions of the target vehicle. This weight is defined as a mask, and the mask can be represented as [F1, F2, F3, F4], where F1, F2, F3, and F4 represent the weights of the target speech information in the four voice regions (i.e., the voice regions where the driver's seat, passenger seat, driver's rear seat, and passenger's rear seat are located).
[0082] The preset sound zone can be set according to the actual situation (such as the placement position of the microphone in the actual target vehicle). This application embodiment does not limit this. That is, the preset sound zone can be any one of the sound zones of the driver's seat, passenger seat, driver's rear seat, and passenger's rear seat.
[0083] Specifically, when there is only one user, taking the user in the main driver's seat as an example, the preset sound zone can be set to the sound zone where the main driver's seat is located, and s12=s13=s14=0, then F1=1, F2=F3=F4=0, that is, the model input combination vector (feature) is: [Y1,Y2,Y3,Y4,1,0,0,0], and the model output vector can be approximately: [1,0,0,0]; when there are two users, taking the user in the main driver's seat and the user in the passenger seat as examples, and the preset sound zone is still set to the sound zone where the main driver's seat is located, then s12=s14=0, F2=F4=0, that is, the model input combination vector (feature) is: [Y1,Y2,Y3,Y4,1,0,1,0], and the model output vector can be: [F1,0,F3,0].
[0084] Then, the weights [F1,F2,F3,F4] of the target speech information in the four voice zones of the target vehicle can be multiplied with the corresponding transformed frequency domain target speech information [Y1,Y2,Y3,Y4] respectively, and the results can be subjected to inverse Fourier transform to obtain the enhanced target speech information, which are defined as y1', y2', y3', and y4' respectively.
[0085] In another optional implementation, the vehicle assistance information of the target vehicle may also include the vehicle speed information and window status information. In this case, the implementation process of step 202 may include: First, performing a Fourier transform on the target speech information to obtain the transformed target speech information. Then, using the vehicle speed information, window status information, and transformed target speech information of the target vehicle, a combined vector is constructed. This combined vector is then input into a pre-built speech enhancement model to predict the weights of the target speech information in each voice region of the target vehicle. Next, each weight can be multiplied by the target speech information of the corresponding voice region, and the calculation results are subjected to an inverse Fourier transform to obtain the enhanced target speech information, which is then used to execute the subsequent step S203.
[0086] In this implementation, there are no restrictions on the method of obtaining the target vehicle's speed information and window status information. Furthermore, after obtaining the target vehicle's speed information and window status information, they can be converted into representation vectors of corresponding dimensions. For example, the target vehicle's speed information can be converted into a 4-dimensional representation vector, with each dimension representing the target vehicle being stationary (0 km / h), at low speed (0–40 km / h), at medium speed (40–80 km / h), or at high speed (above 80 km / h). Similarly, the window status information can be converted into a 4-dimensional representation vector, with each dimension representing the window opening status of each audio frequency range (i.e., whether the window for the corresponding audio frequency range is open).
[0087] For example: still using Figure 1 Taking a target vehicle with four frequency bands as an example, when the vehicle speed information is obtained using the onboard speedometer and it is determined that the target vehicle is stationary, the vehicle speed information can be converted into a 4-dimensional representation vector: Information_speed = [1, 0, 0, 0]. Similarly, when the vehicle speed information is obtained using the onboard speedometer and it is determined that the target vehicle is at high speed, the vehicle speed information can be converted into a 4-dimensional representation vector: Information_speed = [0, 0, 0, 1].
[0088] When the window status information is obtained by detecting the window opening status and it is determined that only the driver's side window of the target vehicle is open, the window status information of the target vehicle can be converted into a corresponding 4-dimensional representation vector: Information_window = [1, 0, 0, 0]. Similarly, when the window status information is obtained by detecting the window opening status and it is determined that both the driver's side and passenger side windows of the target vehicle are open, the window status information of the target vehicle can be converted into a 4-dimensional representation vector: Information_window = [1, 1, 0, 0].
[0089] Based on this, in order to improve the voice interaction experience of in-vehicle users, after obtaining the target voice information of in-vehicle users in each voice zone of the target vehicle, the above formulas (1) and (2) can still be used to perform Fourier transform on it to obtain the transformed frequency domain target voice information, which will not be elaborated here. (Note: The last sentence about "excretion" is unrelated and appears to be a separate, incomplete thought.) Figure 1 The example shown is a target vehicle containing four frequency bands.
[0090] Furthermore, existing or future vector combination methods can be utilized to fuse the 4D representation vectors corresponding to the vehicle speed information and window status information of the target vehicle with the transformed frequency domain target speech information (i.e., Yi) to construct a combined vector, which is defined as a feature. Specifically, the vectors corresponding to the vehicle speed information, window status information, and transformed target speech information can be concatenated to obtain the combined vector as: feature = [Y1, Y2, Y3, Y4, Information_speed, Information_window]; or, gating can be applied to the input layer or intermediate layer of the speech enhancement model to guide the model to better separate speech by utilizing the amplitude and phase differences between microphones and seat-assisted information, thereby achieving vector combination of the target vehicle's seat information and the transformed target speech information.
[0091] The obtained combined vector (feature) is then input into the pre-built speech enhancement model to calculate the ratio between the frequency domain signal of the target speech information collected in each voice region and the frequency domain signal with noise collected in the preset voice region. This ratio is used as the weight of the target speech information in each voice region, thereby predicting the weight of the target speech information in the four voice regions of the target vehicle. This weight is defined as a mask, and the mask can be represented as [F1, F2, F3, F4], where F1, F2, F3, and F4 represent the weights of the target speech information in the four voice regions (i.e., the voice regions where the driver's seat, passenger seat, driver's rear seat, and passenger's rear seat are located).
[0092] It should be noted that the preset sound zone here can also be set according to the actual situation (such as the placement position of the microphone in the actual target vehicle). This application embodiment does not limit this, but usually the sound zone where the seat without the window is located is selected as the preset sound zone, that is, the sound zone with the best signal-to-noise ratio is selected as the preset sound zone to reduce the impact of noise on the speech enhancement result.
[0093] For example, when only the driver's side window is open, based on the actual microphone position distribution, it can be concluded that the noise energy of the microphone near the driver's side is relatively stronger than the noise energy of other microphones far away from the driver's side. In this case, the sound zone where the driver's side is located can be selected as the preset sound zone, so as to obtain the signal-to-noise ratio gain brought by the microphone position distribution.
[0094] Alternatively, an auxiliary microphone can be installed at the armrest between the left and right audio zones. When the signal-to-noise ratio of the microphone near the window is too low, the speech information collected by the microphone in that audio zone can be added to the combination vector input to the model.
[0095] Then, the weights [F1,F2,F3,F4] of the target speech information in the four voice zones of the target vehicle can be multiplied with the corresponding transformed frequency domain target speech information [Y1,Y2,Y3,Y4] respectively, and the results can be subjected to inverse Fourier transform to obtain the enhanced target speech information, which are defined as y1', y2', y3', and y4' respectively.
[0096] It should be noted that the above voice enhancement process is described with the target vehicle being a four-seater vehicle. However, this application does not limit the specific composition of the vehicle. For example, the target vehicle can also be a seven-seater vehicle or other models. For the voice enhancement process of other models, the above-mentioned in-vehicle voice enhancement process for four-seater vehicles can be referred to, and will not be described in detail here.
[0097] Next, this embodiment will introduce the construction process of the speech enhancement model mentioned in the above steps. In one optional implementation, the construction process of the speech enhancement model may specifically include: first, obtaining the sample vehicle auxiliary information of the sample vehicle, and obtaining the sample speech information of the vehicle users in each voice zone of the sample vehicle; then, using the sample vehicle auxiliary information, the sample speech information and the target loss function, training the initial speech enhancement model to obtain the speech enhancement model.
[0098] Specifically, in this implementation, a significant amount of preparatory work is required to construct the speech enhancement model. First, a large amount of in-vehicle auxiliary information from sample vehicles and speech information from in-vehicle users in various audio zones needs to be collected. These are used as sample in-vehicle auxiliary information (including sample seat information, sample vehicle speed information, and sample window status information) and sample speech information to form the model training data. For example, a large amount of seat information from sample vehicles under different conditions can be collected in advance, such as the different numbers of users and corresponding seat information in a four-seater sample vehicle with one or more users. Similarly, a large amount of vehicle speed information (including stationary, low-speed, medium-speed, and high-speed states) and window status information (i.e., the open state of windows for different seats) from sample vehicles under different conditions also needs to be collected in advance. Furthermore, the speech information emitted by in-vehicle users in various audio zones of the sample vehicles under these conditions, with the permission of the in-vehicle users, collectively constitutes the model training data. The weighted recognition results corresponding to the sample speech information under these conditions are then manually labeled. Next, the initial speech enhancement model can be trained based on these sample vehicle assistance information, sample speech information, the weight recognition results corresponding to the sample speech information, and the target loss function (the specific function is not limited and can be set according to the actual situation and experience value), thereby generating a speech enhancement model.
[0099] One possible implementation is that the initial speech enhancement model can be (but is not limited to) at least one of Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), real or complex networks.
[0100] Specifically, during model training, sample vehicle-mounted auxiliary information (such as sample seat information) and sample speech information under the target vehicle's operating conditions can be extracted sequentially from the training data. The sample speech information is then subjected to Fourier transform to convert it into frequency domain sample speech information. The sample vehicle-mounted auxiliary information and the converted frequency domain sample speech information are then converted into corresponding representation vectors. The two representation vectors are fused and used as the model input, with the corresponding weight recognition result as the output. Multiple rounds of model training are performed, and the weight recognition result obtained in each round of training is compared with the corresponding manually labeled result. The model parameters are updated based on the differences between the two until a preset condition is met, such as the target loss function being very small and basically unchanged. At this point, the update of the model parameters is stopped, the training of the speech enhancement model is completed, and a trained speech enhancement model is generated.
[0101] Based on this, after training and generating a speech enhancement model using sample vehicle assistance information and sample speech information, the generated speech enhancement model can be further validated using verification vehicle assistance information and verification speech information. The specific validation process may include the following steps (1)-(3):
[0102] Step (1): Obtain the verification vehicle assistance information of the verification vehicle, and obtain the verification voice information of the vehicle users in each audio zone of the verification vehicle.
[0103] In this embodiment, in order to verify the voice enhancement model, it is first necessary to obtain the verification vehicle auxiliary information (including verification seat information, verification speed information, and verification window status information) of the verification vehicle, as well as the verification voice information of the vehicle users in each voice zone of the verification vehicle. For example, the number of different users and their corresponding seat information can be obtained for a four-seat verification vehicle with one or more users. Similarly, a large amount of vehicle speed information (including stationary state, low speed state, medium speed state, and high speed state) and window status information (i.e., the opening state of windows of different seats) of the verification vehicle can be collected in advance, respectively as verification seat information, vehicle speed information, and verification window status information. With the permission of the vehicle users in the verification vehicle, 1,000 voice data from different vehicle users in each voice zone are collected as verification voice information. Among them, the verification vehicle auxiliary information and verification voice information refer to the vehicle auxiliary information and voice information that can be used to verify the voice enhancement model. After obtaining these verification vehicle auxiliary information and verification voice information and the weight recognition label corresponding to each verification voice information, the subsequent steps (2) can be executed.
[0104] Step (2): Input the verification combination vector constructed using the verification vehicle auxiliary information and the transformed frequency domain verification speech information into the speech enhancement model to obtain the weight prediction result corresponding to the verification speech information.
[0105] After obtaining the verification vehicle assistance information and verification voice information through step (1), the verification voice information can be further transformed into frequency domain verification voice information by Fourier transform. Then, the verification vehicle assistance information and the transformed frequency domain verification voice information can be converted into corresponding representation vectors respectively. The two representation vectors are then fused and input into the voice enhancement model to obtain the weight prediction result corresponding to the verification voice information, which is used to execute the subsequent step (3).
[0106] Step (3): When the weight prediction result of the verification speech information is inconsistent with the weight label result corresponding to the verification speech information, the verification speech information and the verification vehicle auxiliary information are respectively used as sample speech information and sample vehicle auxiliary information to update the speech enhancement model.
[0107] After obtaining the weight prediction result of the verification speech information through step (2), if the weight prediction result of the verification speech information is inconsistent with the real weight recognition result (such as the manually labeled weight label result) corresponding to the verification speech information, the verification speech information and the verification vehicle auxiliary information can be used as sample speech information and sample vehicle auxiliary information respectively to update the parameters of the speech enhancement model.
[0108] Through the above embodiments, the voice enhancement model can be effectively validated by verifying in-vehicle assistance information and verification voice information. When the weight prediction result of the verification voice information is inconsistent with the actual weight recognition result (such as the manually labeled weight label result) of the verification voice information, the voice enhancement model can be adjusted and updated in a timely manner, thereby helping to improve the prediction accuracy and precision of the voice enhancement model.
[0109] S203: Based on the enhanced target voice information, perform preset operation processing on the in-vehicle user and / or target vehicle to obtain the processing result.
[0110] In this embodiment, after obtaining the enhanced target voice information in step S202, the system can further activate preset devices in the target vehicle (such as the vehicle's air conditioning or audio player) based on the enhanced target voice information. Pre-defined operation processing is then performed to locate and identify the vehicle user who issued the wake-up voice (e.g., "It's too hot, please turn on the air conditioning"). For example, the wake-up result can be compared with the energy of the enhanced voice in each voice region, and the user in the voice region with the highest energy can be selected as the target wake-up person. Speech recognition is then performed to obtain the recognition result of the target speaker. This improves the wake-up, location, and recognition effects, thereby enhancing the user's voice interaction experience while driving.
[0111] In this way, by executing the above steps S201-S203, the voice enhancement of the vehicle's voice zone can be achieved by combining the vehicle seat information, vehicle speed information and window information, which can greatly improve the wake-up and positioning effect in subsequent processing, and further improve the user's voice interaction experience while driving.
[0112] In summary, the in-vehicle voice enhancement method provided in this embodiment first acquires the in-vehicle auxiliary information of the target vehicle and the target voice information of the in-vehicle users in each voice zone of the target vehicle. Then, it uses the in-vehicle auxiliary information to enhance the target voice information, obtaining enhanced target voice information. Next, based on the enhanced target voice information, it performs preset operation processing on the in-vehicle users and / or the target vehicle to obtain the processing result. It can be seen that this application first enhances the voice of the in-vehicle users in each voice zone of the vehicle based on the in-vehicle auxiliary information, and then uses the enhanced voice to perform subsequent preset operation processing such as vehicle wake-up, user location, and recognition, thereby improving the wake-up, location, and recognition effects, and thus enhancing the user's voice interaction experience while the target vehicle is in motion.
[0113] Second Embodiment
[0114] This embodiment will introduce an in-vehicle voice enhancement device; please refer to the above method embodiment for related content.
[0115] See Figure 3 This is a schematic diagram of the composition of an in-vehicle voice enhancement device provided in this embodiment. The device 300 includes:
[0116] The acquisition unit 301 is used to acquire the vehicle assistance information of the target vehicle and the target voice information of the vehicle users in each voice zone of the target vehicle.
[0117] Enhancement unit 302 is used to enhance the target speech information using the vehicle-mounted auxiliary information to obtain enhanced target speech information;
[0118] The processing unit 303 is used to perform preset operation processing on the vehicle user and / or target vehicle based on the enhanced target voice information to obtain the processing result.
[0119] In one implementation of this embodiment, the vehicle assistance information of the target vehicle includes the seat information of the target vehicle; the enhancement unit 302 includes:
[0120] The first transformation subunit is used to perform a Fourier transform on the target speech information to obtain the transformed target speech information;
[0121] The first prediction subunit is used to construct a combined vector using the seat information of the target vehicle and the converted target speech information, and input the combined vector into a pre-built speech enhancement model to predict the weights of the target speech information in each voice region of the target vehicle.
[0122] The first calculation subunit is used to multiply the weights with the corresponding target speech information respectively, and perform inverse Fourier transform on the calculation results to obtain the enhanced target speech information.
[0123] In one implementation of this embodiment, the first prediction subunit is specifically used for:
[0124] The vector corresponding to the seat information of the target vehicle and the vector corresponding to the converted target speech information are concatenated to obtain a combined vector; or, a combined vector is constructed using the seat information of the target vehicle and the converted target speech information through gating.
[0125] In one implementation of this embodiment, the vehicle assistance information of the target vehicle includes the vehicle speed information and window status information; the enhancement unit 302 includes:
[0126] The second transformation subunit is used to perform Fourier transform on the target speech information to obtain the transformed target speech information.
[0127] The second prediction subunit is used to construct a combined vector using the vehicle speed information and window status information of the target vehicle and the converted target speech information, and input the combined vector into a pre-built speech enhancement model to predict the weights of the target speech information in each voice region of the target vehicle.
[0128] The second calculation subunit is used to multiply the weights with the corresponding target speech information respectively, and perform inverse Fourier transform on the calculation results to obtain the enhanced target speech information.
[0129] In one implementation of this embodiment, the second prediction subunit is specifically used for:
[0130] The vector corresponding to the vehicle speed information of the target vehicle, the vector corresponding to the window status information, and the vector corresponding to the converted target voice information are concatenated to obtain a combined vector; or, a combined vector is constructed using the vehicle speed information, window status information, and the converted target voice information through gating.
[0131] In one implementation of this embodiment, the speech enhancement model includes at least one of convolutional neural networks (CNN), recurrent neural networks (RNN), and real or complex networks.
[0132] In one implementation of this embodiment, the first prediction subunit or the second prediction subunit is specifically used for:
[0133] The combined vector is input into a pre-built speech enhancement model, and the ratio between the frequency domain signal of the target speech information collected in each voice region and the frequency domain signal with noise frequency collected in the preset voice region is calculated as the weight of the target speech information in each voice region.
[0134] In one implementation of this embodiment, the processing unit 303 is specifically used for:
[0135] Based on the enhanced target voice information, the preset device of the target vehicle is activated, and the vehicle user who issued the activation voice is located and identified to obtain the processing result.
[0136] Furthermore, this application embodiment also provides an in-vehicle voice enhancement device, including: a processor, a memory, and a system bus;
[0137] The processor and the memory are connected via the system bus;
[0138] The memory is used to store one or more programs, the one or more programs including instructions, which, when executed by the processor, cause the processor to perform any of the above-described implementations of the in-vehicle voice enhancement method.
[0139] Furthermore, embodiments of this application also provide a computer-readable storage medium storing instructions that, when executed on a terminal device, cause the terminal device to perform any of the above-described implementations of the in-vehicle voice enhancement method.
[0140] Furthermore, this application embodiment also provides a computer program product, which, when run on a terminal device, causes the terminal device to execute any of the above-described implementation methods of the in-vehicle voice enhancement method.
[0141] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that all or part of the steps in the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network communication device such as a media gateway, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.
[0142] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.
[0143] It should also be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0144] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for enhancing in-vehicle voice, characterized in that, include: Acquire the vehicle assistance information of the target vehicle, and acquire the target voice information of the vehicle users in each audio zone of the target vehicle; The target speech information is subjected to Fourier transform to obtain the transformed target speech information; A combined vector is constructed using the vehicle-mounted auxiliary information of the target vehicle and the converted target speech information, and the combined vector is input into a pre-constructed speech enhancement model to predict the weights of the target speech information in each voice region of the target vehicle. The weights are multiplied by the corresponding target speech information, and the results are subjected to inverse Fourier transform to obtain the enhanced target speech information. Based on the enhanced target voice information, preset operation processing is performed on the in-vehicle user and / or target vehicle to obtain the processing result; When the vehicle assistance information of the target vehicle includes the seat information of the target vehicle, the step of constructing a combined vector using the vehicle assistance information of the target vehicle and the converted target speech information includes: A combined vector is constructed using the seat information of the target vehicle and the converted target speech information; Alternatively, when the target vehicle's onboard assistance information includes the target vehicle's speed information and window status information, the step of constructing a combined vector using the target vehicle's onboard assistance information and the converted target voice information includes: A combined vector is constructed using the target vehicle's speed information, window status information, and the converted target voice information.
2. The method according to claim 1, characterized in that, The step of constructing a combined vector using the seat information of the target vehicle and the converted target speech information includes: The vector corresponding to the seat information of the target vehicle and the vector corresponding to the converted target speech information are concatenated to obtain a combined vector; or, a combined vector is constructed using the seat information of the target vehicle and the converted target speech information through gating.
3. The method according to claim 1, characterized in that, The step of constructing a combined vector using the target vehicle's speed information, window status information, and the converted target speech information includes: The vector corresponding to the vehicle speed information of the target vehicle, the vector corresponding to the window status information, and the vector corresponding to the converted target voice information are concatenated to obtain a combined vector; or, a combined vector is constructed using the vehicle speed information, window status information, and the converted target voice information through gating.
4. The method according to any one of claims 1 to 3, characterized in that, The speech enhancement model includes at least one of convolutional neural networks (CNN), recurrent neural networks (RNN), and real or complex networks.
5. The method according to claims 1 to 3, characterized in that, The step of inputting the combined vector into a pre-built speech enhancement model to predict the weights of the target speech information for each voice region on the target vehicle includes: The combined vector is input into a pre-built speech enhancement model, and the ratio between the frequency domain signal of the target speech information collected in each voice region and the frequency domain signal with noise frequency collected in the preset voice region is calculated as the weight of the target speech information in each voice region.
6. The method according to claim 1, characterized in that, The step of performing preset operation processing on the in-vehicle user and / or target vehicle based on the enhanced target voice information to obtain the processing result includes: Based on the enhanced target voice information, the preset device of the target vehicle is activated, and the vehicle user who issued the activation voice is located and identified to obtain the processing result.
7. A vehicle-mounted voice enhancement device, characterized in that, include: The acquisition unit is used to acquire the vehicle assistance information of the target vehicle and the target voice information of the vehicle users in each audio zone of the target vehicle. The enhancement unit is used to enhance the target speech information using the vehicle-mounted auxiliary information to obtain enhanced target speech information; The processing unit is used to perform preset operation processing on the in-vehicle user and / or target vehicle based on the enhanced target voice information to obtain the processing result; The vehicle assistance information of the target vehicle includes the seat information of the target vehicle; The enhancement unit includes: The first transformation subunit is used to perform a Fourier transform on the target speech information to obtain the transformed target speech information; The first prediction subunit is used to construct a combined vector using the seat information of the target vehicle and the converted target speech information, and input the combined vector into a pre-built speech enhancement model to predict the weights of the target speech information in each voice region of the target vehicle. The first calculation subunit is used to multiply the weights with the corresponding target speech information respectively, and perform inverse Fourier transform on the calculation results to obtain the enhanced target speech information. Alternatively, the vehicle assistance information of the target vehicle includes the vehicle speed information and window status information; the enhancement unit includes: The second transformation subunit is used to perform Fourier transform on the target speech information to obtain the transformed target speech information. The second prediction subunit is used to construct a combined vector using the vehicle speed information and window status information of the target vehicle and the converted target speech information, and input the combined vector into a pre-built speech enhancement model to predict the weights of the target speech information in each voice region of the target vehicle. The second calculation subunit is used to multiply the weights with the corresponding target speech information respectively, and perform inverse Fourier transform on the calculation results to obtain the enhanced target speech information.
8. A vehicle-mounted voice enhancement device, characterized in that, include: Processor, memory, system bus; The processor and the memory are connected via the system bus; The memory is used to store one or more programs, the one or more programs including instructions that, when executed by the processor, cause the processor to perform the method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed on a terminal device, cause the terminal device to perform the method described in any one of claims 1-6.
Citation Information
Patent Citations
Vehicle-mounted multi-tone-zone voice processing method and related device
CN111599366A
Voice processing method and device, storage medium and electronic equipment
CN113270095A