Voice control method, device, system and equipment, storage medium and vehicle

By processing the initial speech signal to generate a target masking signal, the problem of poor auditory experience caused by noise signal masking in the prior art is solved, and the speech signal masking effect and auditory comfort are improved.

CN120932647APending Publication Date: 2025-11-11BYD CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511264880.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing technologies mask speech by mixing it with output noise signals, resulting in a poor auditory experience for drivers and passengers.

Method used

By processing the initial speech signal, a first masking signal and a preset second masking signal are generated. Combined with deep learning prediction weights, a target masking signal is output to mask the speech signal and avoid an unpleasant auditory experience.

Benefits of technology

It effectively masks speech signals, reduces unpleasant auditory experiences, and improves the auditory comfort of drivers and passengers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932647A_ABST
    Figure CN120932647A_ABST
Patent Text Reader

Abstract

The invention relates to a voice control method, device, system and equipment, a storage medium and a vehicle, and the method comprises the steps: outputting a target masking signal based on a first masking signal obtained through the processing of an initial voice signal and a preset second masking signal, and enabling the target masking signal to be used for masking the initial voice signal. According to the invention, while the voice signal masking effect can be effectively ensured, the first masking signal component obtained based on the initial voice signal in the target masking signal can also effectively relieve bad hearing experience brought to a user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of voice control, and more particularly to a voice control method, apparatus, system, device, storage medium, and vehicle. Background Technology

[0002] With the increasing popularity of intelligent electric vehicles, users' demands for in-car experience have shifted from basic comfort to a deeper focus on privacy protection, acoustic environment control, and contextualized interaction. Compared to physical isolation, active acoustic control methods address the space and cost limitations of traditional solutions.

[0003] However, the related technologies can easily lead to a poor auditory experience for drivers and passengers by mixing in noise signals. Summary of the Invention

[0004] This application provides a voice control method, apparatus, system, device, storage medium, and vehicle, aiming to solve the problem of poor user experience in related technologies.

[0005] Firstly, this application provides a voice control method, including:

[0006] Based on the first masking signal obtained from the processing of the initial speech signal and the preset second masking signal, a target masking signal is output, which is used to mask the initial speech signal.

[0007] As a further feasible implementation of this application, the first masking signal is obtained through the following steps:

[0008] The first masking signal is determined based on the initial speech signal and the reversed speech signal of the initial speech signal.

[0009] As a further feasible implementation of this application, determining the first masking signal based on the initial speech signal and the reversed speech signal of the initial speech signal includes:

[0010] The initial speech signal and the reverse speech signal of the initial speech signal are weighted based on a preset first weight to obtain the first masking signal.

[0011] As a further feasible implementation of this application, the voice parameters include at least one of masking ratio, clarity index, and voice isolation.

[0012] As a further feasible implementation of this application, the method further includes:

[0013] The masking signal ratio and the clarity index are determined based on the first masking signal and the initial speech signal;

[0014] The voice isolation is determined based on the front and rear concealment signals in the first masking signal.

[0015] As a further feasible implementation of this application, the method further includes:

[0016] The initial voice signal of the target area of ​​the vehicle is determined based on the energy information and / or time delay information of the multi-channel voice signal inside the vehicle.

[0017] As a further feasible implementation of this application, the method further includes:

[0018] The multi-channel speech signal is preprocessed to determine the initial speech signal of the target area of ​​the vehicle based on the energy information and / or time delay information of the preprocessed multi-channel speech signal; the preprocessing includes at least one of time-frequency domain conversion and filtering.

[0019] As a further feasible implementation of this application, the method further includes:

[0020] If the initial voice signal meets the preset conditions, the step of outputting the target masking signal based on the first masking signal obtained by processing the initial voice signal and the preset second masking signal is executed.

[0021] As a further feasible implementation of this application, the preset conditions include:

[0022] The short-time energy of the initial speech signal exceeds a preset energy threshold, and / or the short-time energy of the weighted speech signal determined based on the initial speech signal exceeds a preset energy threshold.

[0023] As a further feasible implementation of this application, the second masking signal includes a natural sound masking signal and / or a human voice masking signal.

[0024] As a further feasible implementation of this application, the method further includes:

[0025] The target voice masking signal associated with the baseband information is determined based on the baseband information of the initial speech signal.

[0026] As a further feasible implementation of this application, the method further includes:

[0027] In response to the detection of the initial speech signal, the voice masking signal is output.

[0028] As a further feasible implementation of this application, the step of outputting a target masking signal based on a first masking signal obtained from processing the initial speech signal and a preset second masking signal, wherein the target masking signal is used to mask the initial speech signal, includes:

[0029] The first masking signal and the second masking signal are weighted based on a preset second weight to obtain the target masking signal.

[0030] As a further feasible implementation of this application, the method further includes:

[0031] The second weight is predicted based on a preset objective function and through deep learning.

[0032] As a further feasible implementation of this application, the method further includes:

[0033] The target masking signal is output based on the multi-channel voice output device.

[0034] As a further feasible implementation of this application, the method further includes:

[0035] The target masking signal is output based on the prediction of the weight and / or phase of the speech signal output by each channel in the multi-channel speech output device using deep learning.

[0036] Secondly, this application provides a voice control device, including an output module, the output module being used to execute the voice control method as described in any of the preceding claims.

[0037] Thirdly, this application provides a voice control system, comprising:

[0038] Audio acquisition device, used to acquire voice signals inside the vehicle;

[0039] A processor is configured to execute the voice control method as described in any of the preceding claims to process the voice signal to obtain a target masking signal;

[0040] An audio output device for outputting the target masking signal.

[0041] Fourthly, this application provides an electronic device, including a processor; a memory for storing processor-executable instructions; wherein the processor is configured to perform the steps of the voice control method described in any of the preceding claims.

[0042] Fifthly, this application provides a computer-readable storage medium having a computer program stored thereon, the computer program being loaded by a processor to perform the steps of the voice control method described in any of the preceding claims.

[0043] Sixthly, this application provides a vehicle that performs voice control by executing the voice control method described in any of the above claims, or includes the voice control device described in the above claims, or includes the voice control system described in the above claims.

[0044] This application determines and outputs a target masking signal by combining a masking signal obtained from speech signal processing and a preset second masking signal, which is used to mask the speech signal. While effectively ensuring the masking effect of the speech signal, the first masking signal component in the target masking signal, which is based on the initial speech signal, can also effectively alleviate the unpleasant auditory experience for the user. Attached Figure Description

[0045] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0046] To gain a more complete understanding of this application and its beneficial effects, the following description will be provided in conjunction with the accompanying drawings, wherein the same reference numerals in the following description denote the same parts.

[0047] Figure 1 This is a flowchart illustrating the steps of the voice control method provided in the embodiments of this application;

[0048] Figure 2 This application provides a schematic flowchart of a step for screening initial speech signals in an embodiment of the present application.

[0049] Figure 3 This application provides a schematic flowchart of a voice activation detection process.

[0050] Figure 4 This is a flowchart illustrating the steps of an adaptive masking method provided in an embodiment of this application.

[0051] Figure 5 This is a complete flowchart illustrating a voice control method provided in an embodiment of this application;

[0052] Figure 6 A schematic diagram of the functional modules of a voice control system provided in this application;

[0053] Figure 7 A schematic diagram of the specific architecture of a voice control system provided in this application;

[0054] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;

[0055] Figure 9 This is a structural schematic diagram of a vehicle provided in an embodiment of this application. Detailed Implementation

[0056] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the protection scope of this application.

[0057] To clearly understand the voice control method, apparatus, system, device, storage medium, and vehicle provided in the embodiments of this application, the specific application scenarios of the voice control method will be described below. Specifically, with the development of electric vehicles and hybrid vehicles, users' needs for the in-vehicle experience have shifted from basic comfort to deeper requirements for privacy protection, acoustic environment control, and contextualized interaction. Compared to the traditional physical partition method of setting up sound field barriers in the seats of drivers and passengers, acoustic control methods can effectively solve the space and cost pain points of vehicles, while avoiding the adverse experience brought to drivers and passengers by physical partitions.

[0058] However, acoustic-based control methods in related technologies typically mask speech signals by outputting preset noise signals when drivers and passengers require privacy protection, such as during phone calls. However, such solutions are quite annoying and can easily lead to a poor auditory experience for drivers and passengers. To address these technical problems, this application provides a novel speech control method that outputs a target masking signal in an appropriate manner, ensuring effective masking of the speech signal while avoiding a poor auditory experience for drivers and passengers. The following will describe this method in conjunction with specific embodiments.

[0059] For details, please refer to Figure 1 , Figure 1 This is a flowchart illustrating the steps of the voice control method provided in the embodiments of this application. Specifically, it includes step S110:

[0060] S110, based on the first masking signal obtained from the processing of the initial speech signal and the preset second masking signal, output a target masking signal, which is used to mask the initial speech signal.

[0061] In this embodiment, the initial voice signal refers to the voice signal collected inside the vehicle, particularly the voice signal requiring privacy protection. For example, it can typically be the voice signal of users within a specific area, including conversations and calls from rear passengers, as well as navigation prompts and vehicle control voice commands from the driver and passengers. This embodiment does not limit the types of voice signals requiring privacy protection. Specifically, the user can actively or the vehicle can passively detect and activate the masking function for the initial voice signal. This embodiment does not impose any limitations on this.

[0062] Unlike conventional methods that mask the initial speech signal by outputting a noise signal, this embodiment determines and outputs a target masking signal by combining a first masking signal obtained from processing the initial speech signal with a preset second masking signal. This effectively masks the initial speech signal while avoiding any unpleasant auditory experience for the driver or passengers. The specific details will be explained below.

[0063] Specifically, in some embodiments of this application, there are various possible implementation schemes for processing the initial speech signal to obtain the first masking signal, such as filtering the speech signal, time-frequency domain conversion, etc. However, as one possible implementation scheme of this application, in order to reduce the unpleasant auditory experience for drivers and passengers based on the first masking signal, in some embodiments of this application, the first masking signal can be obtained based on processing the time-flipped signal of the initial speech signal. That is, in one embodiment, the first masking signal is obtained through the following steps:

[0064] The first masking signal is determined based on the initial speech signal and the reversed speech signal of the initial speech signal.

[0065] Specifically, in this embodiment, the reversed-order speech signal of the initial speech signal refers to the signal obtained by reversing the signal frames in the initial speech signal. For ease of understanding, taking the initial speech signal as an example, where the initial speech signal is segmented into frames to obtain the nth and (n+1)th signal frames, the reversed-order speech signal of the initial speech signal is composed of the (n+1)th and nth signal frames. Of course, the above implementation is only one possible approach. In fact, the initial speech signal can be segmented into more signal frames, and the corresponding reversed-order speech signal is composed of these signal frames in reverse order. This embodiment does not limit this approach.

[0066] Based on the aforementioned scheme, the first masking signal is ultimately determined by fusing the initial speech signal and the reversed speech signal. For example, taking the two signal frames provided above as an example, the nth frame of the first masking signal is determined by the nth frame of the initial speech signal and the nth frame of the reversed speech signal (corresponding to the n+1th frame of the initial speech signal), and the n+1th frame of the first masking signal is determined by the n+1th frame of the initial speech signal and the n+1th frame of the reversed speech signal (corresponding to the nth frame of the initial speech signal).

[0067] Based on the aforementioned solutions, to avoid the need to record and analyze the recorded speech signals to reconstruct the initial speech signal, in some embodiments of this application, the initial speech signal and its reversed speech signal can be weighted by a certain weighting coefficient to obtain the first masking signal. That is, determining the first masking signal based on the initial speech signal and its reversed speech signal includes:

[0068] The initial speech signal and the reverse speech signal of the initial speech signal are weighted based on a preset first weight to obtain the first masking signal.

[0069] Specifically, in one embodiment, the first masking signal can be obtained by processing the reversed speech signal with a weighting coefficient β and fusing it with the initial speech signal processed by the weighting coefficient γ. Of course, the aforementioned solution is merely one possible embodiment. In fact, in some embodiments of this application, the first weight can also be dynamically and adaptively determined, that is, the first weight can be a dynamic weight. For example, an appropriate weight can be set based on features such as vehicle driving data or speech signals.

[0070] Of course, in order to ensure that the final target masking signal can guarantee the masking effect of the speech signal as much as possible, while ensuring the user experience, in some embodiments of this application, the first weight can also be obtained by constructing a specific objective function to indicate the signal masking effect and predicting it through deep learning. That is, in one embodiment, the method further includes:

[0071] The first weight is predicted based on a preset objective function and through deep learning.

[0072] In the embodiments of this application, the preset objective function can typically be constructed based on some speech feature parameters used to indicate the signal, determined by the first masking signal. For example, in some embodiments, the speech feature parameters may include a masking signal ratio, a clarity index, and a speech isolation degree to represent the masking effect, isolation effect, or the user's intelligibility of the output speech signal. For example, in some embodiments, the masking signal ratio and the clarity index may be determined based on the first masking signal and the initial speech signal, while the speech isolation degree may be determined specifically by the front and rear masking signals in the first masking signal. That is, in one embodiment, the method further includes:

[0073] The masking signal ratio and the clarity index are determined based on the first masking signal and the initial speech signal;

[0074] The voice isolation is determined based on the front and rear concealment signals in the first masking signal.

[0075] Building upon the aforementioned foundation, deep learning can predict the first weights by training a speech model using deep learning. This means that after training the speech model with sample speech signals, inputting the initial speech signal into the model yields the predicted first weights for the initial speech signal and its reversed order. These weights are then used to process the initial and reversed speech signals to obtain the first masking signal, optimizing the objective function of the target masking signal determined by this first masking signal. Furthermore, in some embodiments of this application, boundary conditions can be constructed for the speech feature parameters in the objective function to further ensure the masking effect on the initial speech signal. For example, in one embodiment, configurable boundary conditions include a clarity index <= 0.1, a masking signal ratio >= -8dB, and a speech isolation >= 9dB, etc. This application does not limit the boundary conditions in its embodiments.

[0076] Of course, the above is just an example of using deep learning to predict the weights of the initial speech signal and the reversed speech signal. In fact, in some embodiments, other parameters can also be predicted based on deep learning, such as the weights related to the first masking signal and the second masking signal, or the weights of each channel in the final output target masking signal. The specific implementation scheme will be described in detail in the following embodiments.

[0077] Furthermore, in order to output a target masking signal at an appropriate time to achieve the effect of masking a specific initial speech signal, some embodiments of this application also provide an implementation scheme for determining the initial speech signal requiring privacy protection from speech signals collected inside the vehicle, particularly from speech signals acquired through multi-channel acquisition. Specifically, this can be determined by analyzing the energy information or time delay information of the speech signal. In other words, the method further includes:

[0078] The initial voice signal of the target area of ​​the vehicle is determined based on the energy information and / or time delay information of the multi-channel voice signal inside the vehicle.

[0079] In some embodiments of this application, the multi-channel voice signals within the vehicle can be acquired by multiple voice acquisition devices, such as microphones, installed within the vehicle. Specifically, these voice acquisition devices can be positioned at different locations within the vehicle to acquire voice signals from different locations, thereby selecting the voice of a specific location, such as the rear passenger's voice, for subsequent processing. For example, by analyzing the position of each microphone and combining it with specific time delay information, and comparing the energy information within each channel, useless echoes, noise, and front-seat voice information can be filtered out, thereby extracting the rear-seat voice signal as the initial voice signal required for masking processing in this application.

[0080] Of course, it should be noted that, in order to ensure the extraction effect of the initial speech signal, in some embodiments of this application, the initial speech signal in the target area can also be determined by preprocessing the multi-channel speech signal, thereby analyzing the processed speech signal. The preprocessing typically includes at least one of time-frequency domain transformation and filtering. In other words, the method further includes:

[0081] The multi-channel speech signal is preprocessed to determine the initial speech signal of the target area of ​​the vehicle based on the energy information and / or time delay information of the preprocessed multi-channel speech signal; the preprocessing includes at least one of time-frequency domain conversion and filtering.

[0082] Specifically, to facilitate understanding of the above content, the following description will be provided in conjunction with specific embodiments. For examples from the embodiments of this application, please refer to... Figure 2 , Figure 2 This application provides a flowchart illustrating the steps for screening initial speech signals, specifically including steps S201 to S205:

[0083] S201, framing and windowing.

[0084] Specifically, in this embodiment, the acquired multi-channel speech signal is converted into a frame signal by frame-by-frame windowing, and windowing and Fourier transform are performed on each frame signal to obtain the transformed speech signal. Windowing is used to prevent spectral leakage, while Fourier transforming to the frequency domain facilitates subsequent calculations and reduces latency.

[0085] Specifically, in some embodiments, the frame length during the aforementioned framing process can be selected as 20ms or 10ms, and the frame length can be adjusted as needed. For example, during real-time speech processing, the system can achieve low latency while maintaining a good masking effect by selecting a shorter frame length. It should be noted that in some embodiments, the masking signal output by the above method will still have a certain delay. To reduce the negative user experience caused by this delay, in other embodiments, this delay can be filled by other masking signals, which will be explained in detail in subsequent embodiments.

[0086] Step 202: Echo cancellation.

[0087] Specifically, in this embodiment, adaptive filtering can be used to eliminate echoes in the acquired speech signal. Specifically, Kalman filtering can be used, taking data from the microphones at four locations and reference data from the speaker as input to the Kalman filter for processing, thus outputting echo-free data from four channels.

[0088] Step 203: Noise reduction.

[0089] Specifically, in the embodiments of this application, Wiener filtering can be used to further eliminate noise in the signal, specifically, it can effectively eliminate stationary and non-stationary noise in the speech signal.

[0090] Step 204: Select the microphone.

[0091] Specifically, in the embodiments of this application, since the microphones are set in different positions, the sound intensity and time of the voice signal from the back row transmitted to each microphone are different. Therefore, by comparing the energy and time delay of the four channels, the voice information of the speaker in the back row, that is, the voice signal in a specific area, can be extracted.

[0092] Step 205: Reconstruct the rear-seat audio.

[0093] Specifically, in this embodiment, the speech of the speaker in the back row can be reconstructed by taking the selected frequency domain data back to the time domain through inverse Fourier transform and then adding a synthesis window. This speech signal is the initial language signal for subsequent masking processing.

[0094] The solution provided in this application collects voice signals from multiple locations in a vehicle through multiple channels, and combines echo cancellation and noise reduction to select an initial voice signal from the collected voice signals, ensuring that useful voice information in specific areas, such as the rear seats, is utilized to the maximum extent, effectively improving the final masking effect.

[0095] Furthermore, in some embodiments of this application, in order to complete the masking of the initial speech signal when speech masking is required, a speech activation detection implementation scheme is also provided in some embodiments of this application. That is, by analyzing the initial speech signal to determine whether the speech signal is a speaking speech signal, the subsequent masking function is automatically activated. In other words, in one embodiment, the method further includes:

[0096] If the initial voice signal meets the preset conditions, the step of outputting the target masking signal based on the first masking signal obtained by processing the initial voice signal and the preset second masking signal is executed.

[0097] Specifically, in one embodiment, the initial speech signal meeting a preset condition typically means that there is speaking voice in the back row. Specifically, this can be determined by the short-time energy of the initial speech signal, thereby reducing latency. Of course, to improve detection effectiveness and avoid a decrease in detection accuracy due to potential sudden changes in short-term conditions, it can also be determined by weighting the short-time energy and the reference energy of historical moments. That is, meeting the preset condition includes:

[0098] The short-time energy of the initial speech signal exceeds a preset energy threshold, and / or the short-time energy of the weighted speech signal determined based on the initial speech signal exceeds a preset energy threshold.

[0099] For a clearer understanding of the above, please refer to [link / reference]. Figure 3 , Figure 3 This application provides a flowchart illustrating the steps of voice activation detection, specifically including steps S301 to S303:

[0100] S301, calculate the short-time energy in 10ms.

[0101] In this embodiment, the short-time energy of the selected speech signal frame is determined by a 10ms framing method. Specifically, this embodiment does not limit the implementation scheme for calculating the speech signal energy.

[0102] S302 smooths short-term energy.

[0103] In this embodiment of the application, in order to avoid the decrease in detection accuracy caused by sudden changes in short-term conditions, a certain smoothing process can be performed on the short-term energy. Specifically, a new voice signal energy value can be obtained by weighting the energy of the voice signal at a historical moment and the energy of the voice signal at the current moment, so as to be used for voice activation detection.

[0104] Specifically, in this embodiment, the weighting coefficients of the energy of the speech signal at a historical moment and the energy of the speech signal at the current moment are added together to equal one. The weighting coefficients can be adjusted by the actual vehicle, or they can be adjusted by the aforementioned deep learning-based adaptive method. This embodiment does not limit this.

[0105] Does the S303 have voice control?

[0106] In this embodiment, the energy of the obtained voice signal is compared with a preset energy threshold to determine whether there is voice in the rear seats, thereby selecting to activate or deactivate the masking function. Specifically, in some embodiments, if the energy is less than the energy threshold, no masking sound is output; if the energy is greater than or equal to the threshold, subsequent masking processing is performed. Specifically, the energy threshold selection can be adjusted according to the actual vehicle conditions; of course, it can also be adaptively adjusted based on each frame of signal, and this embodiment does not impose any limitations on this.

[0107] Building upon the aforementioned methods of determining the initial speech signal and determining the first masking signal based on the initial speech signal to mask the initial speech signal, some embodiments of this application further provide an implementation scheme for determining the final target masking signal used to mask the initial speech signal based on the first masking signal and the second masking signal. This will be described in detail below.

[0108] In some embodiments of this application, to improve the masking effect of speech signals, the second masking signal may include a natural sound masking signal and / or a human voice masking signal. The natural sound masking signal refers to improving the masking effect by simulating natural speech signals, while also reducing the unpleasant auditory experience for the user. The human voice masking signal, while avoiding an unpleasant auditory experience for the user, can reduce latency to a certain extent. That is, in the process of processing the initial speech signal to finally determine the target masking signal, outputting the human voice masking signal appropriately ensures the masking effect of the initial speech signal. This will be explained in detail below.

[0109] Specifically, in some embodiments, the natural sound masking signal may include signals such as stream sounds, waterfall sounds, wind sounds, rain sounds, etc., which may be pre-collected and stored in a preset speech signal library. This application embodiment does not limit this.

[0110] Specifically, in some embodiments, the voice masking signal includes male voices, female voices, and other speech signals. Particularly, in some embodiments, to avoid the masking effect being affected by differences between the voice masking signal and the speech signals of the speakers in the back row, some embodiments of this application may also select a target voice masking signal that is as similar as possible to the composition of the speakers in the back row based on the fundamental frequency information of the initial speech signal, such as a bass, tenor, or contralto. That is, in one embodiment, the method further includes:

[0111] The target voice masking signal associated with the baseband information is determined based on the baseband information of the initial speech signal.

[0112] In some embodiments of this application, the voice masking signal can be the same as the natural voice masking signal, which is pre-collected and stored in a preset speech signal library, and the corresponding voice masking signal is selected from the speech signal library based on the fundamental frequency information of the initial speech signal. This application does not impose any restrictions on this.

[0113] In some embodiments of this application, the voice masking signal can also be used to fill the delay of the output target masking signal to improve the masking effect. Specifically, in response to the detection of the initial voice signal or the detection of back-row speech, the voice masking signal can be output directly first, and the target masking signal can be output after determining the target masking signal based on the first masking signal and the second masking signal, thus avoiding the poor user experience caused by the delay.

[0114] Furthermore, in some embodiments of this application, to further improve the masking effect of the target masking signal on the initial speech signal, the target masking signal can also be obtained by weighting the first masking signal and the second masking signal. That is, the target masking signal is output based on the first masking signal obtained from processing the initial speech signal and the preset second masking signal. The target masking signal is used to mask the initial speech signal, including:

[0115] The first masking signal and the second masking signal are weighted based on a preset second weight to obtain the target masking signal.

[0116] Specifically, when the second masking signal includes both a natural sound masking signal and a human voice masking signal, the second weight then includes the weights of the first masking signal, the natural sound masking signal, and the human voice masking signal. Specifically, in some embodiments, similar to the aforementioned determination of the weights of the initial speech signal and the reversed speech signal, the second weight here can also be obtained through the objective function and based on deep learning prediction; that is, the method further includes:

[0117] The second weight is predicted based on a preset objective function and through deep learning.

[0118] Furthermore, in some embodiments of this application, the determined target masking signal can also be output through a multi-channel voice output device; that is, the method further includes:

[0119] The target masking signal is output based on the multi-channel voice output device.

[0120] The weights and / or phases output by each channel of the multi-channel speech output device can also be predicted based on deep learning. In other words, in one embodiment, the method further includes:

[0121] The target masking signal is output based on the prediction of the weight and / or phase of the speech signal output by each channel in the multi-channel speech output device using deep learning.

[0122] Specifically, for a better understanding of the process of determining and outputting the masking signal, please refer to [link to relevant documentation]. Figure 4 , Figure 4 This application provides a flowchart illustrating the steps of an adaptive masking method, specifically including steps S401 to S406:

[0123] S401, pre-stored voice frames.

[0124] In this embodiment of the application, the pre-stored voice frame is the aforementioned initial voice signal, especially the voice signal that has passed the voice activation detection function, which usually includes two voice frames.

[0125] S402, reverse n frames and n+1 frames, forward n frames and n+1 frames.

[0126] In this embodiment of the application, the two pre-stored frames of signals are processed in reverse order. The nth frame of the first masking signal is multiplied by the weight β of the reverse nth frame and multiplied by the weight λ of the forward n+1th frame. The n+1th frame of the first masking signal is multiplied by the weight β of the reverse n+1th frame and multiplied by the weight of the forward nth frame.

[0127] S403, Masking Sound Acquisition and Processing.

[0128] In this embodiment, dedicated microphones are installed at four headrest locations to pick up masking sound signals from these four locations. Since the picked-up speech contains the speaker's voice, the speaker's voice can be eliminated by using the aforementioned initial speech signal as a reference, leaving only the masking sound from the front and back rows.

[0129] S404, establish the objective function (including SII, TMR, ISO).

[0130] In this embodiment, the Target Mask Ratio (TMR), Speech Intelligence Index (SII), and Speech Isolation (ISO) are determined sequentially, and an objective function incorporating these three parameters is established. Weights δ1, δ2, and δ3 are assigned to SII, TMR, and ISO, respectively, and the objective function is used to solve for the optimal parameters. SII and TMR can be calculated from the masking sound and the initial speech signal, particularly based on the front row masking signal and the front row target signal. The isolation (ISO) can be calculated from the masking sound speech signals from both the front and rear rows.

[0131] S405, adaptive parameter adjustment.

[0132] In this embodiment, under the following boundary conditions—SII <= 0.1, TMR >= -8dB, ISO >= 9dB—the parameters are adjusted using deep learning to achieve the optimal value of the objective function. The parameters adjusted in real time include the forward frame weight λ and the reverse frame weight β (used to generate the adaptive masking signal, i.e., the first masking signal), the weights of the speech within the library (i.e., the human voice masking signal), the weights of the natural voice (i.e., the natural voice masking signal), and the weights A1-A12 and phases ω1-ω12 of each channel in the multi-channel output target masking signal. The initial values ​​are arbitrary.

[0133] S406, adaptive masking sound + in-library speech + natural sound.

[0134] The final target masking signal is output by combining the first masking signal, the human voice masking signal, and the natural sound masking signal.

[0135] This application determines and outputs a target masking signal by combining a masking signal obtained from speech signal processing and a preset second masking signal, which is used to mask the speech signal. While effectively ensuring the masking effect of the speech signal, the first masking signal component in the target masking signal, which is based on the initial speech signal, can also effectively alleviate the unpleasant auditory experience for the user.

[0136] In addition, for a better understanding of the complete implementation process of the voice control method provided in this application, please refer to [link / reference needed]. Figure 5 , Figure 5 A complete flowchart of a voice control method provided in an embodiment of this application is described in detail below.

[0137] In this embodiment of the application, after sensing the target speech signal (i.e. the initial speech signal) that needs to be protected, on the one hand, the corresponding speech sound (i.e. the human voice masking signal) in the corresponding speech library is extracted from the corresponding speech library through fundamental frequency detection. On the other hand, an adaptive masking sound (i.e. the first masking signal) is obtained based on the adaptive masking processing of the target speech signal, such as adaptive weighting and time flipping. Then, the adaptive masking sound, the speech sound in the library and the natural sound (i.e. the natural sound masking signal) are synthesized to output the final target speech protection sound (target masking signal).

[0138] Based on the aforementioned voice control method, this application also provides a voice control device, which includes an output module for executing the voice control method provided in any of the foregoing embodiments.

[0139] In addition, this application also provides a voice control system, please refer to [link / reference]. Figure 6 , Figure 6 The functional module diagram of a voice control system provided in this application specifically includes:

[0140] Audio acquisition device 610 is used to acquire voice signals inside the vehicle;

[0141] The controller 620 is used to process the voice signal to obtain a target masking signal;

[0142] Audio output device 630 is used to output the target masking signal.

[0143] The specific steps executed by the controller 620 can be referred to the voice control method provided in any of the foregoing embodiments, and this application does not impose any limitations on them.

[0144] In addition, for a clear understanding of the voice control system provided in this application, please refer to [link / reference needed]. Figure 7 , Figure 7 The specific architecture diagram of a voice control system provided in this application includes:

[0145] Voice acquisition module: mainly to capture the voice of people talking or chatting in the car in real time. Microphones ①-④ pick up voice signals from four positions in the car. At the same time, four additional microphones ⑤-⑧ are provided in the front and rear rows to pick up masking sound signals.

[0146] Multi-zone microphone processing module: Performs echo cancellation and noise reduction on the speech signals from the four microphones to extract clean speaker speech. Then, based on the energy and time delay of each microphone, it performs microphone selection processing to select the speech of the speaker in the back row as the input for the next step.

[0147] VAD detection module: Calculates short-time energy based on the input of the previous speech and smooths the short-time energy. If there is speech, it performs adaptive masking. If there is no speech, this function is not enabled.

[0148] Adaptive masking processing module: Based on the objective function constructed by SII, TMR, and ISO, it uses adaptive weighting and improved time reversal methods to generate adaptive masking sound from the previous speech, and then superimposes it with natural sound and speech in the library to achieve the effect of real-time adjustment without being deciphered by the recording.

[0149] Masking signal replay module: Through the cooperation of in-cabin hardware and software, masking sounds are output from ten speakers in the vehicle audio system to complete the real-time masking operation of the target voice.

[0150] This application determines and outputs a target masking signal by combining a masking signal obtained from speech signal processing and a preset second masking signal, which is used to mask the speech signal. While effectively ensuring the masking effect of the speech signal, the first masking signal component in the target masking signal, which is based on the initial speech signal, can also effectively alleviate the unpleasant auditory experience for the user.

[0151] Figure 8 This is a block diagram illustrating an electronic device 800 according to an exemplary embodiment. For example... Figure 8 As shown, the electronic device 800 may include a processor 801 and a memory 802. The electronic device 800 may also include one or more of a multimedia component 803, an input / output (I / O) component 804, and a communication component 805. In this embodiment, the electronic device 800 may be a device that integrates and implements the voice control method provided in this embodiment.

[0152] The processor 801 controls the overall operation of the electronic device 800 to complete all or part of the steps in the aforementioned voice control method. The memory 802 stores various types of data to support the operation of the electronic device 800. This data may include, for example, instructions for any application or method operating on the electronic device 800, and application-related data such as contact data, sent and received messages, pictures, audio, video, etc. The memory 802 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read Only Memory (PROM), Read Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The multimedia component 803 may include a screen and audio components. The screen may be, for example, a touchscreen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in memory 802 or transmitted via communication component 805. The audio component also includes at least one speaker for outputting audio signals. I / O component 804 provides an interface between processor 801 and other interface modules, such as a keyboard, mouse, buttons, etc. These buttons may be virtual or physical buttons. Communication component 805 is used for wired or wireless communication between the electronic device 800 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, 4G, NB-IoT, eMTC, or other 5G technologies, or combinations thereof, is not limited here. Therefore, the corresponding communication component 805 may include: a WiFi module, a Bluetooth module, an NFC module, etc.

[0153] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described voice control method.

[0154] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the voice control method provided in any of the above embodiments.

[0155] This application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it enables the computer program product to implement the voice control method provided in any of the above embodiments.

[0156] This application also provides a vehicle, such as Figure 9 The diagram shown is a structural schematic of a vehicle provided in this application, wherein the vehicle is voice-controlled by executing the voice control method described in any of the above claims, or includes the voice control device described in the above claims, or includes the voice control system described in the above claims, and the vehicle is equipped with the electronic equipment described in the above claims.

[0157] In one embodiment, the vehicle can be configured for fully or partially autonomous driving. For example, the vehicle can control itself while in autonomous driving mode, and can determine the current state of the vehicle and its surrounding environment through human intervention, determine the possible behaviors of at least one other vehicle in the surrounding environment, and determine the confidence level corresponding to the probability of that other vehicle performing a possible behavior, and control the vehicle based on the determined information. When the vehicle is in autonomous driving mode, it can be configured to operate without human interaction.

[0158] The vehicle may also include various subsystems, such as a driving system, sensor system control system, one or more peripheral devices, as well as power supply, computer system, and user interface. Optionally, the vehicle may include more or fewer subsystems, and each subsystem may include multiple components, such as multiple ECUs (electronic control units, i.e., vehicle computers) per subsystem.

[0159] In addition, each subsystem and component of the vehicle can be interconnected via wired or wireless means.

[0160] A propulsion system may include components that provide powered motion to the vehicle. In one embodiment, the propulsion system may include an engine, an energy source, a transmission, and wheels / tires. The engine may be an internal combustion engine, an electric motor, an air-compressed engine, or a combination of other types of engines, such as a hybrid engine consisting of a gasoline engine and an electric motor, or a hybrid engine consisting of an internal combustion engine and an air-compressed engine. The engine converts energy into mechanical energy.

[0161] Examples of energy sources include gasoline, diesel, other petroleum-based fuels, propane, other compressed gas-based fuels, ethanol, solar panels, batteries, and other sources of electricity. Energy sources can also power other systems in the vehicle.

[0162] A transmission system can transmit mechanical power from an engine to the wheels. The transmission system may include a gearbox, a differential, and a drive shaft. In one embodiment, the transmission system may also include other components, such as a clutch. The drive shaft may include one or more axles that can be coupled to one or more wheels.

[0163] A sensor system may include several sensors that sense information about the vehicle's surrounding environment. For example, a sensor system may include a positioning system (which could be GPS, BeiDou, or another positioning system), an inertial measurement unit (IMU), radar, a laser rangefinder, and cameras. The sensor system may also include sensors from the vehicle's internal systems being monitored (e.g., an in-vehicle air quality monitor, fuel gauge, oil temperature gauge, etc.). Sensor data from one or more of these sensors can be used to detect objects and their corresponding characteristics (position, shape, orientation, speed, etc.). This detection and identification is a critical function for the safe operation of autonomous vehicles.

[0164] A positioning system can be used to estimate a vehicle's geographical location. An IMU is used to sense changes in the vehicle's position and orientation based on inertial acceleration. In one embodiment, the IMU can be a combination of an accelerometer and a gyroscope.

[0165] Radar can use radio signals to sense objects in the vehicle's surrounding environment. In some embodiments, in addition to sensing objects, radar can also be used to sense the speed and / or direction of travel of objects.

[0166] A laser rangefinder can use lasers to sense objects in the environment in which a vehicle is located. In some embodiments, a laser rangefinder may include one or more laser sources, a laser scanner, one or more processing modules, and other system components.

[0167] The camera can be used to capture multiple images of the vehicle's surroundings. The camera can be a still camera or a video camera.

[0168] A control system controls the operation of a vehicle and its components. Control systems can include various elements, including steering systems, throttles, braking units, computer vision systems, route control systems, and obstacle avoidance systems.

[0169] The steering system is operable to adjust the vehicle's direction of travel. For example, in one embodiment, it can be a steering wheel system.

[0170] The throttle is used to control the engine's operating speed and, consequently, the vehicle's speed.

[0171] The braking unit is used to control the vehicle's deceleration. The braking unit uses friction to slow down the wheels.

[0172] In other embodiments, the braking unit can convert the kinetic energy of the wheels into electrical current. The braking unit may also take other forms to slow down the wheel rotation speed, thereby controlling the vehicle speed.

[0173] Computer vision systems can be operated to process and analyze images captured by cameras to identify objects and / or features in the environment surrounding a vehicle. These objects and / or features may include traffic signals, road boundaries, and obstacles. Computer vision systems may use object recognition algorithms, structure from motion (SFM) algorithms, video tracking, and other computer vision techniques. In some embodiments, computer vision systems may be used to map the environment, track objects, estimate object velocities, and so on.

[0174] A route control system is used to determine the driving route of a vehicle. In some embodiments, the route control system may combine data from GPS and one or more predetermined maps to determine the driving route for the vehicle.

[0175] Obstacle avoidance systems are used to identify, assess, and avoid or otherwise traverse potential obstacles in the environment in which a vehicle is located.

[0176] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more features. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0177] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0178] The embodiments, implementation methods, and related technical features of this application can be combined and substituted for each other without conflict.

[0179] The above are merely preferred embodiments of this application and are not intended to limit this application in any way. Any simple modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of this application without departing from the scope of the technical solution of this application shall still fall within the scope of the technical solution of this application.

Claims

1. A voice control method, characterized in that, include: Based on the first masking signal obtained from the processing of the initial speech signal and the preset second masking signal, a target masking signal is output, which is used to mask the initial speech signal.

2. The method according to claim 1, characterized in that, The first masking signal is obtained through the following steps: The first masking signal is determined based on the initial speech signal and the reversed speech signal of the initial speech signal.

3. The method according to claim 2, characterized in that, Determining the first masking signal based on the initial speech signal and the reversed speech signal of the initial speech signal includes: The initial speech signal and the reverse speech signal of the initial speech signal are weighted based on a preset first weight to obtain the first masking signal.

4. The method according to claim 3, characterized in that, The method further includes: The first weight is predicted based on a preset objective function and deep learning. The preset objective function is determined based on the speech feature parameters determined by the first masking signal.

5. The method according to claim 4, characterized in that, The speech feature parameters include at least one of masking ratio, clarity index, and speech isolation.

6. The method according to claim 5, characterized in that, The method further includes: The masking signal ratio and the clarity index are determined based on the first masking signal and the initial speech signal; The voice isolation is determined based on the front and rear concealment signals in the first masking signal.

7. The method according to claim 1, characterized in that, The method further includes: The initial voice signal of the target area of ​​the vehicle is determined based on the energy information and / or time delay information of the multi-channel voice signal inside the vehicle.

8. The method according to claim 7, characterized in that, The method further includes: The multi-channel speech signal is preprocessed to determine the initial speech signal of the target area of ​​the vehicle based on the energy information and / or time delay information of the preprocessed multi-channel speech signal; the preprocessing includes at least one of time-frequency domain conversion and filtering.

9. The method according to claim 1, characterized in that, The method further includes: If the initial voice signal meets the preset conditions, the step of outputting the target masking signal based on the first masking signal obtained by processing the initial voice signal and the preset second masking signal is executed.

10. The method according to claim 9, characterized in that, The preset conditions include: The short-time energy of the initial speech signal exceeds a preset energy threshold, and / or the short-time energy of the weighted speech signal determined based on the initial speech signal exceeds a preset energy threshold.

11. The method according to claim 1, characterized in that, The second masking signal includes a natural sound masking signal and / or a human voice masking signal.

12. The method according to claim 11, characterized in that, The method further includes: The target voice masking signal associated with the baseband information is determined based on the baseband information of the initial speech signal.

13. The method according to claim 11, characterized in that, The method further includes: In response to the detection of the initial speech signal, the voice masking signal is output.

14. The method according to claim 1, characterized in that, The method outputs a target masking signal based on a first masking signal obtained from processing the initial speech signal and a preset second masking signal. The target masking signal is used to mask the initial speech signal, including: The first masking signal and the second masking signal are weighted based on a preset second weight to obtain the target masking signal.

15. The method according to claim 14, characterized in that, The method further includes: The second weight is predicted based on a preset objective function and through deep learning.

16. The method according to any one of claims 1 to 15, characterized in that, The method further includes: The target masking signal is output based on the multi-channel voice output device.

17. The method according to claim 16, characterized in that, The method further includes: The target masking signal is output based on the prediction of the weight and / or phase of the speech signal output by each channel in the multi-channel speech output device using deep learning.

18. A voice control device, characterized in that, It includes an output module, which is used to perform the voice control method as described in any one of claims 1 to 17.

19. A voice control system, characterized in that, include: Audio acquisition device, used to acquire voice signals inside the vehicle; A controller is configured to execute the voice control method as described in any one of claims 1 to 17 to process the voice signal and obtain a target masking signal; An audio output device for outputting the target masking signal.

20. An electronic device, characterized in that, The device includes a processor; a memory for storing processor-executable instructions; wherein the processor is configured to perform the steps of the voice control method according to any one of claims 1 to 17.

21. A computer-readable storage medium, characterized in that, It stores a computer program, which is loaded by a processor to execute the steps of the voice control method according to any one of claims 1 to 17.

22. A vehicle, characterized in that, Voice control is performed by executing the voice control method as described in any one of claims 1 to 17, or includes the voice control device as described in claim 18, or includes the voice control system as described in claim 19.