In-vehicle call privacy control method and device, equipment and medium
Through the acoustic masking technology processing of the in-car call system, the voice signals of the target call user are obtained and converted, and the speaker array is driven to output the sound masking signal, solving the problem of poor white noise masking effect, realizing effective masking of leaked voices, and improving call privacy.
Patent Information
- Application Number
- CN202410138905.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-31
- Publication Date
- 2025-08-01
AI Technical Summary
In the existing in-car call system, white noise at non-target locations cannot effectively mask the leaked voice signals of the target call user, making it easy for other passengers to understand the call content and cannot guarantee the privacy of the call content.
By obtaining the pure voice signal of the target call user, performing frame-based windowing and short-time Fourier transformation, the time frequency domain signal is obtained, and the time inversion process is performed to convert it into a time domain masking signal, and the speaker array in the non-target area outputs acoustic masking signals, and using acoustic masking technology to improve the masking effect.
Effectively reduce the voice intelligibility of non-target areas, improve the privacy and security of call content of target users, and ensure that call content in target areas is not understood by passengers in non-target areas.
Smart Images

Figure CN120412518A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of in-vehicle calls, and particularly relates to a method, device, equipment and medium for controlling the privacy of in-vehicle calls. Background Technique
[0002] Currently, in order to control the privacy of in-vehicle calls, most vehicles adopt a distributed layout of microphones and speakers, that is, there is at least one microphone and one speaker near each seat. While the target call user is making a normal call through an electronic device such as a mobile phone, random white noise is played through the speaker at non-target positions outside the target call user.
[0003] In the related technology, when the target call user is making a normal call, although the sound pressure level of the leaked voice at non-target positions is small, the masking effect of the white noise on the leaked voice signal of the target call user is not good. Therefore, other passengers with good hearing can still relatively easily understand the call content, and the privacy of the call content of the target call user cannot be guaranteed. Summary of the Invention
[0004] The purpose of the embodiments of this application is to provide a method and device for controlling the privacy of in-vehicle calls, which can reduce the intelligibility of the leaked voice signal in the non-target call user area, improve the sound masking effect, and enhance the privacy security of the call content of the target call user.
[0005] In a first aspect, the embodiments of this application provide a method for controlling the privacy of in-vehicle calls, and the method includes: obtaining the pure voice signal of the target call user; performing frame addition and windowing processing and short-time Fourier transform on the pure voice signal to obtain a time-frequency domain signal; performing time reversal processing on the time-frequency domain signal to obtain a time-frequency domain masking signal; performing inverse short-time Fourier transform on the time-frequency domain masking signal to convert the time-frequency domain masking signal into a time-domain masking signal; driving the speaker array in the non-target call user area to output the time-domain masking signal.
[0006] In some realizable ways of the first aspect, after converting the time-frequency domain masking signal into a time-domain masking signal, the method further includes: performing filtering processing on the time-domain masking signal by using an adaptive filtering algorithm; driving the speaker array in the non-target call user area to output the time-domain masking signal, including: driving the speaker array in the non-target call user area to output the filtered time-domain masking signal.
[0007] In some realizable ways of the first aspect, obtaining the pure voice signal of the target call user includes: obtaining the original voice signal of the target call user; performing echo cancellation processing, voice separation processing, and voice enhancement processing on the original voice signal to obtain the pure voice signal.
[0008] In some realizable ways of the first aspect, obtaining the original voice signal of the target call user includes: dividing the vehicle interior space into multiple regions through a microphone array; combining the microphone array and the camera to monitor call users in the multiple regions; when it is detected that there is a target call user in a target region among the multiple regions, determining the target region as a private call region, and collecting the original voice signal of the target call user through the microphone array in the private call region.
[0009] In some realizable ways of the first aspect, combining the microphone array and the camera to monitor call users in the multiple regions includes: controlling the microphone array to monitor the multiple regions in real time, performing sound source localization on the target region that generates a voice signal; calling the camera to perform auxiliary monitoring on the target region to determine whether there is a target call user in the target region.
[0010] In some realizable ways of the first aspect, the method is applied to a vehicle, and the vehicle is provided with a target privacy control. Before obtaining the pure voice signal of the target call user, the method further includes: receiving a touch input of the target call user on the target privacy control; in response to the touch input, turning on the private call mode of the vehicle; driving the speaker array in the non-target call user region to output a time-domain masking signal, including: only in the private call mode, driving the speaker array in the non-target call user region to output a time-domain masking signal.
[0011] In some realizable ways of the first aspect, performing echo cancellation processing, voice separation processing, and voice enhancement processing on the original voice signal to obtain a pure voice signal includes: performing echo cancellation processing, voice separation processing, voice enhancement processing, and feedback suppression processing on the original voice signal to obtain a pure voice signal.
[0012] In a second aspect, an embodiment of the present application provides an in-vehicle call privacy control device, which includes: an acquisition module for acquiring the pure voice signal of the target call user; a masking processing module for performing frame addition and windowing processing and short-time Fourier transform on the pure voice signal to obtain a time-frequency domain signal; the masking processing module is further used for performing time reversal processing on the time-frequency domain signal to obtain a time-frequency domain masking signal; the masking processing module is further used for performing inverse short-time Fourier transform on the time-frequency domain masking signal to convert the time-frequency domain masking signal into a time-domain masking signal; a driving control module for driving the speaker array in the non-target call user region to output the time-domain masking signal.
[0013] In some realizable ways of the second aspect, the device further includes: a filtering module for, after converting the time-frequency domain masking signal into a time-domain masking signal, performing filtering processing on the time-domain masking signal using an adaptive filtering algorithm; the driving control module is specifically used for: driving the speaker array in the non-target call user region to output the filtered time-domain masking signal.
[0014] In some realizable ways of the second aspect, the acquisition module includes: an acquisition unit configured to acquire the original voice signal of the target call user; and a voice processing unit configured to perform echo cancellation processing, voice separation processing, and voice enhancement processing on the original voice signal to obtain a pure voice signal.
[0015] In some realizable ways of the second aspect, the acquisition module includes: a division unit configured to divide the vehicle interior space into multiple regions through a microphone array; a monitoring unit configured to monitor call users in the multiple regions in combination with the microphone array and a camera; and an acquisition unit configured to, when it is detected that a target call user exists in a target region among the multiple regions, determine the target region as a private call region, and acquire the original voice signal of the target call user through the microphone array in the private call region.
[0016] In some realizable ways of the second aspect, the monitoring unit is specifically configured to: monitor the multiple regions in real time through the microphone array, perform sound source localization on the target region where a voice signal is generated; and call the camera to perform auxiliary monitoring on the target region to determine whether a target call user exists in the target region.
[0017] In some realizable ways of the second aspect, the method is applied to a vehicle, and the vehicle is provided with a target privacy control. Before acquiring the pure voice signal of the target call user, the device further includes: a receiving module configured to receive a touch input of the target call user on the target privacy control; an enabling module configured to, in response to the touch input, enable the private call mode of the vehicle; and the driving control module is specifically configured to: only in the private call mode, drive the speaker array in the non-target call user region to output a time-domain masking signal.
[0018] In some realizable ways of the second aspect, the voice processing unit is specifically configured to: perform echo cancellation processing, voice separation processing, voice enhancement processing, and feedback suppression processing on the original voice signal to obtain a pure voice signal.
[0019] In a third aspect, an embodiment of the present application provides an electronic device, including: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the steps of the in-vehicle call privacy control method as in the first aspect are implemented.
[0020] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the steps of the in-vehicle call privacy control method as in the first aspect are implemented.
[0021] In a fifth aspect, an embodiment of the present application provides a computer program product, which is stored in a non-volatile storage medium and is executed by at least one processor to implement the steps of the in-vehicle call privacy control method as in the first aspect.
[0022] In a sixth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the steps of the in-vehicle call privacy control method as in the first aspect.
[0023] The present application provides an in-vehicle call privacy control method, device, equipment and medium. A pure voice signal of a target call user is obtained, the pure voice signal is subjected to frame addition and windowing processing and short-time Fourier transform to obtain a time-frequency domain signal, then the time-frequency domain signal is subjected to time reversal processing to obtain a time-frequency domain masking signal, and then the time-frequency domain masking signal is subjected to inverse short-time Fourier transform to convert the time-frequency domain masking signal into a time-domain masking signal, that is, an acoustic masking signal. In this way, based on the pure voice signal, a series of signal processing operations are performed on the pure voice signal based on the acoustic masking technology to obtain an acoustic masking signal, and the acoustic masking signal is output by driving a speaker array in the non-target call user area to mask the leaked voice signal of the target call user and reduce the intelligibility index of the leaked voice signal. Compared with using white noise to mask the leaked voice signal, the acoustic masking signal of the present application is obtained based on the pure voice signal of the target call user. There will be similarity between the acoustic masking signal and the pure voice signal in the power spectrum, and there is also similarity between the acoustic masking signal and the leaked voice signal in the power spectrum. Therefore, when using the acoustic masking signal to mask the leaked voice signal, the information masking amount can be effectively improved, thereby improving the masking effect on the leaked voice signal, effectively reducing the intelligibility index of the leaked voice of the target call user, and finally achieving that the intelligibility in the target call user area remains almost unchanged, and the intelligibility in the non-target call user area decreases significantly, improving the privacy and security of the call content of the target call user. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required to be used in the embodiments of the present application.
[0025] Figure 1 is a schematic flowchart of an in-vehicle call privacy control method provided by an embodiment of the present application;
[0026] Figure 2 is a schematic flowchart of an in-vehicle call privacy control method provided by another embodiment of the present application;
[0027] Figure 3 is a schematic flowchart of an in-vehicle call privacy control method provided by still another embodiment of the present application;
[0028] Figure 4 It is an exemplary schematic diagram of an in-vehicle call privacy control scenario provided by an embodiment of the present application;
[0029] Figure 5 It is a structural schematic diagram of an in-vehicle call privacy control device provided by an embodiment of the present application;
[0030] Figure 6 It is a structural schematic diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0031] The features and exemplary embodiments of various aspects of the present application will be described in detail below. To make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than limiting the present application. For those skilled in the art, the present application can be implemented without some of these specific details. The following description of the embodiments is only intended to provide a better understanding of the present application by showing examples of the present application.
[0032] Currently, in order to control the privacy of in-vehicle calls, most vehicles adopt a distributed layout of microphones and speakers, that is, there is at least one microphone and one speaker near each seat. While the target call user is making a normal call through an electronic device such as a mobile phone, random white noise is played through the speaker at non-target positions other than the target call user. When the target call user is making a normal call, although the sound pressure level of the leaked voice at non-target positions is small, the white noise cannot provide good masking ability to mask the leaked voice signal of the target call user. Therefore, other passengers with good hearing can still relatively easily understand the call content, and the sound masking effect is not good, unable to ensure the privacy of the call content of the target call user.
[0033] The in-vehicle call privacy control method provided by the embodiments of the present application will be described in detail below with reference to the accompanying drawings, specific embodiments and their application scenarios.
[0034] Figure 1 It is a flowchart of the in-vehicle call privacy control method provided by an embodiment of the present application. The execution subject of this in-vehicle call privacy control method can be the vehicle end or the cloud server end. Specifically, the execution subject can be the processor of the vehicle end or the cloud server end.
[0035] The in-vehicle call privacy control method of the present application will be described below by taking the vehicle as the execution subject of the in-vehicle call privacy control method as an example. It should be noted that the above execution subject and application scenario do not constitute a limitation to the present application.
[0036] As shown Figure 1 in the figure, the in-vehicle call privacy control method provided by the embodiment of the present application may include step 110-step 150.
[0037] Step 110: Obtain the pure voice signal of the target call user;
[0038] Step 120: Perform frame addition and windowing processing and short-time Fourier transform on the pure voice signal to obtain a time-frequency domain signal;
[0039] Step 130: Perform time reversal processing on the time-frequency domain signal to obtain a time-frequency domain masking signal;
[0040] Step 140: Perform inverse short-time Fourier transform on the time-frequency domain masking signal to convert the time-frequency domain masking signal into a time-domain masking signal;
[0041] Step 150: Drive the speaker array in the non-target call user area to output the time-domain masking signal.
[0042] For the in-vehicle call privacy control method provided by the embodiment of the present application, the pure voice signal of the target call user is obtained, the pure voice signal is subjected to frame addition and windowing processing and short-time Fourier transform to obtain a time-frequency domain signal, then the time-frequency domain signal is subjected to time reversal processing to obtain a time-frequency domain masking signal, and then the inverse short-time Fourier transform is performed on the time-frequency domain masking signal to convert the time-frequency domain masking signal into a time-domain masking signal, that is, an acoustic masking signal. In this way, based on the pure voice signal, a series of signal processing operations are performed on the pure voice signal based on the acoustic masking technology to obtain an acoustic masking signal, and the acoustic masking signal is output by driving the speaker array in the non-target call user area to mask the leaked voice signal of the target call user and reduce the intelligibility index of the leaked voice signal. Compared with using white noise to mask the leaked voice signal, the acoustic masking signal of the present application is obtained based on the pure voice signal of the target call user, and there will be similarity between the acoustic masking signal and the pure voice signal in the power spectrum, and there is also similarity between the acoustic masking signal and the leaked voice signal in the power spectrum. Therefore, when using the acoustic masking signal to mask the leaked voice signal, the information masking amount can be effectively improved, and then the masking effect on the leaked voice signal can be improved, effectively reducing the intelligibility index of the leaked voice of the target call user, and finally achieving that the intelligibility in the target call user area remains almost unchanged, the intelligibility in the non-target call user area drops significantly, and the privacy security of the call content of the target call user is improved.
[0043] The following will introduce the specific implementation manners of the above steps in detail in combination with specific embodiments.
[0044] Regarding step 110, obtain the pure voice signal of the target call user.
[0045] In step 110, the target call user is the user in the vehicle who is having a voice call or a video call. The pure voice signal is obtained by performing noise reduction processing on the original voice signal of the target call user, and the original voice signal can be collected by a microphone within the area of the target call user.
[0046] In some embodiments of the present application, Figure 2 is a schematic flowchart of an in-vehicle call privacy control method provided in another embodiment of the present application. The above step 110 may specifically include Figure 2 the steps 210 and 220 shown in
[0047] Step 210, obtaining the original voice signal of the target call user;
[0048] Step 220, performing echo cancellation processing, voice separation processing, and voice enhancement processing on the original voice signal to obtain a pure voice signal.
[0049] In the embodiments of the present application, voice separation and enhancement can extract the voice signal of the target call user and improve the signal-to-noise ratio of the signal to be used as a reference input signal for the time-domain masking signal; while echo cancellation can eliminate the remote call voice played through the speaker array, improve the effect and quality of the output voice of the target call user, and enhance the user call experience.
[0050] In some embodiments of the present application, in order to accurately identify the privacy call scenario and accurately collect the voice signal of the call user, the above step 210 of obtaining the original voice signal of the target call user may specifically include the following steps:
[0051] Dividing the in-vehicle space into multiple regions by a microphone array;
[0052] Combining the microphone array and a camera to monitor call users in multiple regions;
[0053] When it is detected that there is a target call user in the target region among the multiple regions, determining the target region as a privacy call region, and collecting the original voice signal of the target call user through the microphone array within the privacy call region.
[0054] Specifically, the in-vehicle space is divided into multiple regions by the microphone array, and the multiple regions are multiple sound zones. When it is detected that there is a target call user in the target region among the multiple regions, determining the target region as a privacy call region, which is the region where the target call user is located, that is, the target call user region, and the non-target call user region is the region other than the target region among the multiple regions.
[0055] In some embodiments of the present application, the above-mentioned monitoring of call users in multiple regions by combining a microphone array and a camera may specifically include:
[0056] Real-time monitor multiple regions through the microphone array, and perform sound source localization on the target region that generates a voice signal;
[0057] Call the camera to assist in monitoring the target region to determine whether there is a target call user in the target region.
[0058] Specifically, calling the camera to assist in monitoring the target region may specifically include: obtaining image information corresponding to the target region; performing image recognition on the image information to determine whether the image information meets the preset privacy call condition; when the image information meets the preset privacy call condition, determining that there is a target call user in the target region, and turning on the privacy call mode of the vehicle, where only in the privacy call mode, the speaker array in the non-target call user region is driven to output the time-domain masking signal; when the image information does not meet the preset privacy call condition, determining that there is no target call user in the target region.
[0059] The preset privacy call condition may include that the speaking user holds an electronic device and the electronic device is close to the ear, or the speaking user wears in-ear devices, or there are no other users within a preset range of the speaking user. The preset range may be centered on the user's location and radiate a preset distance to the surrounding area. The preset distance can be set according to specific requirements, and the present application does not make specific limitations on this.
[0060] In the embodiments of the present application, through the microphone array, in the scenario where a target call user in the target region makes a voice call or a video call, the voice signal can be timely sensed, and then the camera is called to assist in monitoring the target region to determine whether there is really a call user with a privacy call requirement in the target region, accurately identify the privacy call scenario, and perform sound masking in this privacy call scenario, avoiding turning on the privacy call control when the user talks to other users in the vehicle and causing trouble to the users in the vehicle.
[0061] In some embodiments of the present application, this method can be applied to a vehicle. The vehicle is provided with a target privacy control. Before obtaining the pure voice signal of the target call user in step 110 above, this method may further include the following steps:
[0062] Receive a touch input from the target call user to the target privacy control;
[0063] In response to the touch input, turn on the privacy call mode of the vehicle;
[0064] Driving the speaker array in the non-target call user region to output a time-domain masking signal includes:
[0065] Only in the private call mode, drive the speaker array in the area of non-target call users to output a time-domain masking signal.
[0066] Specifically, the target privacy control is used to trigger the vehicle to enter the private call mode. In this private call mode, the speaker array in the area of non-target call users can be driven to output a time-domain masking signal. The target privacy control can be a physical button installed on the vehicle or a touch button displayed on the touch screen. Therefore, the touch input can be the pressing operation of the physical button by the target call user, or the touch screen input to the touch button. Such touch screen input can be, for example, click input, swipe input, double-click input, etc. This application does not make specific limitations on this.
[0067] In the embodiments of this application, by setting a target privacy control on the vehicle, when a passenger (i.e., the target call user) makes a voice call or a video call through an electronic device in the vehicle, if the passenger does not want the other passengers to hear the content of the call, the passenger can turn on the private call mode through the touch input to the target privacy control. In this private call mode, the vehicle can drive the speaker array in the area of non-target call users to output a time-domain masking signal, that is, perform acoustic masking on the area other than the calling passenger (i.e., the area of non-target call users), reduce the intelligibility of the call voice emitted by the target call user, and improve the privacy and security of the call content. In some embodiments of this application, the above step 220 performs echo cancellation processing, voice separation processing, and voice enhancement processing on the original voice signal to obtain a pure voice signal, which may specifically include:
[0068] Perform echo cancellation processing on the original voice signal by using an adaptive filtering algorithm or a convolutional neural network algorithm;
[0069] Perform voice separation processing and voice enhancement processing on the original voice signal by using any one of Wiener filtering method, spectral subtraction method, recurrent neural network algorithm, and convolutional neural network algorithm.
[0070] Among them, the adaptive filtering algorithm may include at least one of the Least Mean Square (LMS) algorithm and the Multidelay Block Frequency Domain Adaptive Filter (MDF) algorithm; the LMS algorithm is an improved algorithm of the steepest descent algorithm, which is an optimized extension after applying the steepest descent method in the Wiener filtering theory. It has characteristics such as low computational complexity, good convergence in an environment where the signal is a stationary signal, its expected value converging unbiasedly to the Wiener solution, and stability when implementing the algorithm with finite precision, making the LMS algorithm the most stable and widely used algorithm among adaptive algorithms; the MDF algorithm is a method that divides the original multi-order filter into K equal sub-blocks and can perform adaptive filtering on each sub-block with a length of N.
[0071] In some embodiments of the present application, the above-mentioned processing of the original speech signal for echo cancellation, speech separation, and speech enhancement to obtain a pure speech signal may specifically include:
[0072] Perform echo cancellation processing, speech separation processing, speech enhancement processing, and feedback suppression processing on the original speech signal to obtain a pure speech signal.
[0073] In some embodiments of the present application, the above-mentioned processing of the original speech signal for echo cancellation, speech separation, speech enhancement, and feedback suppression may specifically include:
[0074] Use a howling suppression algorithm to perform feedback suppression processing on the original speech signal after speech enhancement processing.
[0075] In the embodiments of the present application, the feedback suppression processing of the original speech signal can be implemented by a howling suppression algorithm, which may be an adaptive filtering algorithm, a phase interference method, a convolutional neural network algorithm, etc. The present application does not make specific limitations in this regard. By performing feedback suppression processing on the original speech signal, howling caused by self-excited amplification can be prevented, and the user's call experience can be improved.
[0076] Referring to step 120, perform frame addition and windowing processing and short-time Fourier transform on the pure speech signal to obtain a time-frequency domain signal.
[0077] In step 120, the purpose of voice signal framing is to divide a number of voice sampling points into one frame. Within this frame, the characteristics of the voice signal can be regarded as stable. Generally, the length of voice framing is about 10 - 40 ms. Windowing is achieved through a window function. Different window functions have different degrees of alleviating spectral leakage, and its total leakage is measured by the equivalent noise bandwidth (ENBW). A good window function design should ensure that the energy of the spectrum is mainly concentrated in the main lobe, and the energy of the sidelobes is minimized as much as possible, so that the signal within the window is approximately periodic. The Short Time Fourier Transform (STFT) is a general tool for voice signal processing. The basic idea of STFT is to divide the time-domain signal into many small time windows and perform Fourier transform on each time window, so as to obtain the spectral information of each time window, thereby revealing the spectral characteristics of the signal in different time periods.
[0078] In the embodiments of the present application, the characteristics of the voice signal change over time and are a non-stationary random process. However, on the other hand, although voice has time-varying properties, its characteristics can be regarded as stable within a short period. This is because the muscle inertia of the human vocal organs makes it impossible to instantaneously switch from one state to another. The fact that voice characteristics remain unchanged within a short period is called the short-time stationary characteristic of voice, and the short-time analysis technology runs through the entire process of the voice signal. The short-time method is the key to analyzing non-stationary signals using the processing method of stationary signals, and the process of short-time analysis generally includes steps of framing, windowing, and DFT.
[0079] Regarding step 130, the time-frequency domain signal is subjected to time reversal processing to obtain a time-frequency domain masking signal.
[0080] In step 130, time reversal processing is time reversal signal processing. Time reversal signal processing means that the received target reflected echo time-domain signal is reversed in time sequence to obtain a backward transmission signal, which is then transmitted into the calculation area where the target is located, equivalent to the "last-in, first-out" of the signal. After the inverse wave is virtually retransmitted, since the clutter environment experienced by the inverse wave during backpropagation is the same as that of the actual received echo, the signal will achieve energy focusing at the target position. This makes the time reversal technology have the characteristics of anti-multipath and can focus synchronously in time and space, that is, the spatio-temporal focusing characteristics possessed by the time reversal technology.
[0081] Regarding step 140, the inverse short-time Fourier transform is performed on the time-frequency domain masking signal to convert the time-frequency domain masking signal into a time-domain masking signal.
[0082] In some embodiments of the present application, Figure 3It is a schematic flowchart of the in-vehicle call privacy control method provided by another embodiment of this application. After the above step 140, the method may further include Figure 3 step 310 shown in the figure, and step 140 may specifically be step 320.
[0083] Step 310, filtering the time-domain masking signal by using an adaptive filtering algorithm;
[0084] Step 320, driving the speaker array in the non-target call user area to output the filtered time-domain masking signal.
[0085] In the embodiments of this application, the frequency response and gain of the time-domain masking signal can be adjusted through the adaptive filtering algorithm, avoiding the time-domain masking signal from being too noisy, so as to ensure that while the passengers in other positions cannot understand the call content, the impact on their driving and riding experience is reduced.
[0086] In some embodiments of this application, the adaptive filtering algorithm may include at least one of the least mean square (LMS) algorithm and the multi-delay block frequency domain (MDF) algorithm.
[0087] It relates to step 150, driving the speaker array in the non-target call user area to output the time-domain masking signal.
[0088] In step 150, the target call user area is the area where the target call user is located, and the non-target call user area is the area other than the target call area. The users in the non-target call user area are non-call users. Therefore, it is necessary to play the time-domain masking signal for the users in this area to prevent the call content of the target call user from being known by the non-call users in the vehicle.
[0089] Exemplarily, a distributed speaker array and a distributed microphone array as shown in Figure 4 the figure may be provided on the vehicle. The target bright area is the target call user area, and the non-target dark area is the non-target call user area. After the original voice signal of the target call user in the target bright area is collected by the distributed microphone array, the digital signal processing chip processes the original voice signal, extracts the pure voice signal from it, and further converts the pure voice signal into a time-domain masking signal. Thus, through the distributed speaker array, the time-domain masking signal is played in all non-target dark areas, interfering with the non-target call users through the time-domain masking signal, covering the output voice signal of the target call user, and reducing the intelligibility index of the call content output by the target call user.
[0090] It can be understood that for the in-vehicle call privacy control method provided by the embodiments of this application, the execution subject may be a terminal, or a control module in the terminal for executing the in-vehicle call privacy control method. The in-vehicle call privacy control device will be introduced in detail below.
[0091] Figure 5 This is a schematic structural diagram of an in-vehicle call privacy control device provided by an embodiment of the present application. As Figure 5 shown, the in-vehicle call privacy control device 500 may include: an acquisition module 510, a masking processing module 520, and a drive control module 530.
[0092] Among them, the acquisition module 510 is used to acquire the pure voice signal of the target call user; the masking processing module 520 is used to perform frame-by-frame windowing processing and short-time Fourier transform on the pure voice signal to obtain a time-frequency domain signal; the masking processing module 520 is further used to perform time reversal processing on the time-frequency domain signal to obtain a time-frequency domain masking signal; the masking processing module 520 is further used to perform inverse short-time Fourier transform on the time-frequency domain masking signal to convert the time-frequency domain masking signal into a time-domain masking signal; the drive control module 530 is used to drive the speaker array in the non-target call user area to output the time-domain masking signal.
[0093] The in-vehicle call privacy control device provided by the present application acquires the pure voice signal of the target call user, performs frame-by-frame windowing processing and short-time Fourier transform on the pure voice signal to obtain a time-frequency domain signal, then performs time reversal processing on the time-frequency domain signal to obtain a time-frequency domain masking signal, and then performs inverse short-time Fourier transform on the time-frequency domain masking signal to convert the time-frequency domain masking signal into a time-domain masking signal, that is, an acoustic masking signal. In this way, based on the pure voice signal, a series of signal processing operations are performed on the pure voice signal based on the acoustic masking technology to obtain an acoustic masking signal, and the acoustic masking signal is output by driving the speaker array in the non-target call user area to mask the leaked voice signal of the target call user and reduce the intelligibility index of the leaked voice signal. Compared with using white noise to mask the leaked voice signal, the acoustic masking signal of the present application is obtained based on the pure voice signal of the target call user. There will be similarity between the acoustic masking signal and the pure voice signal in the power spectrum, and there is also similarity between the acoustic masking signal and the leaked voice signal in the power spectrum. Therefore, when using the acoustic masking signal to mask the leaked voice signal, the information masking amount can be effectively improved, thereby improving the masking effect on the leaked voice signal, effectively reducing the intelligibility index of the leaked voice of the target call user, and finally realizing that the intelligibility in the target call user area remains almost unchanged, and the intelligibility in the non-target call user area decreases significantly, improving the privacy and security of the call content of the target call user.
[0094] In some embodiments of the present application, the device further includes: a filtering module, which is used to perform filtering processing on the time-domain masking signal by using an adaptive filtering algorithm after converting the time-frequency domain masking signal into a time-domain masking signal; the drive control module 530 is specifically used to: drive the speaker array in the non-target call user area to output the filtered time-domain masking signal.
[0095] In some embodiments of the present application, the adaptive filtering algorithm includes at least one of the least mean square (LMS) algorithm and the multi-delay block frequency domain (MDF) algorithm.
[0096] In some embodiments of the present application, the acquisition module 510 includes: an acquisition unit configured to acquire the original voice signal of the target call user; and a voice processing unit configured to perform echo cancellation processing, voice separation processing, and voice enhancement processing on the original voice signal to obtain a pure voice signal.
[0097] In some implementable ways of the second aspect, the acquisition module 510 includes: a division unit configured to divide the vehicle interior space into multiple regions through a microphone array; a monitoring unit configured to monitor call users in the multiple regions in combination with the microphone array and a camera; and an acquisition unit configured to, when it is detected that a target call user exists in a target region among the multiple regions, determine the target region as a private call region and acquire the original voice signal of the target call user through the microphone array in the private call region.
[0098] In some implementable ways of the second aspect, the monitoring unit is specifically configured to: monitor the multiple regions in real time through the control of the microphone array, perform sound source localization on the target region that generates a voice signal; and call the camera to perform auxiliary monitoring on the target region to determine whether a target call user exists in the target region.
[0099] In some implementable ways of the second aspect, the method is applied to a vehicle, and the vehicle is provided with a target privacy control. Before acquiring the pure voice signal of the target call user, the device further includes: a receiving module configured to receive a touch input of the target call user on the target privacy control; an enabling module configured to, in response to the touch input, enable the private call mode of the vehicle; and the drive control module 530 is specifically configured to: only in the private call mode, drive the speaker array in the non-target call user region to output a time-domain masking signal.
[0100] In some embodiments of the present application, the voice processing unit is specifically configured to: perform echo cancellation processing on the original voice signal by using an adaptive filtering algorithm or a convolutional neural network algorithm; and perform voice separation processing and voice enhancement processing on the original voice signal by using any one of a Wiener filtering method, a spectral subtraction method, a recurrent neural network algorithm, and a convolutional neural network algorithm.
[0101] In some embodiments of the present application, the voice processing unit is specifically configured to: perform echo cancellation processing, voice separation processing, voice enhancement processing, and feedback suppression processing on the original voice signal to obtain a pure voice signal.
[0102] In some embodiments of the present application, the voice processing unit is specifically configured to: perform feedback suppression processing on the original voice signal after voice enhancement processing by using a howling suppression algorithm.
[0103] The in-vehicle call privacy control device provided by the embodiments of the present application can implement Figures 1-3 each process implemented by the electronic device in the method embodiments and can achieve the same technical effects. To avoid repetition, they will not be elaborated here.
[0104] Figure 6 It is a schematic diagram of the hardware structure of an electronic device provided by the embodiments of the present application.
[0105] As Figure 6 shown, the electronic device 600 includes a memory 601, a processor 602, and a computer program stored on the memory 601 and executable on the processor 602.
[0106] In one example, the above-mentioned processor 602 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0107] The memory 601 may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage medium device, an optical storage medium device, a flash memory device, an electrical, optical, or other physical / tangible memory storage device. Therefore, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (such as memory devices) encoded with software including computer-executable instructions, and when the software is executed (for example, by one or more processors), it is operable to perform the operations described with reference to the card opening method in the embodiments of the first aspect of the present application.
[0108] The processor 602 runs the computer program corresponding to the executable program code by reading the executable program code stored in the memory 601 to implement the card opening method in the embodiments of the first aspect above.
[0109] In some examples, the electronic device 600 may further include a communication interface 603 and a bus 604. Among them, as Figure 6 shown, the memory 601, the processor 602, and the communication interface 603 are connected through the bus 604 and complete communication with each other.
[0110] The communication interface 603 is mainly used to implement communication between various modules, devices, units, and / or equipment in the embodiments of the present application. The input device and / or output device can also be accessed through the communication interface 603.
[0111] The bus 604 includes hardware, software, or both, and couples the components of the electronic device 600 to each other. By way of example and not limitation, the bus 604 can include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-E) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses or a combination of two or more of these. In a suitable case, the bus 604 can include one or more buses. Although the embodiments of the present application describe and illustrate specific buses, the present application contemplates any suitable bus or interconnect.
[0112] The electronic device provided by the embodiments of the present application can implement Figures 1-3 each process implemented by the electronic device in the method embodiments, and can achieve the same technical effects. To avoid repetition, it will not be described in detail here.
[0113] Combined with the in-vehicle call privacy control method in the above embodiments, the embodiments of the present application can provide a computer storage medium to implement. Computer program instructions are stored on the computer storage medium; when the computer program instructions are executed by a processor, the steps of any one of the in-vehicle call privacy control methods in the above embodiments are implemented.
[0114] Combined with the in-vehicle call privacy control method in the above embodiments, an embodiment of the present application can provide a computer program product to implement. The (computer) program product is stored in a non-volatile storage medium, and when the program product is executed by at least one processor, it implements the steps of any one of the in-vehicle call privacy control methods in the above embodiments.
[0115] Another embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement each process of the above embodiment of the in-vehicle call privacy control method, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0116] It should be understood that the chip mentioned in the embodiment of the present application can also be referred to as a system-level chip, system chip, chip system, or system-on-chip, etc.
[0117] It should be clear that the present application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, the detailed description of known methods is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present application.
[0118] It should also be noted that the functional blocks shown in the above structure block diagrams can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted through a data signal carried in a carrier wave on a transmission medium or a communication link. A "machine-readable medium" can include any medium that can store or transmit information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical discs, hard disks, fiber optic media, radio frequency (RF) links, and so on. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.
[0119] It also needs to be noted that the exemplary embodiments mentioned in the present application describe some methods or systems based on a series of steps or devices. However, the present application is not limited to the order of the above steps. That is, the steps can be executed in the order mentioned in the embodiments, or different from the order in the embodiments, or several steps can be executed simultaneously.
[0120] As described above with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block in the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the functions / actions specified in one or more blocks of the flowcharts and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It should also be understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by dedicated hardware that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0121] As described above, this is only the specific implementation manner of the present application. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein. It should be understood that the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application.
Claims
1. A method for controlling in-vehicle call privacy, characterized in that, The method includes: Obtaining a pure voice signal of a target call user; Performing frame addition and windowing processing and short-time Fourier transform on the pure voice signal to obtain a time-frequency domain signal; Performing time reversal processing on the time-frequency domain signal to obtain a time-frequency domain masking signal; Performing inverse short-time Fourier transform on the time-frequency domain masking signal to convert the time-frequency domain masking signal into a time-domain masking signal; Driving a speaker array in a non-target call user area to output the time-domain masking signal.
2. The method according to claim 1, characterized in that, After converting the time-frequency domain masking signal into a time-domain masking signal, the method further includes: Filtering the time-domain masking signal by using an adaptive filtering algorithm; The driving the speaker array in the non-target call user area to output the time-domain masking signal includes: Driving the speaker array in the non-target call user area to output the filtered time-domain masking signal.
3. The method according to claim 1, characterized in that, The obtaining a pure voice signal of a target call user includes: Obtaining an original voice signal of the target call user; Performing echo cancellation processing, voice separation processing, and voice enhancement processing on the original voice signal to obtain the pure voice signal.
4. The method according to claim 3, wherein The obtaining an original voice signal of the target call user includes: Dividing the vehicle interior space into multiple areas by using a microphone array; Combining the microphone array and a camera to monitor call users in the multiple areas; When it is monitored that the target call user exists in a target area among the multiple areas, determining the target area as a private call area, and collecting the original voice signal of the target call user through the microphone array in the private call area.
5. The method according to claim 4, wherein The combining the microphone array and the camera to monitor call users in the multiple areas includes: Real-time monitoring the multiple areas by controlling the microphone array, and performing sound source localization on a target area where a voice signal is generated; Invoking the camera to perform auxiliary monitoring on the target area to determine whether the target call user exists in the target area.
6. The method according to claim 1, wherein The method is applied to a vehicle, and the vehicle is provided with a target privacy control. Before obtaining a pure voice signal of a target call user, the method further includes: Receiving a touch input of the target call user on the target privacy control; Responding to the touch input, and turning on the private call mode of the vehicle; The driving the speaker array in the non-target call user area to output the time-domain masking signal includes: Only in the private call mode, driving the speaker array in the non-target call user area to output the time-domain masking signal.
7. The method according to claim 3, characterized in that, The performing echo cancellation processing, voice separation processing, and voice enhancement processing on the original voice signal to obtain the pure voice signal includes: Performing echo cancellation processing, voice separation processing, voice enhancement processing, and feedback suppression processing on the original voice signal to obtain the pure voice signal.
8. An in-vehicle call privacy control device, characterized in that, The device includes: An obtaining module, configured to obtain a pure voice signal of a target call user; A masking processing module, configured to perform frame addition and windowing processing and short-time Fourier transform on the pure voice signal to obtain a time-frequency domain signal; The masking processing module is further configured to perform time reversal processing on the time-frequency domain signal to obtain a time-frequency domain masking signal; The masking processing module is further configured to perform an inverse short-time Fourier transform on the time-frequency domain masking signal to convert the time-frequency domain masking signal into a time-domain masking signal; The driving control module is configured to drive the speaker array within the non-target call user area to output the time-domain masking signal.
9. An electronic device, characterized in that, The device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the steps of the in-vehicle call privacy control method according to any one of claims 1-7 are implemented.
10. A computer-readable storage medium, characterized in that, Computer program instructions are stored on the computer-readable storage medium, and when the computer program instructions are executed by the processor, the steps of the in-vehicle call privacy control method according to any one of claims 1-7 are implemented.
Citation Information
Cited By
Vehicle-mounted voice protection method, device and equipment and computer readable storage medium
CN121662010A