Voice masking method, apparatus, and system, and vehicle
The voice masking method in intelligent vehicles creates private zones by determining sound source and masking positions and using sound field control to prevent unwanted audio leakage, enhancing privacy and audio clarity.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-04-14
- Publication Date
- 2026-05-18
AI Technical Summary
Intelligent vehicle cabins face challenges in protecting privacy and preventing information leakage through voice communication, as occupants may not want their conversations to be overheard by others in the vehicle.
A voice masking method that determines sound source and target masking positions, generates masking sounds based on detection results, and uses sound field control to create zones of privacy within the vehicle cabin.
Effectively masks private conversations, ensuring occupants at target locations cannot hear the audio content from the sound source, while minimizing interference and maintaining a clear audio experience for the driver.
Smart Images

Figure 2026515480000001_ABST
Abstract
Description
Technical Field
[0001] This application is related to the field of intelligent vehicles, and particularly relates to a method, apparatus, and system for vehicle cabin voice masking, as well as a vehicle.
Background Art
[0002] With the development of vehicle intelligence, the intelligent vehicle cabin is used as an independent space, and both drivers and passengers have requirements for privacy protection regarding the voice communication content in the cabin space. For example, in a business scenario, two cooperating parties are talking in the rear row of the cabin, but do not want the conversation content to be known to the driver or passengers in the front row; or when the driver is on the phone with others, the driver does not want the conversation content to be known to other passengers in the vehicle. In scenarios such as the above, how to avoid privacy or information leakage caused by the voices of passengers has become a problem that needs to be solved.
Summary of the Invention
[0003] This application provides a voice masking method, apparatus, and system, as well as a vehicle, which satisfy the requirements of passengers for private voice communication in the vehicle cabin space and improve information security.
[0004] According to a first aspect, this application provides a voice masking method applied to a vehicle. The method includes: determining a sound source position and a target masking position; receiving a sound signal from the sound source position; detecting whether there is a voice signal in the sound signal to generate a detection result; and when the detection result indicates that there is a voice signal in the sound signal, generating a masking sound for the voice signal and outputting the masking sound to a first speaker at the sound source position and a second speaker at the target masking position; or when the detection result indicates that there is no voice signal in the sound signal, skipping the step of generating a masking sound.
[0005] Based on the aforementioned solution, if it is necessary to mask the audio signal, the masking sound can be output to a speaker at the target masking location. As a result, the occupant at the target masking location will not be able to understand or clearly hear the audio content from the sound source location, thus protecting the privacy of the occupant at the sound source location. Furthermore, in this solution, the audio signal from the sound source location is detected, and the generation of the masking sound is controlled based on the detection result. As a result, if there is no audio input at the sound source location, it is possible to prevent the masking sound from continuously interfering with the occupant at the target masking location.
[0006] In a possible implementation, first, an audio masking activation command is received, then the sound source location is determined based on the source location of the activation command, and the target masking location is determined based on occupant information acquired by sensors in the vehicle. The sound source location can be precisely determined based on the command source location, and the distribution of occupants in the vehicle can be intelligently identified based on captured occupant photographs to determine the target masking location, thereby preventing the transmission of masking sounds to irrelevant locations.
[0007] In possible implementations, the sound source location and target masking location are determined based on input from the occupants in the vehicle. The sound source location and target masking location are determined based on the dynamic input of the occupants, resulting in a better experience for them. Furthermore, this is applicable to different scenarios, such as those where occupants do not wish to be filmed by cameras in the vehicle.
[0008] In possible implementations, enhancement processing can be performed on the sound signal at the sound source location to provide a clearer sound signal for subsequent speech detection, thereby improving the accuracy of speech detection. Optionally, speech enhancement processing may include echo cancellation and / or adaptive speech noise reduction.
[0009] In possible implementations, time-domain inversion is performed on the audio signal to generate a masking sound.
[0010] In possible implementations, noise data is retrieved from a noise database, and masking sounds are generated using the noise data, or masking sounds are generated using the audio signal and noise data. Optionally, the noise data in the noise database is pre-configured. .sound The speech characteristics of the voice signal are analyzed, and noise data corresponding to those characteristics is retrieved from a noise database.
[0011] This application provides multiple masking sound generation methods, offering greater flexibility in implementation.
[0012] In possible implementations, automatic gain control adjustment may be performed on the masking sound, and as a result, the volume of the masking sound may be within a specified range. Fits inside This avoids volume fluctuations or sound leakage of the masking sound, thereby ensuring the masking effect at the target masking location.
[0013] In possible implementations, sound field control processing is performed on the masking sound based on the sound source position and the target masking position, and a first masking sound and a second masking sound are output, where the sound field control processing includes adjusting the phase and amplitude of each frequency signal in the masking sound, and then outputting the first masking sound to a first speaker at the sound source position and the second masking sound to a second speaker at the target masking position.
[0014] In this application, sound field control processing is performed on the masking sound, and as a result, different masking sounds are reproduced by the first speaker at the sound source location and the second speaker at the target masking location. In this way, the masking sound at the target masking location satisfies the masking requirements, the masking sound reproduced by the speaker at the sound source location is canceled out by the masking sound at another location, and interference caused by the masking sound from another location is avoided for the occupant at the sound source location. .sound Based on the source position and the target masking position, sound field control processing can be performed on the masking sound to output N channels of masking sound, where the N channels of masking sound include a first masking sound and a second masking sound, N is the number of speakers in the vehicle interior, and the volume of the masking sound from the first speaker at the sound source position is lower than the volume of the masking sound from the second speaker at the target masking position.
[0015] Multiple speakers in the vehicle cabin are coordinated and controlled through sound field control processing to reproduce masking sound. As a result, a sound field dark zone can be formed at the sound source location (to avoid interference of the masking sound to the sound source location), and a sound field bright zone can be formed at the target masking location (to ensure the masking effect at the target masking location). Furthermore, various requirements for private voice communication within the vehicle cabin can be met. For example, when a rear-row passenger performs voice communication, it is desirable that the voice be masked from the front-row passenger. In another example, when a rear-row passenger communicates with a passenger in the passenger seat, the voice is masked from the driver's seat. Moreover, if the target masking location includes the driver's seat, a second speaker is included, located in the driver's seat headrest. The masking sound is reproduced through the headrest speaker in the target masking seat, resulting in the masking sound having stronger directivity and a better masking effect.
[0016] In possible implementations, if the target masking location includes the driver's seat, the volume of the vehicle safety warning sound in the driver's seat is increased. This increased volume ensures that the driver can hear the safety warning sound.
[0017] Furthermore, external sounds may be acquired, and specific types of identification may be performed on these external sounds to identify specific sounds within the external noise, which are then output to speakers near the driver's seat. These specific types of sounds may include ambient alarm sounds, such as sirens. By capturing and identifying specific types of external sounds, and playing specific types of sounds inside the vehicle, it is possible to ensure that the driver reacts to the external environment, thereby improving driving safety.
[0018] In possible implementations, a further audio masking stop command may be received, and the reception of audio signals at the sound source location is stopped based on the stop command.
[0019] According to a second aspect, the present application provides a voice masking device, the device: A positioning module configured to determine the sound source position and the target masking position; A sound detection module configured to receive an audio signal from a sound source location, detect whether an audio signal is present in the audio signal, and generate a detection result; A masking sound generation module configured to generate a masking sound based on detection results, wherein if the detection result indicates the presence of an audio signal in the sound signal, it generates a masking sound for the audio signal, or if the detection result indicates the absence of an audio signal in the sound signal, it skips generating a masking sound. ni kamo The masking sound generation module that has been implemented; and It includes a masking sound post-processing module configured to output a masking sound to a first speaker at the sound source location and a second speaker at the target masking location.
[0020] In possible implementations, the position determination module is: Receive an audio masking activation command; Determine the sound source position based on the command source position of the activation command; and Determine the target masking position based on the passenger information acquired by the sensors in the vehicle ni kamo is made.
[0021] In a possible implementation, the positioning module determines the sound source position and the target masking position based on the sound source position and the target masking position input by the passengers in the vehicle ni kamo is made.
[0022] In a possible implementation, the audio detection module is specifically configured to perform enhancement processing on the audio signal after receiving the audio signal from the sound source position. The enhancement processing includes echo cancellation processing and / or adaptive audio noise reduction processing.
[0023] In a possible implementation, the masking sound generation module generates a masking sound by performing time-domain inversion processing on the audio signal ni kamo is made.
[0024] In a possible implementation, the masking sound generation module acquires noise data in the noise database and generates a masking sound by using the noise data, or generates a masking sound by using the audio signal and the noise data ni kamo is made. Optionally, the noise data is preset .Ma The skinning sound generation module analyzes the audio characteristics of the audio signal, acquires the noise data corresponding to the audio characteristics from the noise database, and is configured to generate a masking sound by using the noise data.
[0025] In possible implementations, the masking sound post-processing module is further configured to perform automatic gain control adjustment on the masking sound before outputting the masking sound to a first speaker at the sound source location and a second speaker at the target masking location, so that the volume of the masking sound is within a specified range. Fits inside It will become like that.
[0026] In possible implementations, the masking sound post-processing module performs sound field control processing on the masking sound based on the sound source position and target masking position, and outputs a first masking sound and a second masking sound. ni kamo The sound field control process includes adjusting the phase and amplitude of each frequency signal in the masking sound, outputting the first masking sound to the first speaker at the sound source location, and outputting the second masking sound to the second speaker at the target masking location.
[0027] sound Performing sound field control processing on a masking sound based on the source position and target masking position to output a first masking sound and a second masking sound includes: performing sound field control processing on a masking sound based on the sound source position and target masking position to output N channels of masking sound, where the N channels of masking sound include a first masking sound and a second masking sound, N is the number of speakers in the vehicle cabin, and the volume of the masking sound received by the first speaker is lower than the volume of the masking sound received by the second speaker. Furthermore, if the target masking position includes the driver's seat, the second speaker includes a speaker located in the headrest of the driver's seat.
[0028] In possible implementations, if the target masking location includes the driver's seat, the masking sound post-processing module is further configured to increase the volume of the vehicle safety warning sound. Furthermore, the masking sound post-processing module is further configured to acquire external sounds, perform specific type identification on the external sounds to identify specific sounds in the external sounds, and output the specific sounds to the speaker in the driver's seat.
[0029] In possible implementations, the positioning module is further configured to receive a voice masking stop command, and the voice detection module is further configured to stop receiving sound signals from the sound source location based on the stop command.
[0030] According to a third aspect, the present application provides a voice masking device. The voice masking device includes a processor and a memory. The memory is configured to store instructions, and the processor is configured to execute the instructions stored in the memory to carry out a method according to the first aspect or any one of the possible implementations of the first aspect.
[0031] According to a fourth aspect, the present application provides a voice masking system for use in a vehicle. The voice masking system is: A voice masking activation and / or deactivation device configured to control the activation and / or deactivation of a voice masking function; A microphone positioned at the sound source location and configured to capture voices at the sound source location; A first speaker positioned at the sound source location; and A second speaker positioned at the target masking location; The system includes a first speaker and a second speaker configured to reproduce masking sounds.
[0032] In possible implementations, the voice masking system further includes a camera, which is configured to capture images of the occupants inside the vehicle.
[0033] In possible implementations, the voice masking system further includes an information input device, which is used by occupants in the vehicle to input the sound source location and the target masking location.
[0034] In possible implementations, the microphone further includes a microphone located outside the vehicle, which is configured to capture external sounds.
[0035] According to a fifth aspect, the present application provides a computer-readable storage medium. The computer-readable storage medium is configured to store a computer program. When the computer program is executed on a computer, the computer is able to perform any one of the methods of the first aspect or an implementation thereof.
[0036] According to the sixth aspect, the present application provides a vehicle, which includes any one of the devices according to the second aspect, the devices according to the third aspect, and the system according to the fourth aspect.
[0037] According to the seventh aspect, the present application provides a chip including a processor configured to read instructions for performing a method according to the first aspect or any one of possible implementations of the first aspect.
[0038] According to the eighth aspect, the present application provides a computer program product including computer program code. When the computer program code is executed on a computer, the computer becomes capable of performing any one of the methods of the first aspect or an implementation of the first aspect.
[0039] For the technical effects brought about by the second through eighth aspects or possible implementations, please refer to the description of the technical effects brought about by the first aspect or the corresponding implementation. [Brief explanation of the drawing]
[0040] To more clearly explain the technical solution in this application, the accompanying drawings of the embodiments in this application will be briefly described below.
[0041] [Figure 1] Figure 1 is a diagram showing the seating distribution inside a vehicle according to an embodiment of the present application.
[0042] [Figure 2] Figure 2 is a diagram of a vehicle voice masking system according to an embodiment of the present application.
[0043] [Figure 3] Figure 3 shows the arrangement of the microphone and speaker in the vehicle according to the embodiment of the present application.
[0044] [Figure 4] Figure 4 is a schematic flowchart of the voice masking method according to the embodiment of the present application.
[0045] [Figure 5] Figure 5 shows the input control interface for the sound source position and / or target masking position according to an embodiment of the present application.
[0046] [Figure 6] Figure 6 is a diagram illustrating a method for performing enhancement processing on an audio signal according to an embodiment of the present application.
[0047] [Figure 7A] Figure 7A is a diagram of a masking sound generation method according to an embodiment of the present application.
[0048] [Figure 7B] Figure 7B is a diagram of another masking sound generation method according to an embodiment of the present application.
[0049] [Figure 7C] Figure 7C is a diagram of yet another masking sound generation method according to an embodiment of the present application.
[0050] [Figure 8] Figure 8 is a diagram illustrating the sound masking principle according to an embodiment of the present application.
[0051] [Figure 9] Figure 9 is a diagram illustrating the sound wave cancellation principle according to an embodiment of the present application.
[0052] [Figure 10] Figure 10 is a diagram illustrating sound field control according to an embodiment of the present application.
[0053] [Figure 11] Figure 11 is a diagram of a voice masking device according to an embodiment of the present application.
[0054] [Figure 12] Figure 12 is a diagram of another voice masking device according to an embodiment of the present application. [Modes for carrying out the invention]
[0055] The technical solutions in the embodiments of this application will be described in detail below with reference to the attached drawings. The same reference numerals in the attached drawings indicate elements having the same or similar function. Various aspects of the embodiments are shown in the attached drawings, which are not necessarily drawn to scale unless otherwise specified.
[0056] References to “embodiments,” “some embodiments,” or similar terms described in this specification indicate that one or more embodiments of this application include certain features, structures, or characteristics described by reference to the embodiments. Therefore, phrases such as “in some embodiments,” “in some other embodiments,” and “in other embodiments,” appearing in different parts of this specification, do not necessarily refer to the same embodiment. Rather, they refer to different embodiments. strong Unless otherwise specified, it means "one or more of the embodiments, but not all of them." The terms "includes," "equip," "have," and their variations are all different in nature. strong Unless otherwise specified, it means "includes, but not limited to."
[0057] In this application, terms such as "First" and "Second" are used to distinguish between the same or similar items that essentially have the same role and function. It should be understood that there is no logical or timing dependency between "First," "Second," and "nth," and that there are no limitations on the quantity or execution sequence.
[0058] In this application, "at least one" means one or more, and "multiple" means two or more. "And / or" describes the relationship between related subjects and indicates that there may be three possible relationships. For example, A and / or B may indicate that: only A exists, both A and B exist, and only B exists, where A and B may be singular or plural. The character " / " generally indicates an "or" relationship between related subjects. At least one of the following items(parts) or similar expressions means any combination of these items, including any combination of singular or plural items(parts). For example, at least one item(part) of a, b, or c may indicate a, b, c, a and b, a and c, b and c, or a, b, and c, where a, b, and c may be singular or plural.
[0059] Furthermore, in order to better illustrate this application, numerous specific details are given in the following particular implementations. Those skilled in the art should understand that this application can be implemented without some of these specific details.
[0060] The method provided in the embodiments of this application is applicable to scenarios where it is necessary to mask the voice when an occupant in the vehicle is performing voice communication or making a phone call, and in particular is applicable to scenarios where it is necessary to mask the voice of a rear-row occupant from the driver in the front row when the occupant is having a conversation.
[0061] Figure 1 is a diagram showing the seating distribution inside a vehicle according to an embodiment of the present application. As shown in Figure 1, the vehicle 100 has front and rear seats, which include the driver's seat 101, passenger seat 102, rear left seat 103, rear middle seat 104, and rear right seat 105. For example, if the occupant in the rear left seat 103 is having a conversation with the occupant in the rear right seat 105, and the content of the conversation relates to privacy or business matters, the rear occupants do not want the occupants in the front seats, the driver's seat 101 and passenger seat 102, to overhear the conversation. Another example is when the occupant in the driver's seat 101 (i.e., the driver) is making a phone call, the content of the call may be overheard by the occupant in the passenger seat 102 and the rear occupants. If the content of the call relates to personal privacy, this is not something the occupant in the driver's seat 101 would want.
[0062] In relation to the aforementioned problem, this application provides a solution. First, the sound source location and the target masking location are determined. Next, an audio signal is received at the sound source location, and it is detected whether an audio signal is present in the audio signal. Based on the detection result, a masking sound is generated. Furthermore, the masking sound is played back through a first speaker at the sound source location and a second speaker at the target masking location. As a result, the occupant at the target masking location cannot understand or clearly hear what the occupant at the sound source location is saying.
[0063] The solution provided in the embodiments of this application can be implemented through the voice masking system 200 shown in Figure 2. As shown in Figure 2, the voice masking system 200 includes, but is not limited to, a control device 210, a microphone device 220, and a speaker device 230. The control device 210 can be connected to the microphone device 220 and the speaker device 230. The control device 210 may be a processing-capable hardware platform, a processing-capable software platform, or a platform integrating processing-capable hardware and software, for example, an in-vehicle computing platform and / or a cabin domain control platform. This is not limited to this application. The speaker device 230 may include one or more speakers positioned at different locations.
[0064] The voice masking system 200 may further include an interaction device 240, which may include a sensor device 241. The sensor device 241 may include, but is not limited to, some or all of an image sensor 241-1, a radar 241-2, and a seat sensor 241-3. The image sensor 241-1 may be a camera, and the image sensor 241-1 may be configured to capture images of the interior of a vehicle. The number and arrangement of the image sensors 241-1 are as specified in this application. is limited Not specified. Radar 241-2 may include one or more ultrasonic radars, millimeter-wave radars, and similar devices, and is configured to detect occupants in the vehicle cabin. Seat sensors 241-3 are configured to detect whether an occupant is in a seat in the vehicle, and seat sensors 241-3 may be gravity sensors. The control device 210 can acquire relevant information about the occupant's position in the cabin via the sensor device 241.
[0065] The interaction device 240 may further include a display device 242. The display device 242 includes a mobile terminal or an in-cabin display and is configured to interact with the occupant and receive input from the occupant. The interaction device 240 may also include a function activation / deactivation device 243 in the cabin, which may include physical buttons, switches, and similar devices.
[0066] With the development of intelligent vehicles and the increasing demands of people for interaction and audio quality in vehicle cabins, the number of microphones and speakers in vehicles is increasing. Speaker placement is also becoming more precise to achieve a superior sound field and acoustic experience. Microphone placement is also being adjusted to be as close as possible to the occupants in the cabin, facilitating the capture of their sound signals.
[0067] Figure 3 is a diagram showing the placement of microphones and speakers in a vehicle according to an embodiment of the present application. As shown in Figure 3, the vehicle's speaker system 230 may include speakers 230-1 to 230-8. The speakers may be arranged in a surround sound manner around the passenger compartment, or they may be placed on the headrests of the seats inside the passenger compartment, for example, speaker 230-4, in which case the speaker is placed on the headrest of the driver's seat. The placement and directivity of the speakers are appropriately set so that a sound field with excellent acoustic effect can be formed inside the passenger compartment. The vehicle's microphone system 220 may include in-cabin microphones 220-1 to 220-4. The in-cabin microphones are placed above or to the side of the occupants' seats and are as close as possible to the occupants' heads. The vehicle's microphone system may further include an external microphone 220-5, which is configured primarily to capture sounds from outside the vehicle. It should be understood that the placement and number of speakers and microphones are merely examples and do not represent all possible arrangements. The placement and number of speakers and microphones are as specified in this application. is limited It is not determined. However, it should be understood that all configurations that meet the requirements of this application fall within the scope of protection of this application.
[0068] For ease of understanding and explanation, it should be understood that the following describes the method provided in the embodiments of this application using a control device within a vehicle as the implementing body. For example, the control device may be the control device 210 in Figure 2. The control device may be a component within a vehicle, such as a chip, chip system, or other functional module capable of calling and executing a program. However, it should be understood that this should not constitute any limitation on the implementing body of the method provided in this application. In this application, the control device 210 may also be referred to as a voice masking device.
[0069] Figure 4 is a schematic flowchart of the voice masking method according to an embodiment of the present application. As shown in Figure 4, the method may include steps S410 to S430.
[0070] S410: Determine the sound source location and the target masking location.
[0071] Optionally, S410 may be implemented in any one of the following ways. The specific method to be used may depend on the implementation of control devices and equipment within the vehicle cabin.
[0072] In an embodiment, the method further includes receiving a voice masking activation command prior to method S410. Determining the sound source location and target masking location includes: determining the sound source location based on the command source location of the activation command; and obtaining occupant information acquired by a sensor device in the vehicle and determining the target masking location based on the occupant information. The sensor device may include sensor device 241. The activation command may be triggered by a function activation / deactivation device located at the sound source location in the vehicle cabin. The function activation / deactivation device may be function activation / deactivation device 243. The position of the function activation / deactivation device may be determined based on the activation command to determine the command source location of the command.
[0073] For example, an image sensing device (e.g., a camera) inside the vehicle may be used to acquire images of the occupants inside the vehicle, and the distribution of occupant positions inside the vehicle may be determined from the occupant images to determine the target masking position.
[0074] For example, the occupants inside the vehicle may be detected via a radar inside the vehicle (e.g., millimeter-wave radar), and the distribution of the occupants' positions inside the vehicle may be determined to determine the target masking position.
[0075] For example, whether an occupant is in a seat in the vehicle may be detected via pressure sensors in the seats inside the vehicle, the distribution of occupant positions inside the vehicle may be determined, and the target masking position may be determined accordingly.
[0076] For example, a combination of multiple sensors, such as a pressure sensor and an image sensor, may be used to more accurately determine the distribution of occupant positions within the vehicle cabin, thereby more accurately determining the target masking position.
[0077] In the embodiment, the sound source location and / or target masking location is the sound source location and / or target masking location input by an occupant in the vehicle. For example, the occupant may input the sound source location and / or target masking location via a display device. The display device may include a display device 242. The display device may be located in a position such as the center console, behind the seat headrest, or in the rear center console.
[0078] Figure 5 shows an input control interface 500 for sound source position and / or target masking position according to an embodiment of the present application. In the input control interface shown in Figure 5, the microphone icon indicates the sound source position, and the mute icon indicates the target masking position. As shown in Figure 5, the position of seat 501 is the sound source position, and seats 502, 503, 504, and 505 are target masking positions. A crew member can tap an icon to switch the corresponding seat position to either the sound source position or the target masking position. For example, a crew member can tap the mute icon for seat 502 in the input control interface to switch the position of seat 502 to the sound source position. In this case, both seat 501 and seat 502 are sound source locations, the occupant at seat 502 can understand the content of the voice from the occupant at seat 501, and the voices of the occupants at seats 501 and 502 are not clearly audible or understandable to the occupants at seats 503, 504, and 505.
[0079] An input control interface is provided, which allows occupants to dynamically control and adjust the locations where voice masking needs to be performed, thereby improving the occupant experience by more flexibly adapting to various scenarios within the vehicle cabin. For example, if a rear-row occupant is making a phone call or having a conversation, the content of the conversation needs to be masked from the front row. In another example, if a passenger in the front passenger seat needs to join a conversation between rear-row occupants, the input interface can be flexibly adjusted. It should be understood that the input control interface provided in this embodiment is merely an example of an input control interface, and the sound source location and target masking location may be indicated by using other icons instead. The input control interface may also further indicate the current occupant location within the vehicle cabin, which can be identified via various sensing devices within the cabin. Alternatively, the input control interface may be a single / multiple selection switch for the sound source location and / or target masking location, a voice receiver, or the like. The sound source position and / or target masking position, input control interface, and one / more adjustment methods for the input control method are as defined in the embodiments of this application. is limited It is not determined.
[0080] In some embodiments, the sound source location may be determined based on the command source location of the voice masking activation command, and the target masking location may be determined based on input from an occupant in the vehicle. For example, the voice masking activation / deactivation device may be located above or to the side of an occupant's seat. When an occupant in the vehicle triggers the activation and / or deactivation of voice masking, the control device determines the physical location of the voice masking activation and / or deactivation device based on the received voice masking activation command. In this case, the seat location corresponding to the physical location of the voice masking activation and / or deactivation device is used as the sound source location.
[0081] The aforementioned methods for determining the sound source position and the target masking position may be combined. For example, the sound source position can be determined by combining an image sensor and a command source position. It should be understood that the combined position determination method should also fall within the scope of protection of this application.
[0082] S420: Receives an audio signal from the sound source location, detects whether an audio signal is present in the audio signal, and generates a detection result.
[0083] In one embodiment, after receiving an audio signal from the sound source location in step S420, the method further includes performing enhancement processing on the audio signal.
[0084] Specifically, step S420 This is a diagram. This includes the following steps as shown in 6.
[0085] 601: Receives an audio signal from the sound source location.
[0086] 602: Perform enhancement processing on the audio signal to generate an enhanced audio signal.
[0087] 603: Detect whether an audio signal is present within the enhanced audio signal and generate a detection result.
[0088] In some embodiments, the enhancement process may include echo cancellation and / or adaptive speech noise reduction.
[0089] Enhancement processing is performed on the audio signal, which reduces noise in the signal and makes it clearer. This helps improve the accuracy of detection results indicating whether an audio signal is present in the sound signal.
[0090] In this embodiment, the sound signal may be detected using a voice activity detection (VAD) method, and a detection result indicating whether or not a voice signal is present in the sound signal may be generated.
[0091] S430: If the detection result indicates that an audio signal is present in the sound signal, a masking sound is generated for the audio signal and output to the first speaker at the sound source location and the second speaker at the target masking location; or, if the detection result indicates that no audio signal is present in the sound signal, the generation of the masking sound is skipped.
[0092] Based on the detection result indicating whether an audio signal is present, the generation of a masking sound is controlled, and as a result, if no audio signal is present, the masking sound will not interfere with the target masking location. For example, if the occupant in the driver's seat is on a phone call, the peer end on the call may be speaking for a long time, and the occupant in the driver's seat is listening. In this case, no audio signal is detected. Ino Therefore, it is not necessary to generate a masking sound. In this way, interference from the masking sound to other occupants in the vehicle can be avoided.
[0093] In this embodiment, an audio signal buffer mechanism is set up to generate a masking sound using the buffered audio signal during a gap (e.g., 2s) in conversation where the occupant at the sound source location is paused, thereby ensuring the continuity of the masking effect. For example, the buffering may be set to include audio signal content with a duration of 500 ms. The buffering setting method and buffer size are described in this application. is limited It should be understood that this is not fixed. Furthermore, the content of the buffered audio signal may be updated depending on whether an audio signal is present. If an audio signal is detected, the buffered audio signal content is updated based on the audio signal; if no audio signal is detected, the buffered audio signal content is not updated.
[0094] Similarly, the generated masking sound may be further buffered, and the masking sound buffer mechanism is used to ensure the continuity of the masking effect in gaps where occupants at the sound source location pause during conversation. This mechanism is consistent with the buffering of the voice signal described above.
[0095] In this embodiment, as shown in Figure 7A, the masking sound is generated by performing a time-domain inversion process on the audio signal.
[0096] In one embodiment, as shown in Figure 7B, the masking sound is generated by using noise data obtained from a noise database. The noise data in the noise database may include one or more of the following: white noise, narrowband noise, speech noise, and the like. The noise data may be pre-configured in the system or downloaded from the cloud and updated in real time. The noise data in the noise database and the data sources are described in this application. is limited It is not determined.
[0097] In this embodiment, noise data is acquired from a noise database, and a masking sound is generated using the noise data and the audio signal. For example, a time-domain inversion process may be performed on the audio signal to obtain a processed audio signal, and the processed audio signal and the noise data acquired from the noise database may be fused to generate the masking sound.
[0098] In one embodiment, as shown in Figure 7C, the audio features of the audio signal are analyzed, noise data corresponding to the audio features is retrieved from a noise database, and a masking sound is generated using the noise data. The noise data in the noise database is matched based on the audio features, and as a result, the retrieved noise data can better match the audio signal, making the masking sound more comfortable for the occupant at the target masking position, and improving the masking effect. Alternatively, the masking sound may be generated through a neural network model by analyzing the audio features of the audio signal.
[0099] It should be understood that the aforementioned embodiments for generating masking sounds may be combined with each other to generate masking sounds. Other methods for generating masking sounds also exist. This is described in the present application. is limited It is not determined.
[0100] In some embodiments, automatic gain control (AGC) adjustment may be further performed on the generated masking sound, and as a result, the volume of the masking sound may be within a specified range. Fits inside This allows for the control of the masking sound volume, maintaining it within a specified range, thereby minimizing the masking sound volume and avoiding discomfort for occupants at the target masking location. It also stabilizes the masking sound volume, avoiding fluctuations or sound leakage, and thereby ensuring the masking effect.
[0101] In this embodiment, sound field control processing is performed on the masking sound based on the sound source position and the target masking position to output a first masking sound and a second masking sound, wherein the sound field control processing includes adjusting the phase and amplitude of each frequency signal in the masking sound, outputting the first masking sound to a first speaker at the sound source position, and outputting the second masking sound to a second speaker at the target masking position.
[0102] The speakers used to reproduce the masking sound are determined based on the location information of the sound source and the target masking location. The first speaker at the sound source location and the second speaker at the target masking location are coordinated and controlled through sound field control. As a result, the first and second speakers reproduce different masking sounds, creating a bright zone at the target masking location and a dark zone at the sound source location. In this way, the occupant at the target masking location cannot understand or clearly hear the audio content from the occupant at the sound source location, thus avoiding interference caused to the occupant at the sound source location by the masking sound reproduced at the target masking location.
[0103] The following briefly explains sound field control from the perspective of sound masking and noise reduction principles.
[0104] The phenomenon in which the perception of a weak sound (masked sound) is affected by another strong sound (masking sound) is called the masking effect of the human ear. Generally, the closer the frequencies of two sounds are to each other, the greater the amount of masking they exhibit. Also, high-frequency sounds are easily masked by low-frequency sounds, and low-frequency sounds are not easily masked by high-frequency sounds. For example, in a concert setting, the sound pressure level of the bass drum may not be high, but people can still clearly hear the bass drum in the concert music, while the sound of a violin is easily masked by another low-frequency instrument.
[0105] Based on the aforementioned principle, in this embodiment of the present application, a speaker near the target masking location plays a masking sound, and the sound from the sound source location is masked at the position of the occupant's human ear at the target masking location, so that the occupant at the target masking location cannot understand or clearly hear the sound content at the sound source location.
[0106] As shown in Figure 8, there is a certain distance between the crew member at the sound source and the crew member at the target masking position. Therefore, the voice of the crew member at the sound source position propagates to the ear at the target masking position via an airborne propagation path (direct sound waves). A microphone near the crew member at the sound source position captures the sound signal (direct sound waves) from the crew member at the sound source position, detects the voice signal from the sound signal, and then generates a masking sound via a masking sound generator. This masking sound is output to a speaker near the target masking position, and finally, the speaker near the target masking position plays the masking sound. In this case, at the ear of the crew member at the target masking position, the masking sound interferes with the voice of the crew member at the sound source position, achieving the purpose of voice masking.
[0107] When a masking sound is played at the target masking location, the masking sound at the target masking location propagates to the sound source location, potentially causing interference to occupants at the sound source location. Therefore, in this application, noise reduction processing may be further performed at the sound source location. For example, an active noise cancellation solution may be used. The principle of the active noise cancellation solution is as follows: Each sound contains a specific spectrum, and active noise is found whose spectrum is the same as the spectrum of the noise to be canceled and whose phase is exactly opposite (180-degree difference) to the phase of the noise to be canceled, thereby canceling out the noise to be canceled.
[0108] Figure 9 is a diagram illustrating the sound wave cancellation principle according to an embodiment of the present application. As shown in Figure 9, the first sound wave 901 and the second sound wave 902 have the same spectrum, but their phases are exactly opposite (the difference is 180 degrees). The point where the first sound wave 901 and the second sound wave 902 intersect is controlled by precise calculation. For example, the point where the first sound wave 901 and the second sound wave 902 intersect is the human ear, and the sound wave obtained after the two sound waves intersect, superimpose, and cancel each other out is the third sound wave 903, and the amplitude of the third sound wave 903 is very small, so that the human ear can hardly hear any noise at the intersection.
[0109] Based on the aforementioned principle, the objective of creating an audio dark zone at the sound source location and an audio bright zone at the target masking location can be achieved through sound field control.
[0110] Figure 10 is a diagram of sound field control according to an embodiment of the present application. As shown in Figure 10, the masking sound generation device generates a masking sound, outputs the masking sound to a first filter and a second filter, controls the parameters of the first filter and the second filter to generate a first masking sound and a second masking sound, then outputs the first masking sound to a first speaker and the second masking sound to a second speaker. The parameters of the first filter and the second filter are controlled to adjust the phase and amplitude of the respective frequency signals in the masking sound, so that the first masking sound and the second masking sound have different phases and amplitudes. The first speaker may be a speaker near the sound source, for example, a speaker on the side of the sound source. The second speaker may be a speaker near the target masking position, for example, a speaker on the side of the target masking position.
[0111] In this embodiment, the volume of the first masking sound received by the first speaker is lower than the volume of the second masking sound received by the second speaker.
[0112] In this embodiment, the masking sound is output to a third filter to generate a third masking sound, which is then output to a third speaker at the sound source location. The third speaker and the first speaker at the sound source location operate simultaneously to improve the noise reduction effect at the sound source location.
[0113] In this embodiment, the masking sound is output to a fourth filter to generate a fourth masking sound, which is then output to a fourth speaker at the target masking position. The fourth speaker and the second speaker at the target masking position operate simultaneously to improve the sound masking effect.
[0114] In this embodiment, the masking sound may be further output to N filters (the first to the Nth filter) to generate N channels of masking sound (the first to the Nth masking sound), where N is the number of speakers in the vehicle cabin. The N masking sounds are output to N speakers (the first to the Nth speakers), and the overall sound field control of the cabin is achieved by using all the speakers in the vehicle cabin. As a result, there is no sound leakage in voice masking at the target masking position, interference is low, and noise at the sound source position is reduced.
[0115] In the embodiment, multiple positions / single positions of the human ear at the sound source location and / or target masking location may be further identified, and the directivity of one or more speakers in the vehicle cabin may be dynamically adjusted for better noise reduction. reduction Effective and better sound masking effects can be achieved. The identification method may use a millimeter-wave radar, camera / sensor or similar device inside the vehicle. The specific identification method is described in this application. is limited It is not determined.
[0116] In one embodiment, if the target masking position is the driver's seat and a headrest speaker is provided in the driver's seat, a better sound masking effect can be achieved by playing the masking sound through the headrest speaker, because the headrest speaker is closer to the human ear and has stronger directivity.
[0117] Similarly, if headrest speakers are provided in the driver's seat, passenger seat, and rear seats, the masking sound can be reproduced through the headrest speakers in the driver's seat, passenger seat, and rear seats to achieve a better sound masking effect.
[0118] In the embodiment, the target masking position can be dynamically adjusted based on changes in the occupants inside the vehicle (which may include changes in occupant positions, changes in the number of occupants, and similar events), and sound field control can be performed adaptively. For example, when a new occupant boards the vehicle, sound field control is dynamically performed after the new occupant's position is detected, and sound masking is performed at the new occupant's position. In the embodiment, when an occupant disembarks from the vehicle and a change in the occupant's position is detected, the occupant at the sound source position is prompted, such as by a display or voice, to adjust the sound source position or the target masking position.
[0119] In some embodiments, sound field control may be performed by using a variable span trade-off (VAST) algorithm.
[0120] In embodiments, if the target masking location includes the driver's seat, the volume of the vehicle alert sound may be increased. The vehicle alert sound may include vehicle safety warning sounds, which may include the vehicle's battery level alarm, fuel level alarm, and similar sounds. The vehicle alert sound may further include obstacle alarm sounds and safety warning sounds, such as obstacle collision avoidance alarms and tire pressure alarms. The vehicle may further include function alert sounds, such as navigation alert sounds. The volume of the vehicle alert sound is increased, and as a result, driving safety can be ensured, while ensuring voice masking, as well as ensuring that the driver can identify the alarm in a timely manner.
[0121] In embodiments, if the target masking position includes the driver's seat, external sound signals may also be acquired, and specific types of identification may be performed on the external sound signals to identify specific sounds in the external sound, and these specific sounds may be output to the speaker in the driver's seat. External sound may also be captured via an external microphone, and specific types of identification may be performed on the external sound to identify specific sounds, which are then played back through the speaker in the driver's seat. For example, if a vehicle is moving and a vehicle behind is sounding its horn as it overtakes, the horn of the vehicle behind can be identified by identifying the external sound, and then the horn can be played back through the speaker near the driver's seat, alerting the driver and enabling them to take appropriate safe driving action. Specific types of sounds may include horns, sirens, sound signals of a specific direction, emergency sound signals, and similar types.
[0122] Only a few vehicle alert sounds, a few vehicle alarm sounds, and a few specific types of sounds are shown here, and it should be understood that more scenarios, as well as corresponding alert sounds and specific types of sounds, may be included. This is stated in the present application here. teeth Not limited.
[0123] In the embodiment, a voice masking stop command may be received, and the reception of sound signals at the sound source location is stopped based on the stop command. For example, the stop command may be triggered by an occupant at the sound source location by turning off the voice masking function activation / deactivation device.
[0124] In this embodiment, the voice masking stop command may be further automatically triggered when it is detected that an occupant at the sound source location has disembarked from the vehicle.
[0125] In some embodiments, a timeout mechanism may be further configured. For example, the timeout mechanism may be set to 2 minutes. If no voice signal is detected in the sound signal captured from the sound source location within 2 minutes, a voice masking stop command may be automatically triggered.
[0126] In one embodiment, the voice masking stop command may be further automatically triggered when it is detected that an occupant at the target masking position has disembarked from the vehicle.
[0127] The embodiments described in this specification may be independent solutions or may be combined based on internal logic. All of these solutions fall within the scope of protection of this application.
[0128] It will be understood that the methods and operations performed by the control device in the embodiments of the above-described method may, alternatively, be performed by a component (e.g., a chip or circuit) that can be used for the control device.
[0129] The voice masking device provided in the embodiments of this application will be described in detail below with reference to Figures 11 and 12. It should be understood that the description of the device embodiments corresponds to the description of the method embodiments. Therefore, for details not described in detail, please refer to the method embodiments described above.
[0130] As shown in Figure 11, embodiments of the present application further provide a voice masking device 1100 configured to perform the functions of the control device in the method described above, the device being used in the flowcharts shown in Figures 4, 6, 7A, 7B, 7C, and 10, and capable of performing the functions of the embodiments of the method described above. For example, the device may be a software module or a chip system. In this embodiment of the present application, the chip system may include a chip or include a chip and other components. The voice masking device 1100 may be the control device 210 shown in Figure 2.
[0131] In one embodiment, the voice masking device 1100 is used inside a vehicle, and the voice masking device 1100 is: A position determination module 1101 configured to determine the sound source position and the target masking position; A voice detection module 1102 is configured to receive an audio signal from a sound source location, detect whether an audio signal is present in the audio signal, and generate a detection result; and A masking sound generation module 1103 is configured to generate a masking sound based on detection results, and if the detection result indicates that an audio signal is present in the sound signal, it generates a masking sound for the audio signal, or if the detection result indicates that no audio signal is present in the sound signal, it skips generating a masking sound. ni kamo The masking sound generation module 1103 is implemented; and The system includes a masking sound post-processing module 1104 configured to output a masking sound to a first speaker at the sound source location and a second speaker at the target masking location.
[0132] In an embodiment, the position determination module 1101 is: Received a command to activate the voice masking function; The sound source position is determined based on the command source position of the activation command; and The system is specifically configured to determine the target masking position based on occupant information acquired by sensors within the vehicle.
[0133] In the embodiments, the sensors include one or more of an image sensor, a radar sensor, and a seat sensor, where the image sensor may include a camera, the radar sensor may include one or more of an ultrasonic radar, a millimeter-wave radar, and the like, and the seat sensor may include one or more of a gravity sensor, a pressure sensor, and the like.
[0134] In one embodiment, the position determination module 1101 is configured to receive information about occupants in the vehicle and to determine the sound source position and / or target masking position based on occupant input, the occupant input including the sound source position and / or target masking position entered by the occupant.
[0135] In this embodiment, the sound detection module 1102 is further configured to perform enhancement processing on the sound signal after receiving the sound signal from the sound source location.
[0136] In the embodiment, the enhancement process includes echo cancellation and / or adaptive speech noise reduction.
[0137] In this embodiment, the masking sound generation module 1103 generates a masking sound by performing a time-domain inversion process on the audio signal. ni kamo It has been done.
[0138] In one embodiment, the masking sound generation module 1103 is: Obtain noise data from the noise database; and To generate a masking sound by using noise data, or to generate a masking sound by using an audio signal and noise data. ni kamo This has been done. Noise data in the database may be pre-configured.
[0139] In one embodiment, the masking sound generation module 1103 is configured to acquire noise data from a noise database, and the masking sound generation module 1103: The system analyzes the audio features of an audio signal and retrieves noise data corresponding to those features from a noise database. ni kamo It has been done.
[0140] In one embodiment, the masking sound post-processing module 1104 is further configured to perform automatic gain control adjustment on the masking sound before outputting the masking sound to the first speaker at the sound source location and the second speaker at the target masking location, so that the volume of the masking sound is within a specified range. Fits inside It will become like that.
[0141] In one embodiment, the masking sound post-processing module 1104 is: A step of performing sound field control processing on a masking sound based on the sound source position and the target masking position, and outputting a first masking sound and a second masking sound, wherein the sound field control processing includes adjusting the phase and amplitude of each frequency signal in the masking sound; and The process involves outputting a first masking sound to a first speaker at the sound source location, and outputting a second masking sound to a second speaker at the target masking location. ni kamo It has been done.
[0142] In one embodiment, sound field control processing is performed on the masking sound based on the sound source position and the target masking position to output a first masking sound and a second masking sound: The process includes performing sound field control processing on the masking sound based on the sound source position and the target masking position, and outputting N channels of masking sound, where the N channels of masking sound include a first masking sound and a second masking sound, and N is the number of speakers in the vehicle cabin.
[0143] In this embodiment, the masking sound received by the first speaker at the sound source location is lower in pitch than the masking sound received by the second speaker at the masking location.
[0144] In one embodiment, when the target masking position is the driver's seat, the second speaker includes a speaker located on the headrest of the driver's seat.
[0145] In one embodiment, if the target masking location includes the driver's seat, the masking sound post-processing module 1104 is further configured to increase the volume of the vehicle safety warning sound.
[0146] In one embodiment, if the target masking location includes the driver's seat, the masking sound post-processing module 1104 further: The system is configured to acquire external sounds, perform specific type identification on those external sounds to identify specific sounds within the external soundscape, and output those specific sounds to the speaker in the driver's seat.
[0147] In this embodiment, the position determination module 1101 can further receive a voice masking stop command and transmit the stop command to the voice detection module. The voice detection module 1102 is configured to stop receiving sound signals from the sound source location based on the stop command.
[0148] In this embodiment, the position determination module 1101 can further directly control the sound detection module based on a stop command, and as a result, the sound detection module stops receiving sound signals at the sound source location.
[0149] In one embodiment, the voice detection module may be configured to directly receive a voice masking stop command and, based on the stop command, to stop receiving the sound signal at the sound source location.
[0150] In some embodiments provided in this application, it should be understood that the disclosed apparatus and methods may be carried out in other ways. For example, the embodiments of the apparatus described are merely examples. For example, the division into units, components, or modules is merely a logical functional division, and other division methods may be used in actual implementations. For example, multiple units, components, or modules may be combined or integrated into another system, or some features may be ignored or not performed. In addition, the mutual coupling, direct coupling, or communication connection shown or described may be carried out through some interface. Indirect coupling or communication connection between apparatus or units may be carried out in electrical, mechanical, or other forms.
[0151] Units described as separate parts may or may not be physically separate, and parts shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected based on the actual requirements in order to achieve the objectives of the solution of the embodiment.
[0152] Furthermore, the functional units in the embodiments of this application may be integrated into a single processing unit, each unit may exist physically independently, or two or more units may be integrated into a single unit.
[0153] Embodiments of the present application further provide another voice masking device 1200. The voice masking device 1200 shown in Figure 12 may be a hardware circuit implementation of the device shown in Figure 11 and is used in the flowcharts shown in Figures 4, 6, 7A, 7B, 7C, and 10 to perform the functions of the embodiments of the method described above.
[0154] As shown in Figure 12, the voice masking device 1200 includes at least one processor 1201. The at least one processor 1201 may be a central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute a relevant program to carry out the voice masking method in the embodiment of the method of the present application.
[0155] The processor 1201 may alternatively be an integrated circuit chip having signal processing capabilities. In the implementation process, the steps of the voice masking method in this application may be carried out through hardware logic circuits in the processor 1201 or through instructions in the form of software.
[0156] The voice masking device 1200 may further include at least one memory 1202, which is configured to store instructions and / or data. The processor 1201 may be configured to execute instructions stored in the memory 1202. The memory 1202 may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (erasable PROM, EPROM), electrically erasable programmable read-only memory (electrically EPROM, EEPROM), or flash memory. The volatile memory may be random access memory (RAM). As an example that is not limited to this, RAM may have many forms, such as static random access memory (static RAM, SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (synchronous DRAM, SDRAM), double data rate synchronous dynamic random access memory (double data rate SDRAM, DDR SDRAM), enhanced synchronous dynamic random access memory (enhanced SDRAM, ESDRAM), synchlink dynamic random access memory (synchlink DRAM, SLDRAM), and direct rambus dynamic random access memory (direct rambus RAM, DR RAM). It should be understood that the memory in the embodiments of this application may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. It should be noted that the memory of the systems and methods described in this specification includes, but is not limited to, these and any other suitable type of memory.
[0157] It should be noted that while only memory and a processor are shown for the voice masking device 1200, those skilled in the art will understand that in a particular implementation process, the voice masking device 1200 may further include other components required to implement normal operation, such as a power supply, communication ports, or input / output devices.
[0158] Embodiments of the present application further provide a voice masking system 200. The voice masking system 200 may include a control device 210, a microphone device 220, and a speaker device 230, as shown in Figure 2. The control device 210 may be the voice masking device 1100 or 1200 in the embodiments described above. The microphone device 220 may include a microphone positioned at the sound source location and configured to capture voice signals at the sound source location. The speaker device 230 may include a first speaker positioned at the sound source location and a second speaker positioned at the target masking location and configured to reproduce a masking sound.
[0159] In an embodiment, the voice masking system 200 may further include an interaction device 240 shown in Figure 2, the interaction device may include a function activation / deactivation device 243. The function activation / deactivation device 243 may include a voice masking function activation / deactivation device, which is configured to control the activation and / or deactivation of the voice masking function.
[0160] In an embodiment, the interaction device 240 may further include a sensor device 241 shown in Figure 2. The sensor device 241 may include one or more of an image sensor, a radar sensor, and a seat sensor, and is configured to collect information about occupants in the vehicle. For example, the image sensor includes a camera, and based on images of occupants in the vehicle captured by the camera, images are transmitted to the control device 210. The control device 210 obtains the sound source location and / or target masking location through image-based identification.
[0161] In this embodiment, the interaction device 240 provides information input The device further includes an information input device configured to allow occupants in the vehicle to input the sound source location and / or target masking location. The information input device may be the display device 242 shown in Figure 2, or a physical button-type input device, or a terminal device. The terminal device may be a mobile phone, tablet, wristwatch, or the like. The terminal device may communicate directly or indirectly with the control device 210 via Wi-Fi, Bluetooth®, a mobile network, or the like. The terminal device is described in this application. is limited The interaction method between the terminal device and the control device 210 is not defined. Limited It should be understood that it is not fixed.
[0162] In one embodiment, the microphone device 220 may further include one or more microphones located outside the vehicle, the microphones being configured to capture sound signals from outside the vehicle.
[0163] Embodiments of the present application further provide a computer-readable storage medium configured to store a computer program, wherein when the computer program is executed on the computer, the computer is able to perform a voice masking method.
[0164] Embodiments of the present application further provide a computer program product including computer program code, which, when executed on a computer, enables the computer to perform a voice masking method.
[0165] Embodiments of the present application further provide a vehicle comprising either a voice masking device 1100 or a voice masking device 1200, or a voice masking system 200.
[0166] Those skilled in the art will recognize, by referring to the examples described in the embodiments disclosed in this specification, that the units and algorithmic steps may be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether the function is performed by hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art may use various methods to implement the described function for each specific application, but such implementation should not be considered to extend beyond the scope of this application.
[0167] In this document, the specific term "example" means "used as an example, embodiment, or illustration." Any embodiment described as an "example" is not necessarily described as being superior or better than any other embodiment.
[0168] It should be understood that the process sequence number does not mean the execution sequence in the embodiments of this application. The execution sequence of a process is determined based on the function and internal logic of the process and should not be construed as any limitation to the implementation process of the embodiments of this application.
[0169] It should be understood that determining B based on A does not mean that B is determined solely on A; B may, alternatively, be determined based on A and / or other information.
[0170] As used in this specification, terms such as “component,” “module,” and “system” refer to computer-related entities, hardware, firmware, combinations of hardware and software, software, or running software. For example, a component may be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. As illustrated by the drawings, both computing devices and applications running on computing devices can be components. One or more components may reside within a process and / or an execution thread, and components may be located on one computer and / or distributed across two or more computers. Furthermore, these components may run from various computer-readable media that store various data structures. For example, components may communicate by using local and / or remote processes, and on the basis of signals, for example, having one or more data packets (e.g., data from two components, communicating via a network such as the Internet, communicating with a local system, another component in a distributed system, and / or communicating with other systems using signals).
[0171] Relevant parts between embodiments of the method of this application can be referenced from one another; an apparatus provided in an embodiment of an apparatus is configured to perform the method provided in the corresponding embodiment of the method. Thus, an embodiment of an apparatus can be understood by referring to the relevant parts in the relevant embodiment of the method.
[0172] The foregoing description is merely a specific implementation of the present application and is not intended to limit the scope of protection of the present application. Any modification or substitution that is readily conceived by a person skilled in the art within the scope of the technical scope disclosed in the present application shall fall within the scope of protection of the present application. Accordingly, the scope of protection of the present application shall be subject to the scope of protection of the claims.
Claims
1. A method for masking sounds applied to vehicles: Steps to determine the sound source position and the target masking position; A step of receiving an audio signal from the aforementioned sound source location; A step of detecting whether an audio signal is present in the sound signal and generating a detection result; and If the detection result indicates that an audio signal is present in the sound signal, the process generates a masking sound for the audio signal and outputs the masking sound to a first speaker at the sound source location and a second speaker at the target masking location; or, if the detection result indicates that no audio signal is present in the sound signal, the process of generating a masking sound is skipped. A method that includes this.
2. In the method according to claim 1, Prior to the step of determining the sound source position and the target masking position, the method further: The step includes receiving a voice masking activation command; The steps for determining the sound source position and the target masking position are as follows: A step of determining the sound source position based on the command source position of the activation command; and A step of determining the target masking position based on occupant information acquired by sensors in the vehicle; A method that includes this.
3. The method according to claim 1, wherein the sound source position and the target masking position are the sound source position and target masking position input by the occupants in the vehicle.
4. A method according to any one of claims 1 to 3, wherein, after the step of receiving an audio signal from the sound source location, the method further includes: performing an enhancement process on the audio signal.
5. The method according to claim 4, wherein the enhancement process includes echo cancellation and / or adaptive speech noise reduction.
6. In the method according to any one of claims 1 to 5, generating a masking sound for the audio signal is: A method comprising generating the masking sound by performing a time-domain inversion process on the audio signal.
7. The method according to any one of claims 1 to 6, wherein the step of outputting the masking sound to a first speaker at the sound source location and a second speaker at the target masking location is: A step of performing sound field control processing on the masking sound based on the sound source position and the target masking position, and outputting a first masking sound and a second masking sound, wherein the sound field control processing includes adjusting the phase and amplitude of each frequency signal in the masking sound; and The steps include outputting the first masking sound to a first speaker at the sound source location and outputting the second masking sound to a second speaker at the target masking location; A method that includes this.
8. The method according to claim 7, the step of performing sound field control processing on the masking sound based on the sound source position and the target masking position to output a first masking sound and a second masking sound is: The process includes the step of performing sound field control processing on the masking sound based on the sound source position and the target masking position, and outputting N channels of masking sound, wherein the N channels of masking sound include the first masking sound and the second masking sound, and N is the number of speakers in the vehicle cabin. A method wherein the volume of the masking sound received by the first speaker is lower than the volume of the masking sound received by the second speaker.
9. The method according to claim 7, wherein if the target masking position includes the driver's seat, the second speaker includes a speaker located on the headrest of the driver's seat.
10. In the method according to any one of claims 1 to 9, the step of generating a masking sound for the audio signal is: Steps to acquire noise data in a noise database; and A method comprising the steps of generating the masking sound by using the noise data, or generating the masking sound by using the audio signal and the noise data.
11. The method according to claim 10, the step of acquiring noise data in the noise database is: A method comprising the steps of analyzing the audio characteristics of the audio signal and obtaining noise data corresponding to the audio characteristics from the noise database.
12. In the method according to any one of claims 1 to 11, prior to the step of outputting the masking sound to the first speaker at the sound source location and the second speaker at the target masking location, the method further: A method comprising the step of performing automatic gain control adjustment on the masking sound so that the volume of the masking sound fits within a specified range.
13. In the method according to any one of claims 1 to 12, if the target masking position includes the driver's seat, the method further: A method that includes a step of increasing the volume of the vehicle safety warning sound.
14. The method according to claim 13, further: A method comprising the steps of acquiring external sounds, performing specific type identification on the external sounds to identify specific sounds in the external sounds, and outputting the specific sounds to the speaker in the driver's seat.
15. The method according to any one of claims 2 to 14, further: Steps include receiving a command to stop voice masking; and A step of stopping the reception of the sound signal at the sound source location based on the stop command; A method that includes this.
16. It is a voice masking device: A positioning module configured to determine the sound source position and the target masking position; A sound detection module configured to receive an audio signal from the sound source location, detect whether an audio signal is present in the audio signal, and generate a detection result; and A masking sound generation module configured to generate a masking sound based on the detection result, wherein if the detection result indicates that an audio signal is present in the sound signal, it generates a masking sound for the audio signal, or if the detection result indicates that an audio signal is not present in the sound signal, it skips generating a masking sound; and A masking sound post-processing module configured to output the masking sound to a first speaker at the sound source location and a second speaker at the target masking location; A device that includes this.
17. In the apparatus according to claim 16, the position determination module is: Received a voice masking activation command; The sound source position is determined based on the command source position of the aforementioned activation command; and A device specifically configured to determine the target masking position based on occupant information acquired by sensors within the vehicle.
18. The apparatus according to claim 16, wherein the position determination module is specifically configured to determine the sound source position and the target masking position based on the sound source position and the target masking position input by the occupants in the vehicle.
19. The apparatus according to any one of claims 16 to 18, wherein the sound detection module is further configured to perform enhancement processing on the sound signal after receiving the sound signal from the sound source location.
20. The apparatus according to claim 19, wherein the enhancement process includes echo cancellation and / or adaptive speech noise reduction.
21. The apparatus according to any one of claims 16 to 20, wherein the masking sound generation module is specifically configured to generate the masking sound by performing a time-domain inversion process on the audio signal.
22. In the apparatus according to any one of claims 16 to 21, the output of the masking sound to the first speaker at the sound source location and the second speaker at the target masking location by the masking sound post-processing module specifically means: A step of performing sound field control processing on the masking sound based on the sound source position and the target masking position, and outputting a first masking sound and a second masking sound, wherein the sound field control processing includes adjusting the phase and amplitude of each frequency signal in the masking sound; and The steps include outputting the first masking sound to a first speaker at the sound source location and outputting the second masking sound to a second speaker at the target masking location; An apparatus that includes performing the following actions.
23. In the apparatus according to claim 22, when sound field control processing is performed on the masking sound based on the sound source position and the target masking position to output a first masking sound and a second masking sound, the masking sound post-processing module: The system is specifically configured to perform sound field control processing on the masking sound based on the sound source position and the target masking position, and to output N channels of masking sound, wherein the N channels of masking sound include the first masking sound and the second masking sound, and N is the number of speakers in the vehicle cabin. A device in which the volume of the masking sound received by the first speaker is lower than the volume of the masking sound received by the second speaker.
24. The apparatus according to claim 22, wherein if the target masking position includes the driver's seat, the second speaker includes a speaker located on the headrest of the driver's seat.
25. In the apparatus according to any one of claims 16 to 24, the masking sound generation module is: Obtain noise data from the noise database; and A device specifically configured to generate the masking sound by using the noise data, or to generate the masking sound by using the audio signal and the noise data.
26. In the apparatus according to claim 25, when acquiring noise data in the noise database, the masking sound generation module: A device specifically configured to analyze the audio characteristics of the aforementioned audio signal and to acquire noise data corresponding to the aforementioned audio characteristics from the noise database.
27. In the apparatus according to any one of claims 16 to 26, before outputting the masking sound to the first speaker at the sound source location and the second speaker at the target masking location, the masking sound post-processing module further: A device configured to perform automatic gain control adjustment on the masking sound so that the volume of the masking sound falls within a specified range.
28. In the apparatus according to any one of claims 16 to 27, if the target masking position includes the driver's seat, the masking sound post-processing module further: A device configured to increase the volume of vehicle safety warning sounds.
29. In the apparatus according to claim 28, the masking sound post-processing module further comprises: A device configured to acquire external sounds, perform specific type identification on the external sounds to identify specific sounds within the external sounds, and output the specific sounds to the speaker in the driver's seat.
30. In the apparatus according to any one of claims 17 to 29: The position determination module is further configured to receive a voice masking stop command; and The sound detection module is further configured to stop receiving the sound signal from the sound source location based on the stop command.
31. A voice masking device comprising a processor and a memory, wherein the memory is configured to store instructions, and the processor is configured to execute instructions stored in the memory to perform the method according to any one of claims 1 to 15.
32. A voice masking system used in vehicles: A voice masking activation and / or deactivation device configured to control the activation and / or deactivation of a voice masking function; The apparatus according to any one of claims 16 to 31; A microphone positioned at a sound source location and configured to capture sound signals at the sound source location; A first speaker positioned at the sound source location; and A second speaker positioned at the aforementioned target masking location; A system comprising the first speaker and the second speaker configured to reproduce a masking sound.
33. The system according to claim 32, further comprising a sensor device, wherein the sensor device is configured to collect information relating to occupants in the vehicle.
34. The system according to claim 32, further comprising an information input device, the information input device being used by an occupant in the vehicle to input the sound source location and the target masking location.
35. A system according to any one of claims 32 to 34, wherein the microphone further includes a third microphone located outside the vehicle, the third microphone located outside the vehicle being configured to capture external sounds.
36. A vehicle comprising the device according to any one of claims 16 to 31, or the system according to any one of claims 32 to 35.