Sound masking method, device and terminal equipment

By generating and transmitting masked sound signals in mobile communication devices, the problem of privacy leaks for eavesdroppers is solved while ensuring that the call quality of the listener is not affected, thus achieving privacy protection without compromising the call quality of the listener.

CN116684514BActive Publication Date: 2026-07-31HUAWEI TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2020-03-20
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

During a call on a mobile communication device, an eavesdropper may hear the voice transmitted from the receiver, leading to the leakage of private or confidential information. At the same time, existing technologies affect the call quality of the listener when processing voice.

Method used

By generating a masking sound signal in the terminal device and emitting the signal through a speaker to mask the audio signal output by the receiver, the masking sound signal is ensured to be correlated with the audio signal and effectively masked in the far field, reducing interference to the listener's ears.

Benefits of technology

It effectively prevents eavesdroppers from learning about the private or confidential information contained in the call, without affecting the call quality for the listener.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116684514B_ABST
    Figure CN116684514B_ABST
Patent Text Reader

Abstract

This application discloses a sound masking method, apparatus, and terminal device. When the terminal device uses a receiver as the output terminal for audio signals, a masking sound signal is determined based on the audio signal, and then emitted through a speaker. Since the masking sound signal is determined based on the audio signal, and the distance difference between the speaker and the receiver relative to the far field is small, the masking sound signal can effectively mask sound leakage from the receiver, preventing information leakage in the voice call. Furthermore, since the masking sound signal and the audio signal are output by the speaker and receiver respectively, when the listener listens to the audio signal through the receiver, the distance difference between the speaker and the receiver relative to the listener's ear is large. Therefore, the masking sound signal has minimal interference with the listener's audio signal reception and will not affect the listener's call quality.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese application entitled "A method, apparatus and terminal device for masking sound", application number 2020102020575, and application date March 20, 2020. Technical Field

[0002] This application relates to the field of electronic equipment technology, and in particular to a sound masking method, apparatus and terminal device. Background Technology

[0003] With the continuous development of mobile communication technology, mobile communication devices have become widely used in people's production and daily lives. People establish communication connections and conduct calls through mobile communication devices. Typically, when using a mobile communication device for a call, the listener hears the other party's voice through the receiver. However, when the listener is in a quiet room, the other party's volume is loud, or the receiver's playback volume is loud, other people nearby (hereinafter referred to as eavesdroppers) may also hear the voice transmitted through the receiver. On the one hand, the voice transmitted through the receiver can easily interfere with eavesdroppers; on the other hand, this may lead to some or all of the voice being leaked to the eavesdroppers' ears, and the security of private or confidential information in the voice cannot be guaranteed.

[0004] To address the aforementioned issues, current methods can prevent the leakage of private or confidential information during phone calls by processing the audio to be played on the receiver. However, this method requires processing the audio to be played, which can negatively impact the listener's call quality.

[0005] Currently, ensuring the privacy of call content and avoiding interference with the call quality of the listener when making calls using mobile communication devices has become an urgent technical problem to be solved in this field. Summary of the Invention

[0006] This application provides a sound masking method, apparatus, and terminal device to prevent bystanders from understanding information leaked from the receiver without affecting the call quality of the listener.

[0007] Firstly, this application provides a sound masking method. This method is applicable to any terminal device with communication capabilities and including a receiver and a speaker. The method includes:

[0008] When the terminal device uses the receiver as the output of the audio signal, the masking sound signal is determined according to the audio signal.

[0009] A masking sound signal is emitted through a speaker, which is used to mask the audio signal output by a far-field receiver. For example, the speaker can be controlled to output this masking sound signal.

[0010] In this method, the masking sound signal is determined based on the audio signal. Furthermore, for the far field, the distances between the speaker and the far field are relatively small, similar to the distances between the receiver and the far field, thus the audio signal can be effectively masked by the masking sound signal. Additionally, the attenuation of the audio signal is significantly less than that of the listener's ear; the masking sound signal is actually masked by the audio signal. Therefore, the masking sound signal emitted by the speaker will not interfere with the listener's ear, thereby avoiding any impact on the listener's call quality.

[0011] The audio signal played by the receiver can include various types of audio signals, such as human voice signals, animal sound signals, or music.

[0012] The masking sound signal can be determined based on the audio signal. For example, in the first implementation of the first aspect, the corresponding masking sound signal can be selected or matched from a pre-generated audio library based on the audio signal. In the second implementation of the first aspect, the masking sound signal can also be generated based on the real-time downlink audio signal.

[0013] Spectral characteristic analysis can be used to match pink noise or white noise signals as masking sound signals, and it can also generate pink noise or white noise signals as masking sound signals. Furthermore, specific analysis can be performed on the audio signal to extract features and match targeted masking sound signals; or targeted processing operations can be performed on the audio signal, such as time-domain inversion and enhancement processing, to generate masking sound signals.

[0014] Combining the second implementation method of the first aspect, the masking sound signal is determined based on the audio signal, which may specifically include:

[0015] Perform spectral analysis on the audio signal to obtain the spectral response;

[0016] The masked sound signal is generated based on the spectral response.

[0017] A masking sound signal is generated based on the spectral response of the audio signal, ensuring that the spectral response of the masking sound signal matches that of the audio signal, thereby enabling the masking sound signal to effectively mask the audio signal output by the receiver.

[0018] Combining the second implementation method of the first aspect, the masking sound signal is determined based on the audio signal, which may specifically include:

[0019] The audio signal is extracted according to the preset frame length to obtain the extracted sound segment;

[0020] Reverse the time domain of a sound clip to obtain the reversed sound;

[0021] The masked sound signal can be generated by directly splicing the inverted sounds, or by splicing them after applying a window function.

[0022] By inverting the generated masking sound signal, the masking sound signal is difficult for bystanders to understand. Thus, when the masking sound signal is transmitted to the far field, the leakage sound from the receiver can be masked by the masking sound signal, which is difficult to understand.

[0023] Combining the second implementation method of the first aspect, the masking sound signal is determined based on the audio signal, which may specifically include:

[0024] The audio signal is extracted according to the preset frame length to obtain the extracted sound segment;

[0025] The sound segments are interpolated to obtain the supplemented sound signal; or subsequent segments are matched from a preset audio library to obtain the supplemented sound signal.

[0026] The corresponding masking sound signal is generated based on the supplemented sound signal.

[0027] By obtaining supplementary audio signals through interpolation or matching, frequent audio segmentation is avoided, reducing the processing load on terminal devices and improving the efficiency of generating masked audio signals.

[0028] Optionally, in conjunction with the second implementation of the first aspect, the method may further include:

[0029] Obtain the duration of the masking sound signal generated from the audio signal;

[0030] The audio signal is delayed according to the duration of the delay so that the audio signal output by the receiver matches the masking sound signal output by the speaker.

[0031] For example, after the delay, the audio signal and the masking sound signal are partially or completely aligned. By aligning the masking sound signal and the audio signal, the synchronization between the masking sound signal and the audio signal is ensured, thus improving the masking effect.

[0032] In combination with any of the implementations of the first aspect, the method may also include:

[0033] The masking sound signal is phase-inverted to obtain an inverted sound signal;

[0034] The amplitude of the inverted sound signal is reduced and then mixed with the audio signal to obtain a mixed sound signal.

[0035] The receiver outputs a mixed audio signal.

[0036] By inverting the phase, the mixed audio signal can, to some extent, offset the effect of the masking audio signal emitted by the speaker on the ears of near-field listeners, thereby further improving call quality while ensuring call privacy.

[0037] In combination with any implementation of the first aspect, this method transmits a masking sound signal through a speaker, specifically including:

[0038] The method detects ambient sound signals, and when the amplitude of the ambient sound signals falls below a first preset threshold, it emits a masking sound signal through a speaker. This reduces the risk of leaked audio content.

[0039] In combination with any implementation of the first aspect, this method transmits a masking sound signal through a speaker, specifically including:

[0040] When a downlink audio signal is detected, if the amplitude of the downlink audio signal exceeds a second preset threshold, a masking sound signal is emitted through the speaker. This method avoids causing unnecessary noise interference to listeners.

[0041] Optionally, the preset frame length can be set to a value between 10ms and 300ms.

[0042] Optionally, the phase range of the inversion process is 90 degrees to 270 degrees.

[0043] The distance between the receiver and the speaker is greater than the width of the terminal device;

[0044] The distance between the receiver and the speaker is greater than half the length of the terminal device;

[0045] The distance between the receiver and the speaker is greater than 100mm;

[0046] The distance between the receiver and the speaker should be at least 20 times the distance between the receiver and the listener's ear.

[0047] Secondly, this application provides a sound masking device. This device can be applied to any terminal device with communication capabilities that includes a receiver and a speaker.

[0048] The device includes: a judgment module, a determination module, and a first control module.

[0049] The judgment module is used to determine whether the terminal device uses the receiver as the output end of the audio signal.

[0050] The determination module is used to determine the masking sound signal based on the audio signal when the judgment result of the judgment module is yes;

[0051] The first control module is used to control the loudspeaker to emit a masking sound signal in order to mask the audio signal output by the receiver in the far field.

[0052] In a first implementation of the second aspect, the determining module is used to select or match a corresponding masking sound signal from a pre-generated audio library based on the audio signal when the receiver outputs an audio signal.

[0053] In the second implementation of the second aspect, the determining module is used to generate a masking sound signal in real time based on the audio signal.

[0054] Combining the second implementation method of the second aspect, the module can specifically include:

[0055] The spectrum analysis unit is used to perform spectrum analysis on audio signals and obtain the spectrum response;

[0056] The first generation unit is used to generate a masked sound signal based on the spectral response.

[0057] Combining the second implementation method of the second aspect, the module can specifically include:

[0058] The signal extraction unit is used to extract audio signals according to a preset frame length to obtain the extracted sound segments.

[0059] The signal inversion unit is used to invert the time domain of a sound segment to obtain an inverted sound.

[0060] The second generation unit is used to generate a corresponding masking sound signal based on the reversed sound.

[0061] Combining the second implementation method of the second aspect, the module can specifically include:

[0062] The signal extraction unit is used to extract audio signals according to a preset frame length to obtain the extracted sound segments.

[0063] The signal supplementation unit is used to interpolate sound segments to obtain a supplemented sound signal; or to match subsequent segments from a preset audio library to obtain a supplemented sound signal.

[0064] The third generation unit is used to generate a corresponding masking sound signal based on the supplemented sound signal.

[0065] In conjunction with the second implementation of the second aspect, the device may further include:

[0066] The time length acquisition module is used to obtain the time length for generating the masked sound signal based on the audio signal;

[0067] The delay module is used to delay the audio signal according to the time length so that the audio signal output by the receiver matches the masking sound signal output by the speaker.

[0068] In conjunction with any implementation of the second aspect, the above-mentioned apparatus may further include:

[0069] The phase inversion processing module is used to invert the phase of the masking sound signal to obtain an inverted sound signal;

[0070] The mixing module is used to reduce the amplitude of the inverted audio signal and then mix it with the audio signal to obtain a mixed audio signal.

[0071] The second control module is used to control the output of the mixed audio signal from the receiver.

[0072] In combination with any of the implementation methods of the second aspect, the first control module may specifically include:

[0073] The first detection unit is used to detect sound signals from the surrounding environment;

[0074] The first judgment unit is used to determine whether the amplitude of the ambient sound is lower than the first preset threshold.

[0075] The first control unit is used to transmit a masking sound signal through a speaker when the first judgment unit determines that the result is yes.

[0076] In combination with any of the implementation methods of the second aspect, the first control module may specifically include:

[0077] The second detection unit is used to detect whether there is a downlink audio signal;

[0078] The second judgment unit is used to determine whether the amplitude of the downlink audio signal is greater than a second preset threshold when the second detection unit detects a downlink audio signal.

[0079] The second control unit is used to emit a masking sound signal through a speaker when the second judgment unit determines that the result is yes.

[0080] In conjunction with any implementation of the second aspect, the apparatus may further include:

[0081] The signal enhancement module is used to enhance the masked sound signal, obtain the enhanced masked sound signal, and then provide it to the speaker.

[0082] Thirdly, this application provides a terminal device. This terminal device can be any terminal device with communication capabilities and including a receiver and speaker, such as a mobile phone, tablet computer, personal digital assistant (PDA), point of sales (POS), or in-vehicle computer.

[0083] The terminal equipment provided in the third aspect of this application includes: a receiver, a speaker, and a processor;

[0084] A processor, used to determine or generate a masking sound signal based on an audio signal when an audio signal is output from a receiver;

[0085] A loudspeaker is used to emit masking sound signals to mask the audio signals output by a receiver in the far field.

[0086] In the first implementation of the third aspect, the processor is specifically used to select or match a corresponding masking sound signal from a pre-generated audio library based on the audio signal when the receiver outputs an audio signal.

[0087] In the second implementation of the third aspect, the processor is specifically used to generate a masking sound signal in real time based on the audio signal.

[0088] In conjunction with the second implementation method of the third aspect, the processor is specifically used to perform spectral analysis on the audio signal to obtain the spectral response; and to generate a masked sound signal based on the spectral response.

[0089] In conjunction with the second implementation method of the third aspect, the processor is specifically used to extract audio signals according to a preset frame length to obtain the extracted sound segments; to perform time-domain inversion on the sound segments to obtain inverted sound; and to generate corresponding masking sound signals based on the inverted sound.

[0090] In conjunction with the second implementation method of the third aspect, the processor is specifically used to extract audio signals according to a preset frame length to obtain the extracted sound segments; interpolate the sound segments to obtain the supplemented sound signals; or match subsequent segments from a preset audio library to obtain the supplemented sound signals; and generate corresponding masking sound signals based on the supplemented sound signals.

[0091] In conjunction with the second implementation of the third aspect, the processor is also used to obtain the time length for generating the masking sound signal based on the audio signal; and to delay the audio signal based on the time length so that the audio signal output by the receiver is adapted to the masking sound signal output by the speaker.

[0092] In conjunction with any of the implementation methods of the third aspect, the processor is also used to perform phase inversion processing on the masking sound signal to obtain an inverted sound signal; to perform amplitude reduction processing on the inverted sound signal and mix it with the audio signal to obtain a mixed sound signal; and to control the receiver to output the mixed sound signal.

[0093] In combination with any of the implementation methods of the third aspect, the processor is specifically used to detect the sound signals of the surrounding environment. When the amplitude of the sound signals of the surrounding environment is lower than a first preset threshold, it emits a masking sound signal through a speaker.

[0094] In conjunction with the second implementation of the third aspect, the processor is specifically used to emit a masking sound signal through the speaker when it is determined that the amplitude of the downlink audio signal is greater than a second preset threshold.

[0095] This application has at least the following advantages:

[0096] The sound masking method provided in this application is applied to a terminal device with a receiver and a speaker. When the terminal device outputs an audio signal through the receiver, a masking sound signal is determined based on the audio signal, and then the speaker is controlled to emit the masking sound signal. Since the masking sound signal is determined based on the audio signal, and the distance difference between the speaker and the receiver relative to the far field is small, the masking sound signal can effectively mask the receiver's sound leakage, preventing information leakage in the voice call. Furthermore, since the masking sound signal and the audio signal are output by the speaker and receiver respectively, when the listener listens to the audio signal through the receiver, the distance difference between the speaker and the receiver relative to the listener's ear is large. Therefore, the masking sound signal has minimal interference with the listener's audio signal reception and will not affect the listener's call quality. Attached Figure Description

[0097] Figure 1 A schematic diagram illustrating an application scenario of the sound masking method provided in this application embodiment;

[0098] Figure 2 for Figure 1 The diagram shows the distance between the first terminal device and the ears of the near-field listener and the far-field bystander.

[0099] Figure 3 A flowchart illustrating a sound masking method provided in an embodiment of this application;

[0100] Figure 4 A flowchart illustrating another sound masking method provided in an embodiment of this application;

[0101] Figure 5 This is a schematic diagram of the spectral response obtained by performing spectral analysis on an audio signal, provided in an embodiment of this application.

[0102] Figure 6 A schematic diagram illustrating the masking of sound signals and their spectral characteristics.

[0103] Figure 7 This is a schematic diagram of signal processing provided for an embodiment of this application;

[0104] Figure 8 A schematic diagram of a truncated audio segment provided in an embodiment of this application;

[0105] Figure 9 for Figure 8 The diagram shows the inverted sound after time-domain reversal of the sound clip shown.

[0106] Figure 10 A schematic diagram of a masking sound signal and an inverted sound signal provided in an embodiment of this application;

[0107] Figure 11 A schematic diagram showing the comparison of the masked sound signal before and after enhancement, provided in an embodiment of this application;

[0108] Figure 12 A schematic diagram of the structure of a sound masking device provided in an embodiment of this application;

[0109] Figure 13 This is a schematic diagram of the structure of a signal generation module provided in an embodiment of this application;

[0110] Figure 14 This is a schematic diagram of another signal generation module provided in an embodiment of this application;

[0111] Figure 15 This is a schematic diagram of the structure of another signal generation module provided in an embodiment of this application;

[0112] Figure 16 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application;

[0113] Figure 17 This is a hardware architecture diagram of a mobile phone provided in an embodiment of this application. Detailed Implementation

[0114] Mobile communication devices, as commonly used communication devices, meet people's communication needs in various scenarios due to their portability. For example, people can use mobile communication devices to make calls on crowded subways, in bustling commercial streets, or in empty changing rooms. When receiving audio signals with a receiver, if the listener is in a quiet room, the other party's volume is too loud, or the receiver's playback volume is too high, the problem of audio leakage is unavoidable. The call may contain content that the listener does not want others to hear, which may involve private or confidential information. Audio leakage from the receiver can easily lead to the leakage of private or confidential information during the call.

[0115] Adjusting the receiver volume to prevent audio leakage lacks controllability. When the listener turns down the receiver volume, leakage may already be occurring. While processing the audio signal to be played through the receiver can reduce the listener's intelligibility to some extent, it also makes it difficult for the listener to hear the other party's conversation. Therefore, this method negatively impacts the listener's call quality.

[0116] Based on the above problems, this application provides a sound masking method, apparatus, and terminal device. In this application, when the receiver is used as the output of the audio signal (i.e., the masked tone), a masking sound signal (i.e., the masking tone) is determined based on the audio signal, and then the speaker is controlled to emit the masking sound signal. Since the masking tone is determined based on the masked tone, the two are correlated, and the distance between the listener and the receiver is close to the distance between the listener and the speaker, the masking tone can mask the masked tone from the listener. Simultaneously, since the distance from the masking tone to the listener's ear is much greater than the distance from the masked tone to the listener's ear, the masking tone will not affect the listener's normal listening to the audio signal. The masking sound signal emitted by the speaker is actually masked by the audio signal emitted by the receiver at the listener's ear, thus ensuring that the listener's call quality is not interfered with. Generally speaking, a position with a distance greater than a critical value (1 / r-law) from the sound source is defined as the far field, and a position with a distance less than or equal to a critical value (1 / r-law) from the sound source is defined as the near field. To facilitate understanding of the technical solutions in this application, since the distance between the listener and the terminal device is much smaller than the distance between the bystander and the terminal device, the terminal device is used as the sound source, and the listener is set to be in the near field and the bystander to be in the far field.

[0117] To enable those skilled in the art to better understand the technical solutions provided in the embodiments of this application, the application scenarios of the sound masking method provided in the embodiments of this application will be introduced below.

[0118] See Figure 1 The figure is a schematic diagram of an application scenario of the sound masking method provided in the embodiments of this application.

[0119] Figure 1 In this embodiment, the first terminal device 101 and the second terminal device 102 establish a communication connection. The sound masking method provided in this application is applied to the first terminal device 101, and the second terminal device 102 serves as the peer device of the first terminal device 101. In practical applications, the first terminal device 101 can be any mobile communication device with communication functions, such as a mobile phone, tablet computer, or portable laptop computer. Figure 1 The example shown here is only a mobile phone-type first terminal device 101. The specific type of the first terminal device 101 is not limited in this embodiment.

[0120] The first terminal device 101 includes a receiver 1011 and a speaker 1012. In this embodiment, the receiver 1011 and speaker 1012 of the first terminal device 101 may be respectively disposed on both sides of the first terminal device 101. Figure 1 As shown, the receiver 1011 is located on one side of the centerline 1013 in the length direction of the first terminal 101, and the speaker 1012 is located on the other side of the centerline 1013.

[0121] Both the receiver 1011 and the speaker 1012 can output audio signals from the second terminal device 102 to the outside world. During a call, the user (i.e., the listener) of the first terminal device 101 can choose to output audio signals via the receiver 1011 or the speaker 1012 according to their needs. The technical solution of this application embodiment is mainly based on a scenario where the receiver 1011 is used as the audio signal output terminal. When the receiver 1011 is used as the audio signal output terminal, the speaker 1012 outputs a masking sound signal. This masking sound signal is used to mask audio signals to the far field.

[0122] See Figure 2 The image is Figure 1 The diagram shows the distance between the first terminal device 101 and the ears of near-field listeners and far-field bystanders.

[0123] Figure 2 In this embodiment, when the receiver 1011 of the first terminal device 101 is used as the output terminal of the audio signal, the distance between the listener (the user of the first terminal device 101, i.e., the target receiver of the audio signal) and the first terminal device 101 is much smaller than the distance between the bystander (other people located near the listener, i.e., the non-target receiver of the audio signal) and the first terminal device 101. Therefore, in this embodiment, the listener's ear 201 is located in the near field of the first terminal device 101, and the bystander's ear 202 is located in the far field of the first terminal device 101.

[0124] When the receiver 1011 is used as the audio signal output, the listener's ear 201 is positioned close to the receiver 1011. At this time, the first distance d1 between the listener's ear 201's ear reference point (ERP) and the receiver 1011 is approximately 5mm, and the second distance d2 between the listener's ear 201's ERP and the speaker 1012 is approximately 150mm. It is evident that the second distance d2 differs from the first distance d1 by nearly 30 times. The third distance d3 between the listener's ear 202's ERP and the receiver 1011 is approximately 500mm, and the fourth distance d4 between the listener's ear 202's ERP and the speaker 1012 is approximately 500mm. It is evident that the difference between the fourth distance d4 and the third distance d3 is very small, with a difference factor far less than 30.

[0125] Based on the relationship between the sound pressure radiated by the pulsating spherical source in space and the radial distance of the radiation, as well as the conversion relationship between sound pressure level and sound pressure, the relationship between sound pressure level and distance can be obtained: the sound pressure amplitude decreases inversely with the increase of radial distance. Assuming that receiver 1011 and speaker 1012 output audio signals and masking sound signals of equal volume respectively, since the second distance d2 differs from the first distance d1 by approximately 30 times, the attenuation amplitude of the masking sound signal heard by the listener's ear 201 is greater than the attenuation amplitude of the heard audio signal. Therefore, the masking sound signal will not interfere with the listener's ear 201's reception of the audio signal. Furthermore, since the difference between the fourth distance d4 and the third distance d3 is very small, typically less than 2 times, the attenuation amplitude of the masking sound signal heard by the bystander's ear 202 is close to the attenuation amplitude of the heard audio signal, and the loudness of the masking sound signal and the audio signal received by the bystander's ear 202 is basically the same. In addition, since the masking sound signal is determined based on the audio signal, the masking sound signal can mask the audio signal in the far field.

[0126] It should be noted that, Figure 2 The values ​​of the first distance d1, second distance d2, third distance d3, and fourth distance d4 shown are merely examples. In practical applications, the first distance d1 may be related to the listener's hearing ability or the listener's posture when holding the first terminal device. For example, if the listener has poor hearing, the first distance d1 may be less than 5mm; while if the listener has good hearing, the first distance d1 may be greater than 5mm, for example, d1 = 10mm. As the length of the first terminal device 101 changes, the second distance d2 may be greater than or less than 150mm. In addition, the values ​​of the third distance d3 and the fourth distance d4 may also change according to the relative position of the listener and the first terminal device, so the third distance d3 and the fourth distance d4 may also be greater than 500mm. Therefore, the embodiments of this application do not limit the value of d1, d2, d3, and d4.

[0127] See Figure 3 The figure is a flowchart of a sound masking method provided in an embodiment of this application.

[0128] like Figure 3 As shown, sound masking methods include:

[0129] S301: Determine whether the terminal device uses the receiver as the output terminal for audio signals. If so, execute S302.

[0130] The audio signal output by the receiver can include human voice signals, animal sounds, or music, etc. No restrictions are placed on the type or specific content of the audio signal here.

[0131] In one possible implementation, it is determined whether the receiver has received an audio signal output command. When the receiver receives an audio signal output command, it indicates that the receiver is being used as the audio signal output terminal. Similarly, when the speaker receives an audio signal output command, it indicates that the speaker is being used as the audio signal output terminal. In practical applications, the audio signal output command may be sent only to the receiver or only to the speaker.

[0132] Since the speaker provides voice playback when it is working, when the terminal device uses the speaker as the output of the audio signal, it means that the user of the terminal device does not need to keep the content of the call confidential. At this time, there is no need to perform the subsequent operation of the method in this embodiment to mask the audio signal.

[0133] When a terminal device uses a receiver as the audio signal output, the conversation between the user of the terminal device and the user of the other terminal device may involve privacy or confidential information. If the receiver leaks audio, an eavesdropper may be able to learn about the private or confidential information in the audio signal. Therefore, it is necessary to mask the audio signal.

[0134] S302: When the terminal device uses the receiver as the output of the audio signal, the masking sound signal is determined according to the audio signal.

[0135] For example, the audio signal provided by the other party's terminal device has corresponding time-domain and frequency-domain characteristics. In this embodiment of the application, the masking sound signal can be generated in real time after analyzing the time-domain and / or frequency-domain characteristics of the real-time audio signal.

[0136] In addition, an audio library containing various masking sound signals can be constructed by analyzing the time-domain and / or frequency-domain characteristics of historical audio signals in advance. When the terminal device uses the receiver as the output of the audio signal again, a masking sound signal is selected or matched from the preset audio library according to the time-domain and / or frequency-domain characteristics of the current audio signal to mask the current audio signal.

[0137] The masking sound signal determined based on the audio signal has a higher degree of matching with the audio signal in the time domain and / or frequency domain, and can better mask the audio signal at a distance, reducing the understanding of private or confidential information in the leaked sound by listeners in the far field.

[0138] S303: The speaker emits a masking sound signal, which is used to mask the audio signal output by the far-field receiver.

[0139] In practical applications, the speaker can be controlled to emit masking sound signals based on a variety of possible triggering conditions.

[0140] In one possible implementation, the terminal device controls the speaker to continuously emit a masking sound signal throughout the entire call, with the receiver serving as the audio signal output.

[0141] When the environment around a terminal device is noisy, even if the receiver leaks sound, the content of the call is difficult for an eavesdropper to hear. Conversely, when the environment is quiet, receiver leakage can easily lead to the disclosure of private or confidential information during the call. To avoid this problem, another possible implementation involves detecting the ambient sound signal. When the amplitude of the ambient sound signal falls below a first threshold, it indicates that the environment is too quiet. In this case, the speaker needs to be controlled to emit a masking sound signal.

[0142] During a call on a terminal device, there may be gaps or periods of low amplitude in the audio signal. In these situations, the masking signal will not function effectively and may even cause noise interference for listeners. To avoid this problem, in another possible implementation, when a downlink audio signal is detected, its amplitude is checked to see if it exceeds a second preset threshold. If the downlink audio signal amplitude exceeds the second preset threshold, it indicates that the volume of the downlink audio signal is high, and the receiver may leak sound. In this case, the speaker is controlled to emit a masking signal.

[0143] In practical applications, the sound masking method provided in this embodiment can be implemented according to the user's selection. For example, if an application (full name: Application, abbreviation: APP) with the function of masking audio signals is installed on the terminal device, the user can manipulate the function options on the APP to select whether to enable or disable the audio signal masking function. When the function is enabled, the terminal device can execute the sound masking method provided in this embodiment. In addition, a function module can also be embedded in the call interface of the terminal device. This function module can be enabled or disabled according to the user's selection. When the function module is enabled, the terminal device can execute the sound masking method provided in this embodiment.

[0144] In addition, the aforementioned apps or functional modules can also be turned on automatically. For example, they can be turned on automatically when a downlink audio signal is detected, when a call request is received, or when the terminal device is powered on.

[0145] The above describes the sound masking method provided in this application embodiment. This method is applied to a terminal device with a receiver and a speaker. When the terminal device outputs an audio signal through the receiver, a masking sound signal is determined based on the audio signal. The speaker is then controlled to emit the masking sound signal. Since the masking sound signal is determined based on the audio signal, and the distance difference between the speaker and receiver relative to the far field is small, the masking sound signal can effectively mask sound leakage from the receiver, reducing the intelligibility of the leaked sound for onlookers and preventing information leakage in the conversation. Furthermore, since the masking sound signal and the audio signal are output by the speaker and receiver respectively, when the listener listens to the audio signal through the receiver, the distance difference between the speaker and receiver relative to the listener's ear is large. Therefore, the masking sound signal has minimal interference with the listener's audio signal reception and will not affect the listener's call quality.

[0146] In practical applications, a 15dB difference in sound pressure levels between two sound signals reaching the listener's ear can create a noticeable difference. To prevent masking sounds from interfering with the listener's call quality, in one possible implementation, the second distance d2 between the speaker of the terminal device and the listener's ear is greater than the first distance between the receiver and the listener's ear. For example, if d1 is more than 10 times d2, the sound pressure level of the audio signal reaching the listener's ear is more than 20dB higher than the sound pressure level of the masking sound signal. In this case, the listener's hearing of the audio signal will not be interfered with by the masking sound signal.

[0147] by Figure 1 Taking the first terminal device 101 shown as an example, the length of the first terminal device 101 is L1, the width is W1, and the distance between the speaker 1012 and the receiver 1011 is L2. L2 satisfies at least one of the following inequalities (1)-(3):

[0148] L2>W1 (1)

[0149] L2>0.5*L1 (2)

[0150] L2>100mm (3)

[0151] When L2 satisfies at least one of the inequalities (1)-(3), it can be guaranteed that the first distance d1 is much smaller than the second distance d2, thereby ensuring that the masked sound signal reaches the listener's ear at a lower sound pressure level than the audio signal, thus avoiding interference with the listener's call quality caused by the masked sound signal.

[0152] In another possible implementation, L2 satisfies the following inequality (4):

[0153] L2≥20*d1 (4)

[0154] The sound masking method provided in this application includes various implementations for generating masked sound signals. These will be described below in conjunction with embodiments and accompanying drawings.

[0155] See Figure 4 The figure is a flowchart of another sound masking method provided in an embodiment of this application.

[0156] like Figure 4 As shown, the sound masking method includes:

[0157] S401: Determine whether the terminal device uses the receiver as the output terminal for audio signals. If so, execute S402.

[0158] The implementation of S401 is basically the same as the implementation of S301 in the aforementioned method embodiment. The relevant description of S401 can be referred to the aforementioned embodiment, and will not be repeated here.

[0159] S402: When the terminal device uses the receiver as the output of the audio signal, it performs spectrum analysis on the audio signal to obtain the spectrum response.

[0160] Spectral analysis of a sound signal to obtain its spectral response is a relatively mature technique in this field; therefore, the specific implementation of S402 will not be elaborated here. For easier understanding, please refer to [link to relevant documentation]. Figure 5 The figure shows a schematic diagram of the spectral response obtained by executing S402. Figure 5 The horizontal axis represents frequency (unit: Hz), and the vertical axis represents signal amplitude (unit: dBFS).

[0161] S403: Generates a masked sound signal based on the spectral response.

[0162] In practical applications, the spectral response curve obtained from S402 can be used as a filter to generate a masked sound signal. The generated masked sound signal can take many forms, such as random noise signals, like white noise or pink noise signals whose frequency response curves are consistent with the audio signal.

[0163] S404: The speaker emits a masked sound signal.

[0164] The implementation of S404 is basically the same as the implementation of S303 in the aforementioned method embodiment. The relevant description of S404 can be referred to the aforementioned embodiment, and will not be repeated here.

[0165] In the sound masking method provided in this application embodiment, since the masking sound signal is generated based on the spectral response of the audio signal, the masking sound signal has a similar or identical characteristic curve to the audio signal in the spectrum. The amplitude of the masking sound signal can be the same as or different from the amplitude of the audio signal. See also Figure 6 In the figure, the spectral characteristic curve of the masked sound signal 601 is very similar to that of the audio signal 602. Therefore, the masked sound signal 601 has a good effect on masking the audio signal 602 in the far field.

[0166] It should be noted that in the embodiments described above, the masking sound signal can be generated in real time based on the audio signal of the current call, for example, based on the first n milliseconds of the audio signal of the current call (n is a positive number, and n milliseconds is less than the total duration of the downlink audio signal).

[0167] Alternatively, the masked audio signal can be pre-generated based on audio signals from previous calls between the terminal device and the peer device. For example, when the peer device 102 previously communicated with the local device 101, it sent an audio signal to the local device 101. This audio signal contained the spectral characteristics of the voice of user A2 from the peer device 102. Before the peer device 102 re-establishes a communication connection with the local device 101, a masked audio signal V2 corresponding to user A2 is generated based on the spectral response of the audio signal it provided. Similarly, a corresponding masked audio signal V3 can be created for user A3. In this way, a mapping table between each contact in the address book of the terminal device 101 and the masked audio signal can be established, and masked audio signals V2, V3, etc., can be added to the audio library. When a contact in the address book establishes a communication connection with this device 101 through their terminal device, if the receiver of the terminal device 101 is used as the output end of the audio signal, the masking sound signal corresponding to the contact can be obtained directly by selecting or matching from the audio library using the mapping table, and then the audio signal of the contact can be masked to the far field.

[0168] By pre-generating the masking sound signal using the above method, the generation efficiency of the masking sound signal is improved, and the masking effect is more targeted. In this implementation, since S402 and S403 are completed in advance, only S401 and S404 are executed each time the method of this embodiment is implemented after generation.

[0169] In practical applications, to further prevent masked sound signals from interfering with the listener's conversation, the audio signal can be processed to weaken the impact of the masked sound signal on the listener. This will be explained below with reference to the accompanying drawings and embodiments.

[0170] See Figure 7 This figure is a schematic diagram of signal processing provided in an embodiment of this application.

[0171] like Figure 7As shown, the audio signal in the terminal device is divided into two paths, 701 and 702, with identical content. Audio signal 701 is provided to the receiver, and audio signal 702 is provided to the speaker. As one possible implementation, audio signal 702 can be obtained by copying audio signal 701.

[0172] The process of generating the masked audio signal is described below. To reduce the intelligibility of the receiver's leaked audio to the listener, in this embodiment, the audio signal 702 is truncated according to a preset frame length to obtain the truncated audio segment; then, the audio segment is time-reversed to obtain the reversed audio. In one possible implementation, the preset frame length can be a fixed frame length or a floating frame length (i.e., the frame length is variable). It is understandable that if the preset frame length is too large, it may take too long to generate the masked audio signal, affecting the listener's call experience. In this embodiment, the preset frame length ranges from 10ms to 300ms, ensuring that the audio segment is time-reversed at a relatively fast frequency to facilitate real-time masking of the audio signal to the far field.

[0173] See Figure 8 and Figure 9 ,in, Figure 8 The image shown is a clipped audio recording. Figure 9 for Figure 8 The sound segment shown is the inverted sound after time-domain reversal. Understandably, the intelligibility of the inverted sound is significantly reduced compared to the original sound segment. The inverted sound can be used to generate the corresponding masking sound signal 703. For example, the masking sound signal can be generated by directly splicing the inverted sound of each frame, or by processing the inverted sound using a window function and then splicing the processed sound to generate the masking sound signal.

[0174] In practical applications, the generation time of the masking sound signal 703 may be delayed relative to the audio signal 701. For example, the masking sound signal 703 may lag behind the audio signal 701 by several milliseconds. To further improve the masking effect, in this embodiment, the time length for generating the masking sound signal 703 based on the audio signal 702 can be obtained, and the audio signal 701 can be delayed according to this time length so that the audio signal 701 output by the receiver is adapted to the masking sound signal 703 output by the speaker, such as partial alignment or complete alignment. For example, if it takes 10ms to generate the masking sound signal, then the audio signal 701 is delayed by 10ms. In addition, the audio signal 701 can also be delayed according to a preset delay length, the value of which ranges from 10ms to 300ms. It should be noted that in this embodiment, delaying the audio signal 701 is an optional operation, not a necessary one.

[0175] The masking sound signal 703 is directly provided to the speaker so that the speaker can output the masking sound signal. In addition, to reduce the interference of the masking sound signal 703 on the ears of near-field listeners, this embodiment can also use the masking sound signal 703 to process the audio signal 701. Specifically, the masking sound signal is phase-inverted to obtain an inverted sound signal 704. In one possible implementation, the phase range of the phase inversion is 90 degrees to 270 degrees, thereby ensuring that the inverted sound signal 704 has good compensation capability for the masking sound signal 703.

[0176] Figure 10 This is a schematic diagram illustrating the masking of audio signal 703 and inverted audio signal 704. In this embodiment, the inverted audio signal 704 is reduced in amplitude and then mixed with the audio signal 701 to obtain a mixed audio signal 705. Finally, this mixed audio signal 705 is provided to the receiver so that the receiver can play the mixed audio signal 705. The amplitude reduction processing can be implemented using an equalizer, gain control, or filtering. The specific implementation method of the amplitude reduction processing is not limited here.

[0177] Since the mixed audio signal 705 is formed by mixing the inverted audio signal 704 and the audio signal 701, it contains valid call content. Simultaneously, the component of the inverted audio signal 704 in the mixed audio signal 705, after being output, can also compensate for the masking audio signal 703 in the near field, thus canceling the masking audio signal played by the speaker from affecting the listener's call quality. Furthermore, the amplitude reduction processing also weakens the interference effect of the inverted audio signal 704 in the mixed audio signal 705 on the listener's ears.

[0178] like Figure 10 As shown, in one possible implementation, the masking sound signal 703 can also be enhanced to obtain an enhanced masking sound signal, which can then be provided to the speaker for transmission. Specifically, since the mid-to-high frequency range of the leaked sound is closely related to the listener's intelligibility of the leaked sound, an equalizer can be used to enhance the mid-to-high frequency range of the masking sound signal to strengthen the masking effect on the receiver's output audio signal in the mid-to-high frequency range. The enhancement processing can be achieved through equalizers, gain control, or filtering; the implementation method is not limited here.

[0179] See Figure 11In the figure, curve 1101 represents the masked sound signal before enhancement, and curve 1102 represents the masked sound signal after enhancement. By enhancing the masked sound signal 703, the masking effect of the masked sound signal on leaked sound is improved. It is understood that in practical applications, other methods can also be used to enhance the masked sound signal 703. In this embodiment, the specific method of enhancing the masked sound signal 703, the frequency domain of the enhancement, and the magnitude of the enhancement are not limited.

[0180] in the above and Figure 7 The previous section introduced a method for generating a masked sound signal by reversing a segment of sound after extracting it. The following describes another method for generating a masked sound signal using extracted sound segments.

[0181] In this embodiment, an audio signal 702 is extracted according to a preset frame length to obtain an extracted sound segment. The sound segment can then be interpolated to obtain supplemented sound information; or subsequent segments can be matched from a preset audio library to obtain supplemented sound information. Finally, a corresponding masking sound signal is generated based on the supplemented sound information.

[0182] As an example, the extracted audio segment can be preprocessed to extract feature parameters (such as bytes, pitch, etc.), and then the audio segment can be interpolated using the feature parameters and a pre-trained empirical model. After sampling, the supplemented audio information can be obtained.

[0183] As another example, an audio library is pre-built, in which each sound segment matches at least one other sound segment. After obtaining a truncated sound segment, any matching sound segment is retrieved from the pre-built audio library based on that segment; this matched sound segment is called the subsequent segment. Using the original sound segment and the subsequent segment, the supplementary sound information is obtained.

[0184] In this embodiment, a masking sound signal is generated by supplementing the sound information, which reduces the frequency of sound segment extraction and improves the generation efficiency of the masking sound signal. After the masking sound signal is generated, as an optional implementation, in order to improve the masking effect, the audio signal 701 can also be delayed so that the masking sound signal is aligned with the audio signal 701 for playback.

[0185] Testing showed that by implementing the sound masking method provided in the above embodiments, the intelligibility of leaked sound for listeners at 500mm was significantly reduced, decreasing from 90% before implementation to below 10%. In the far field, the intelligibility of individual words from leaked sound was less than 30%, and the intelligibility of sentences was less than 10%. Furthermore, the impact on ambient noise was less than 6dB, with no significant change in loudness compared to before implementation. Implementing this method had almost no impact on the audio intelligibility of near-field listeners; therefore, the sound masking method provided in this embodiment can effectively mask receiver leaks in the far field without altering the listener's call quality.

[0186] Based on the sound masking method provided in the foregoing embodiments, this application also provides a sound masking device. The following description is in conjunction with the accompanying drawings and embodiments.

[0187] See Figure 12 This figure is a schematic diagram of the structure of a sound masking device provided in an embodiment of this application. The sound masking device 120 shown in this figure can be applied to... Figure 1 and Figure 2 In the first terminal device 101 shown.

[0188] like Figure 12 As shown, the device 120 includes:

[0189] The judgment module 1201 is used to determine whether the terminal device uses the receiver as the output terminal of the audio signal;

[0190] The determination module 1203 is used to determine the masking sound signal based on the audio signal when the judgment result of the judgment module is yes;

[0191] The first control module 1202 is used to control the loudspeaker to emit a masking sound signal in order to mask the audio signal output by the receiver in the far field.

[0192] Because the masking sound signal is determined based on the audio signal, and the distance difference between the speaker and the receiver relative to the far field is small, the masking sound signal can effectively mask sound leakage from the receiver, reducing the intelligibility of the leaked sound to bystanders and preventing information leakage in the conversation. Furthermore, since the masking sound signal and the audio signal are output by the speaker and receiver respectively, when the listener listens to the audio signal with the receiver, the distance difference between the speaker and receiver relative to the listener's ear is relatively large. Therefore, the masking sound signal has minimal interference with the listener's audio signal reception and will not affect the listener's call quality.

[0193] In one possible implementation, the determining module 1203 is used to select or match a corresponding masking sound signal from a pre-generated audio library based on the audio signal.

[0194] In one possible implementation, the distance between the speaker and the listener's ear is greater than the distance between the receiver and the listener's ear.

[0195] In one possible implementation, module 1203 is used to generate a masking sound signal based on the audio signal.

[0196] like Figure 13 The schematic diagram of the signal generation module shown illustrates, in one possible implementation, that module 1203 specifically includes:

[0197] The spectrum analysis unit 12031 is used to perform spectrum analysis on audio signals and obtain the spectrum response;

[0198] The first generation unit 12032 is used to generate a masking sound signal based on the spectral response.

[0199] Understandably, since the masking sound signal is generated based on the spectral response of the audio signal, the masking sound and the masked sound have similar or identical spectral characteristics. Consequently, the masking sound signal can effectively mask the audio signal played by the receiver.

[0200] like Figure 14 The schematic diagram of the signal generation module shown illustrates that, in another possible implementation, module 1203 specifically includes:

[0201] The signal interception unit 12033 is used to intercept the audio signal according to a preset frame length to obtain the intercepted sound segment;

[0202] The signal inversion unit 12034 is used to invert the time domain of a sound segment to obtain an inverted sound.

[0203] The second generation unit 12035 is used to generate a corresponding masking sound signal based on the reversed sound.

[0204] By inverting audio segments and generating corresponding masking audio signals based on the inverted audio, the intelligibility of leaked audio can be reduced for far-field listeners. This ensures the security of private or confidential information within the call content.

[0205] like Figure 15 The schematic diagram of the signal generation module shown, in another possible implementation, determines that module 1203 specifically includes:

[0206] The signal interception unit 12033 is used to intercept the audio signal according to a preset frame length to obtain the intercepted sound segment;

[0207] The signal supplementation unit 12036 is used to interpolate the sound segment to obtain the supplemented sound signal; or to match subsequent segments from a preset audio library to obtain the supplemented sound signal.

[0208] The third generation unit 12037 is used to generate a corresponding masking sound signal based on the supplemented sound signal.

[0209] By supplementing the audio signal and reducing the frequency of audio signal interception, the efficiency of generating masked audio signals is improved. This avoids the listener's waiting time for the audio signal, thus enhancing the listener's call experience.

[0210] One possible implementation also includes:

[0211] The time length acquisition module is used to obtain the time length for generating the masked sound signal based on the audio signal;

[0212] The delay module is used to delay the audio signal according to the time length so that the sound signal output by the receiver matches the masked sound signal output by the speaker.

[0213] By delaying the audio signal output by the receiver, the masking audio signal is ensured to synchronously mask the audio signal output by the receiver, thereby improving the masking effect.

[0214] One possible implementation also includes:

[0215] The phase inversion processing module is used to invert the phase of the masking sound signal to obtain an inverted sound signal;

[0216] The mixing module is used to mix the inverted audio signal with the audio signal after EQ or amplitude processing to obtain a mixed audio signal.

[0217] The second control module is used to control the output of the mixed audio signal from the receiver.

[0218] The inverted audio signal obtained through the inverting processing module can compensate for the masking audio signal, thereby offsetting the interference of the masking audio signal on the near-field listener's ears to a certain extent and ensuring the listener's call quality.

[0219] In one possible implementation, the first control module 1202 specifically includes:

[0220] The first detection unit is used to detect sound signals from the surrounding environment;

[0221] The first judgment unit is used to determine whether the amplitude of the sound signal in the surrounding environment is lower than a first preset threshold.

[0222] The first control unit is used to emit a masking sound signal through a speaker when the judgment result of the first judgment unit is yes.

[0223] The system uses a threshold value below a first preset threshold as a trigger condition for transmitting a masking sound signal through the speaker, preventing bystanders in the environment from hearing leaked audio from the receiver due to excessive quietness. This avoids the leakage of private or confidential information during audio leaks.

[0224] In one possible implementation, the first control module 1202 specifically includes:

[0225] The second detection unit is used to detect whether there is a downlink audio signal;

[0226] The second judgment unit is used to determine whether the amplitude of the downlink audio signal is greater than a second preset threshold when the second detection unit detects a downlink audio signal.

[0227] The second control unit is used to emit a masking sound signal through a speaker when the first judgment unit determines that the result is yes.

[0228] The amplitude of the downlink audio signal is set to be greater than a preset threshold as the trigger condition for transmitting a masked sound signal through the speaker, so as to avoid unnecessary noise interference to listeners in the surrounding environment.

[0229] One possible implementation also includes:

[0230] The signal enhancement module is used to enhance the masked audio signal to obtain an enhanced masked audio signal. By enhancing the masked audio signal, the enhanced masked audio signal can more effectively mask the audio signal played by the receiver, reducing the intelligibility of leaked audio for bystanders.

[0231] Based on the sound masking method and sound masking device provided in the foregoing embodiments, this application also provides a terminal device. This terminal device can be... Figure 1 and Figure 2 The first terminal device 101 shown here, for its application scenarios, please refer to [link / reference]. Figure 1 and Figure 2 The details are omitted here. The structural implementation of the terminal device provided in the embodiments of this application is described below with reference to the examples and accompanying drawings.

[0232] See Figure 16 The figure is a schematic diagram of the structure of a terminal device provided in an embodiment of this application.

[0233] like Figure 16 As shown, the terminal device 160 includes: a receiver 1601, a speaker 1602, and a processor 1603.

[0234] The processor 1603 is used to determine the masking sound signal based on the audio signal when the receiver 1601 outputs an audio signal.

[0235] Speaker 1602 is used to transmit a masking sound signal to mask the audio signal output by receiver 1011 in the far field. Speaker 1602 can output the masking sound signal under the control of processor 1603.

[0236] Since the masking sound signal is determined based on the audio signal, and the distance difference between the speaker 1602 and the receiver 1601 relative to the far field is small, the masking sound signal can effectively mask the sound leakage of the receiver 1601, reducing the intelligibility of the leaked sound to the listener and preventing information leakage in the conversation. Furthermore, since the masking sound signal and the audio signal are output by the speaker 1602 and receiver 1601 respectively, when the listener receives the audio signal with the receiver 1601, the distance difference between the speaker 1602 and receiver 1601 relative to the listener's ear is relatively large. The masking sound signal emitted by the speaker 1602 is actually masked by the audio signal emitted by the receiver 1601 at the listener's ear. Therefore, the masking sound signal has minimal interference with the listener's audio signal reception and will not affect the listener's call quality.

[0237] Optionally, the processor 1603 is specifically configured to select or match a corresponding masking sound signal from a pre-generated audio library based on the audio signal when the receiver outputs an audio signal.

[0238] Optionally, the distance between the speaker 1602 and the listener's ear is greater than the distance between the receiver 1601 and the listener's ear.

[0239] Optionally, the processor 1603 is specifically used to perform spectral analysis on the audio signal to obtain the spectral response; and to generate a masked sound signal based on the spectral response.

[0240] Optionally, the processor 1603 is specifically used to extract audio signals according to a preset frame length to obtain the extracted sound segments; to perform time-domain inversion on the sound segments to obtain inverted sound; and to generate a corresponding masking sound signal based on the inverted sound.

[0241] Optionally, the processor 1603 is specifically used to extract audio signals according to a preset frame length to obtain the extracted sound segments; interpolate the sound segments to obtain supplemented sound signals; or match subsequent segments from a preset audio library to obtain supplemented sound signals; and generate corresponding masking sound signals based on the supplemented sound signals.

[0242] Optionally, the processor 1603 is also configured to obtain the duration of the masking sound signal generated based on the audio signal; and to delay the audio signal based on the duration so that the audio signal output by the receiver 1601 is adapted to the masking sound signal output by the speaker 1602.

[0243] Optionally, the processor 1603 is also configured to perform phase inversion processing on the masking sound signal to obtain an inverted sound signal; mix the inverted sound signal and the audio signal to obtain a mixed sound signal; and control the receiver 1601 to output the mixed sound signal.

[0244] Optionally, the processor 1603 is specifically used to detect the sound signal of the surrounding environment, and when the amplitude of the sound signal of the surrounding environment is lower than a first preset threshold, it controls the speaker 1602 to output a masking sound signal.

[0245] Optionally, the processor 1603 is specifically used to control the speaker 1602 to output a masking sound signal when it determines that the amplitude of the downlink audio signal is greater than a second preset threshold.

[0246] Optionally, the processor 1603 is also used to enhance the masking sound signal to obtain an enhanced masking sound signal.

[0247] In the terminal device 160 provided in this application embodiment, the processor 1603 can be used to execute some or all of the steps in the foregoing method embodiments. For a description of the functions of the processor 1603 and the related technical effects of executing the method steps, please refer to the foregoing method embodiments and device embodiments; further details will not be repeated here.

[0248] exist Figure 16 The terminal device 160 shown only illustrates the parts relevant to the embodiments of this application. For specific technical details not disclosed, please refer to the method section of the embodiments of this application. The terminal device 160 can be any terminal device, including mobile phones, tablets, personal digital assistants (PDAs), point-of-sale (POS) terminals, in-vehicle computers, etc. The following description and explanation use a mobile phone as an example to illustrate the terminal device provided in the embodiments of this application.

[0249] Figure 17 This diagram shows a partial structural block diagram of a mobile phone related to the terminal device provided in the embodiments of this application. (Reference) Figure 17The mobile phone 170 includes: a radio frequency (RF) circuit 1710, a memory 1720, an input unit 1730, a display unit 1740, a sensor 1750, an audio circuit 1760, a wireless fidelity (WiFi) module 1770, and a processor 1780 (this processor 1780 can implement...). Figure 16 The processor 1603 shown herein (its functions), as well as components such as the power supply battery, will be understood by those skilled in the art. Figure 17 The mobile phone structure shown does not constitute a limitation on the mobile phone and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0250] The following is combined Figure 17 A detailed introduction to each component of a mobile phone:

[0251] RF circuit 1710 can be used for receiving and transmitting signals during information transmission or calls. Specifically, it receives downlink information from the base station and processes it with processor 1780; additionally, it transmits uplink data to the base station. Typically, RF circuit 1710 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier (LNA), a duplexer, etc. Furthermore, RF circuit 1710 can also communicate wirelessly with networks and other devices. The aforementioned wireless communications may use any communication standard or protocol, including but not limited to Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.

[0252] The memory 1720 can be used to store software programs and modules. The processor 1780 executes various functional applications and data processing of the mobile phone 170 by running the software programs and modules stored in the memory 1720. The memory 1720 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone 170 (such as audio data, phonebook, etc.). In addition, the memory 1720 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0253] The input unit 1730 can be used to receive input numerical or character information, and to generate key signal inputs related to user settings and function control of the mobile phone 170. Specifically, the input unit 1730 may include a touch panel 1731 and other input devices 1732. The touch panel 1731, also known as a touch screen, can collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch panel 1731), and drive corresponding connected devices according to a pre-set program. Optionally, the touch panel 1731 may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, and sends it to the processor 1780, and can also receive and execute commands sent by the processor 1780. In addition, the touch panel 1731 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 1731, the input unit 1730 may also include other input devices 1732. Specifically, other input devices 1732 may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc.

[0254] The display unit 1740 can be used to display information input by the user or information provided to the user, as well as various menus of the mobile phone 170. The display unit 1740 may include a display panel 1741, which may optionally be configured as a Liquid Crystal Display (LCD), Organic Light-Emitting Diode (OLED), or similar display panel 1741. Further, a touch panel 1731 may cover the display panel 1741. When the touch panel 1731 detects a touch operation on or near it, it transmits the information to the processor 1780 to determine the type of touch event. Subsequently, the processor 1780 provides corresponding visual output on the display panel 1741 based on the type of touch event. Although in Figure 17 In this embodiment, the touch panel 1731 and the display panel 1741 are two separate components to realize the input and output functions of the mobile phone 170. However, in some embodiments, the touch panel 1731 and the display panel 1741 can be integrated to realize the input and output functions of the mobile phone 170.

[0255] The mobile phone 170 may also include at least one sensor 1750, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. The ambient light sensor can adjust the brightness of the display panel 1741 according to the ambient light level, and the proximity sensor can turn off the display panel 1741 and / or backlight when the mobile phone 170 is moved to the ear. As a type of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in various directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity, and can be used for applications that recognize the posture of the mobile phone 170 (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc. Other sensors that may be configured on the mobile phone 170, such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, will not be described in detail here.

[0256] Audio circuitry 1760, speaker 1761, microphone 1762, and receiver 1763 provide an audio interface between the user and mobile phone 170. Audio circuitry 1760 converts received audio data into electrical signals and transmits them to speaker 1761 or receiver 1763, where they are converted into sound signals for output. On the other hand, microphone 1762 converts collected sound signals into electrical signals, which are received by audio circuitry 1760, converted into audio data, and then processed by processor 1780 before being transmitted via RF circuitry 1710 to, for example, another mobile phone 170, or output to memory 1720 for further processing.

[0257] WiFi is a short-range wireless transmission technology. The mobile phone 170, through its WiFi module 1770, can help users send and receive emails, browse web pages, and access streaming media, providing users with wireless broadband internet access. Although Figure 17 WiFi module 1770 is shown, but it is understood that it is not a necessary component of mobile phone 170 and can be omitted as needed without changing the essence of the invention.

[0258] The processor 1780 is the control center of the mobile phone 170. It connects various parts of the mobile phone 170 via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 1720, and by calling data stored in the memory 1720, it performs various functions and processes data of the mobile phone 170, thereby providing overall monitoring of the mobile phone 170. Optionally, the processor 1780 may include one or more processing units; preferably, the processor 1780 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the aforementioned modem processor may also not be integrated into the processor 1780.

[0259] The mobile phone 170 also includes a power supply (such as a battery) that supplies power to various components. Preferably, the power supply can be logically connected to the processor 1780 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system.

[0260] Although not shown, the mobile phone 170 may also include a camera, Bluetooth module, etc., which will not be described in detail here.

[0261] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0262] The above description is merely a preferred embodiment of this application and is not intended to limit the application in any way. Although this application has disclosed preferred embodiments above, it is not intended to limit the application. Any person skilled in the art can make many possible variations and modifications to the technical solutions of this application using the methods and techniques disclosed above, or modify them into equivalent embodiments with equivalent changes, without departing from the scope of the technical solutions of this application. Therefore, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of this application without departing from the content of the technical solutions of this application shall still fall within the protection scope of the technical solutions of this application.

Claims

1. A method of masking sound, characterized by, The invention is applied to a terminal device, which includes a receiver and a speaker; the terminal device is an in-vehicle computer, and the distance between the speaker and the listener's ear is greater than the distance between the receiver and the listener's ear. The distance between the listener and the receiver is similar to the distance between the listener and the speaker; The method includes: When the terminal device outputs an audio signal through the receiver, it acquires the time-domain and / or frequency-domain characteristics of the audio signal, and determines or generates a masking sound signal matching the audio signal based on the time-domain and / or frequency-domain characteristics of the audio signal. Specifically, determining the masking sound signal based on the audio signal includes: selecting or matching a corresponding masking sound signal from a pre-generated audio library based on the audio signal; or, performing spectral analysis on the audio signal to obtain a spectral response, and generating the masking sound signal based on the spectral response; or... The audio signal is truncated according to a preset frame length to obtain a truncated sound segment. The sound segment is then time-reversed to obtain an inverted sound. The inverted sound is then directly concatenated to generate the masking sound signal, or concatenated after applying a window function to generate the masking sound signal. Alternatively, the audio signal is truncated according to a preset frame length to obtain a truncated sound segment. The sound segment is then interpolated to obtain a supplemented sound signal, or subsequent segments are matched from a preset audio library to obtain a supplemented sound signal. A corresponding masking sound signal is then generated based on the supplemented sound signal. The speaker emits the masking sound signal, which is used to mask the audio signal output by the receiver to the far field; the masking sound signal is sent when the amplitude of the sound signal in the environment around the terminal device is lower than a first threshold. The masking sound signal is phase-inverted to obtain an inverted sound signal; The amplitude of the inverted sound signal is reduced and then mixed with the audio signal to obtain a mixed sound signal. The receiver outputs the mixed audio signal.

2. The method of claim 1, wherein, Also includes: Obtain the time length during which the masking sound signal is generated based on the audio signal; The audio signal is delayed according to the specified time length so that the audio signal output by the receiver matches the masking sound signal output by the speaker.

3. The method according to any of claims 1-2, characterized in that, The process of emitting the masking sound signal through the speaker specifically includes: When a downlink audio signal is detected, if the amplitude of the downlink audio signal is greater than a second preset threshold, a masking sound signal is sent through the speaker.

4. A terminal device, characterized by comprising: include: receiver, speaker, and processor; The terminal device is an in-vehicle computer, and the distance between the speaker and the listener's ear is greater than the distance between the receiver and the listener's ear; The distance between the listener and the receiver is similar to the distance between the listener and the speaker; The processor is configured to, when an audio signal is output by the receiver, acquire the time-domain characteristics and / or frequency-domain characteristics of the audio signal, and determine or generate a masking sound signal matching the audio signal based on the time-domain characteristics and / or frequency-domain characteristics of the audio signal. The processor is specifically configured to, when the receiver outputs an audio signal, select or match a corresponding masking sound signal from a pre-generated audio library based on the audio signal; or, specifically, perform spectral analysis on the audio signal to obtain a spectral response, and generate the masking sound signal based on the spectral response; or, specifically, truncate the audio signal according to a preset frame length to obtain a truncated sound segment, invert the time domain of the sound segment to obtain an inverted sound, and generate a corresponding masking sound signal based on the inverted sound; or, specifically, truncate the audio signal according to a preset frame length to obtain a truncated sound segment, interpolate the sound segment to obtain a supplemented sound signal, or match subsequent segments from a preset audio library to obtain a supplemented sound signal, and generate a corresponding masking sound signal based on the supplemented sound signal. The loudspeaker is used to emit the masking sound signal to mask the audio signal output by the receiver to the far field; the masking sound signal is transmitted when the amplitude of the sound signal in the environment surrounding the terminal device is lower than a first threshold. The processor is further configured to perform phase inversion processing on the masking sound signal to obtain an inverted sound signal; perform amplitude reduction processing on the inverted sound signal and mix it with the audio signal to obtain a mixed sound signal; and control the receiver to output the mixed sound signal.

5. The terminal device according to claim 4, characterized in that, The processor is further configured to obtain the time length for generating the masking sound signal based on the audio signal; and to delay the audio signal based on the time length so that the audio signal output by the receiver is adapted to the masking sound signal output by the speaker.

6. The terminal device according to any one of claims 4-5, characterized in that, The processor is specifically configured to, when a downlink audio signal is detected, determine that the amplitude of the downlink audio signal is greater than a second preset threshold, and then transmit the masking sound signal through the speaker.

7. The terminal device according to any one of claims 4-5, characterized in that, The vehicle-mounted computer also includes: a call interface; The call interface displays an option to turn audio masking on or off for the user to choose from.