Sound masking method, system, device, storage medium and program product

By re-arranged processing of the ambient sound of the target object and superimposed pseudo-voice signals, a rearranged masked voice signal with similar spectral characteristics is generated, which solves the problem of poor effect of traditional sound masking methods and achieves more efficient information confidentiality protection and equipment flexibility.

CN119993188BActive Publication Date: 2025-07-22BYD CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510480181.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-22
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

The existing technology has poor effect in information confidentiality protection, especially the inability to effectively prevent information leakage through electronic devices, and traditional sound insulation equipment is large in size, complex in design, lacks flexibility, and cannot fully prevent secret theft.

Method used

By rearranging the ambient sound of the target object, a rearranged masked voice signal is generated so that the spectrum characteristics difference between it and the ambient sound is less than or equal to the preset threshold, and a sound masking process is performed based on the signal, including the superposition of the pseudo-voice signal and the interfering noise signal, ensuring masking effect and stability.

Benefits of technology

It improves the protection effect of information confidentiality, prevents target objects from extracting effective information from environmental sounds, reduces noise interference, and improves the flexibility of the equipment and anti-theft ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993188B_ABST
    Figure CN119993188B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a sound masking method, system, device, storage medium and program product. The method includes performing rearrangement processing on the external environmental sound in the sound masking space where the target object is located to obtain a rearranged masking speech signal, and the difference between the rearranged masking speech signal and the spectral characteristics of the external environmental sound is less than or equal to a preset threshold. Based on the rearranged masking speech signal, sound masking processing is performed on the target object. The method is used to achieve the effect of improving information security protection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of acoustic technologies, and particularly to an acoustic masking method, system, device, storage medium, and program product. Background Art

[0002] With the rapid development of science and technology, especially the rapid development and wide application of large-scale integrated circuits and mobile electronic devices, the risk of information leakage has been increasing continuously. For example, by using electronic devices such as smartphones and voice recorders, although work efficiency can be improved, the possibility of information leakage is also increased. Therefore, how to achieve information confidentiality protection in specific places has become increasingly important. Traditional meeting anti-leakage systems mainly use passive sound insulation means for sound protection. By installing sound-absorbing materials and designing wall sound insulation, the sound transmission path is cut off to achieve meeting anti-leakage. However, the current means of information confidentiality protection still have the problem of poor effect.

[0003] Therefore, how to improve the effect of information confidentiality protection is an urgent problem to be solved. Summary of the Invention

[0004] The acoustic masking method, system, device, storage medium, and program product provided by the embodiments of this application are used to improve the effect of information confidentiality protection.

[0005] In a first aspect, an embodiment of this application provides an acoustic masking method, including:

[0006] Rearranging the external environmental sound of the acoustic masking space where the target object is located to obtain a rearranged masking voice signal, where the difference between the rearranged masking voice signal and the spectral characteristics of the external environmental sound is less than or equal to a preset threshold;

[0007] Based on the rearranged masking voice signal, performing acoustic masking processing on the target object.

[0008] Optionally, the rearranging the environmental sound of the target object to obtain a rearranged masking voice signal includes:

[0009] Obtaining a beamforming signal corresponding to the environmental sound, where the beamforming signal is related to the sound source position and / or the number of sound sources of the environmental sound;

[0010] Rearranging the beamforming signal to obtain the rearranged masking voice signal.

[0011] Optionally, the rearranging the beamforming signal to obtain the rearranged masking voice signal includes:

[0012] Performing time slicing on the beamforming signal to obtain at least two initial signal segments;

[0013] Processing the initial signal segments by reversing the time order of the initial signal segments to obtain at least two target signal segments;

[0014] Combining the at least two target signal segments to obtain the rearranged masked speech signal.

[0015] Optionally, the performing acoustic masking processing on the target object based on the rearranged masked speech signal includes:

[0016] Performing acoustic masking processing on the target object based on the rearranged masked speech signal and a target signal; the target signal includes: a pseudo-speech signal and / or an interference noise signal.

[0017] Optionally, when the target signal includes a pseudo-speech signal, the method further includes:

[0018] Obtaining a pseudo-speech signal corresponding to the target speech scene based on the target speech scene where the target object is located.

[0019] Optionally, the obtaining a pseudo-speech signal corresponding to the target speech scene based on the target speech scene where the target object is located includes:

[0020] Obtaining a pseudo-speech signal corresponding to the target speech scene from a corpus based on the target speech scene where the target object is located.

[0021] Optionally, the method further includes:

[0022] Obtaining the target speech scene based on the input information of the user.

[0023] Optionally, the performing acoustic masking processing on the target object based on the rearranged masked speech signal and a target signal includes:

[0024] Obtaining a target masking signal based on a superimposed signal of the rearranged masked speech signal and the target signal;

[0025] Performing acoustic masking processing on the target object based on the target masking signal.

[0026] Optionally, the obtaining a target masking signal based on a superimposed signal of the rearranged masked speech signal and the target signal includes:

[0027] Determining masking intensity information of the target masking signal according to the external ambient sound;

[0028] Compressing the superimposed signal according to the masking intensity information to obtain the target masking signal.

[0029] Optionally, the masking intensity information includes a masking intensity upper limit, and determining the masking intensity information of the target masking signal according to the external ambient sound includes:

[0030] Determining the masking intensity upper limit corresponding to the frequency band according to the energy information of the external ambient sound in at least one frequency band;

[0031] Determining the masking intensity upper limit of the target masking signal according to the masking intensity upper limit corresponding to the frequency band.

[0032] Optionally, determining the masking intensity upper limit corresponding to the frequency band according to the energy information of the external ambient sound in at least one frequency band includes:

[0033] Determining the masking intensity upper limit of each frequency band according to the energy information of the external ambient sound in at least one frequency band and the passive sound insulation isolation degree of the target object.

[0034] Optionally, compressing the superimposed signal according to the masking intensity information to obtain the target masking signal includes:

[0035] If the sound pressure level of the superimposed signal is greater than the masking intensity upper limit, compressing the superimposed signal to obtain the target masking signal; the sound pressure level of the target masking signal is less than or equal to the masking intensity upper limit.

[0036] Optionally, the masking intensity information includes a masking state, and the masking state is used to determine the content included in the target masking signal and / or adjust the masking intensity upper limit;

[0037] Determining the masking intensity of the target masking signal according to the external ambient sound includes:

[0038] Determining whether there is a sound from a sound source belonging to the target type according to the spectrogram feature of the external ambient sound;

[0039] If so, determining that the masking state of the target masking signal is the first masking state;

[0040] If not, determining that the masking state of the target masking signal is the second masking state, the content of the target masking signal corresponding to the second masking state is different from that corresponding to the first masking state, and / or the masking intensity upper limit corresponding to the second masking state is less than the masking intensity upper limit corresponding to the first masking state.

[0041] In a second aspect, an embodiment of the present application provides a sound masking system, and the sound masking system includes: a microphone array, a speaker array, and a processing module;

[0042] The microphone array is located outside the sound masking space, and the speaker array and the target object to be sound masked are located inside the sound masking space;

[0043] The microphone array is configured to acquire the ambient sound of the target object, where the ambient sound is the ambient sound outside the housing;

[0044] The processing module is configured to obtain the rearranged masking speech signal based on the ambient sound by the method described in any item of the first aspect, and play the rearranged masking speech signal through the speaker array, so as to perform sound masking processing on the target object based on the rearranged masking speech signal.

[0045] In a third aspect, an embodiment of the present application provides a sound masking device, where the device includes:

[0046] A processing module rearranges the external ambient sound of the sound masking space where the target object is located to obtain a rearranged masking speech signal, and the difference between the rearranged masking speech signal and the spectral characteristics of the external ambient sound is less than or equal to a preset threshold;

[0047] A control module is configured to perform sound masking processing on the target object based on the rearranged masking speech signal.

[0048] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, where computer-executable instructions are stored in the computer-readable storage medium, and when the computer-executable instructions are executed by a processor, they are used to implement the first aspect and / or various possible implementation manners of the first aspect as described above.

[0049] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the first aspect and / or various possible implementation manners of the first aspect as described above.

[0050] The sound masking method, system, device, storage medium, and program product provided by the embodiments of the present application obtain a rearranged masking speech signal by rearranging the ambient sound of the target object, and perform sound masking processing on the target object based on the rearranged masking speech signal, so as to interfere with the semantics of the sound collected by the target object and prevent the target object from extracting valid information from the ambient sound, thereby improving the effect of information confidentiality protection. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] The drawings here are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0052] Figure 1 It is a schematic flowchart of a sound masking method provided by an embodiment of the present application;

[0053] Figure 2 It is a schematic flowchart of another sound masking method provided by an embodiment of the present application;

[0054] Figure 3 It is a schematic flowchart of yet another sound masking method provided by an embodiment of the present application;

[0055] Figure 4 It is a schematic flowchart of still another sound masking method provided by an embodiment of the present application;

[0056] Figure 5 It is a schematic flowchart of still another sound masking method provided by an embodiment of the present application;

[0057] Figure 6 It is a schematic flowchart of still another sound masking method provided by an embodiment of the present application;

[0058] Figure 7 It is a schematic structural diagram of a sound masking system provided by an embodiment of the present application;

[0059] Figure 8 It is a schematic structural diagram of another sound masking system provided by an embodiment of the present application;

[0060] Figure 9 It is a schematic structural diagram of a sound masking device provided by an embodiment of the present application.

[0061] Through the above-mentioned drawings, specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Specific Embodiments

[0062] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0063] At present, traditional conference anti-disclosure systems mainly use passive sound insulation means for sound protection. By installing sound-absorbing materials and designing sound-insulating walls, etc., the sound transmission path is cut off to achieve conference anti-disclosure. However, this method requires high sound insulation performance of the room, the equipment is bulky, the design is complex, lacks flexibility, and cannot protect against the theft of secrets from electronic devices inside the conference room. Some current emerging technologies install noise-generating devices in ventilation ducts, etc., and use noise to mask leaked voices. However, the noise emitted by such solutions lacks pertinence and the interference efficiency is low. At the same time, it only protects the outside of the conference room and cannot prevent participants from recording and other secret theft behaviors through electronic devices such as mobile phones.

[0064] In the prior art, the theft of secrets from electronic devices is mainly protected by the following methods:

[0065] Method 1: Protect the theft of secrets from electronic devices by setting up a storage box for storing communication devices. Among them, the storage box includes an aluminum alloy box body and a box cover. Transverse or longitudinal partition boards are arranged in the box body. Metal shielding nets are provided on the inner surfaces of the box body and the box cover and on the partition boards and are covered with soft pads. Sliders are installed at both ends of the partition boards, and slideways are arranged at the edges of the inner surface of the box body to adjust the distance between the partition boards. In this way, the electromagnetic shielding affects the communication signal of the communication device to improve the anti-secret theft protection effect for electronic devices.

[0066] However, this method does not consider the problem of sound masking, that is, it cannot prevent information leakage caused by recording through communication devices.

[0067] Method 2: Capture the ambient sound through a microphone subsystem, analyze the signal spectrum through a signal processing subsystem to form a specific directional masking sound, and emit the masking sound through a speaker subsystem.

[0068] However, although this method considers the problem of sound masking, it can only generate a specific directional masking sound according to the ambient sound. When playing the masking sound, if the masking sound is small, the masking effect will be poor; if the masking sound is large, it will cause noise interference to the people in the scene and affect the normal communication of the people in the scene.

[0069] Method 3: Use multi-microphone sound source localization technology. After locating the indoor coordinates of the speaker, according to the volume of the host microphone of the measured sound source and its coordinates, the coordinates of each vibration terminal (or speaker) and the sound energy attenuation law, calculate the sound energy at the coordinates of each vibration terminal, so as to more accurately set the energy of the interference signal, meet the signal-to-noise ratio required by the system, and obtain the best anti-eavesdropping effect with the least noise interference.

[0070] However, although this method meets the signal-to-noise ratio required by the system and reduces the noise interference caused by people in the scene, it determines the interference signal only based on the sound energy and cannot interfere with the semantics in the collected sound. That is, the interference signal cannot comprehensively and effectively destroy the eavesdropper's extraction of voice information, resulting in the problems of low interference effect and low interference efficiency of this method.

[0071] Therefore, how to protect the confidentiality of electronic devices and improve the effect of information confidentiality protection is an urgent problem to be solved.

[0072] In view of this, the present application provides a sound masking method. By collecting the ambient sound of a target object (such as a smart phone), rearranging the ambient sound of the target object, a rearranged masking speech signal capable of disrupting the semantics in the ambient sound is obtained, and the difference in spectral characteristics between the rearranged masking speech signal and the ambient sound is small. By playing the rearranged masking speech signal, sound masking processing is performed on the target object to interfere with the semantics of the sound collected by the target object, preventing the target object from extracting effective information from the ambient sound, thereby improving the effect of information confidentiality protection.

[0073] The technical solution of the present application and how the technical solution of the present application solves the above technical problems will be described in detail below through specific embodiments. These specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the drawings.

[0074] Figure 1 It is a schematic flowchart of a sound masking method provided by an embodiment of the present application. As Figure 1 shown, the method includes:

[0075] S101. Rearrange the external ambient sound of the sound masking space where the target object is located to obtain a rearranged masking speech signal.

[0076] Among them, the difference in spectral characteristics between the rearranged masking speech signal and the ambient sound is less than or equal to a preset threshold. Since the rearranged masking speech signal is directly obtained by rearranging the ambient sound, it has spectral characteristics similar to those of the ambient sound, that is, the difference in spectral characteristics between the rearranged masking speech signal and the ambient sound is less than or equal to the preset threshold.

[0077] Since the human auditory system has a certain adaptability to the surrounding ambient sound. When there is a sound with a certain specific spectral characteristic continuously in the environment, the auditory sense will gradually adapt to this sound pattern, making it less likely to attract special attention to a certain extent. Therefore, if the difference in spectral characteristics between the rearranged masking speech signal and the ambient sound is small, the rearranged masking speech signal is more likely to blend into the existing ambient sound background and reduce the impact on the people in the environment.

[0078] In addition, the spectral characteristics of the rearranged masked speech signal and the ambient sound have small differences, indicating that the rearranged masked speech signal can match the ambient sound well in each frequency component, so that it can more comprehensively cover each frequency range in the ambient sound and effectively mask various sounds in the environment.

[0079] The above-mentioned target object can be, for example, an electronic device with a recording function such as a mobile phone, a tablet computer, a smart wearable device, a voice recorder, etc., and there is a risk of information leakage caused by collecting ambient sound for the target object. The sound masking space can be, for example, a specific location where the target object is stored. For example, the target object is stored in a specific enclosed structure or a semi-enclosed structure (for example, it can be a box that can perform sound masking processing on the target object).

[0080] The external ambient sound of the sound masking space is the sound existing in the environment near the sound masking space, and this environment is an environment that requires information confidentiality protection. For example, it can be a scenario where procurement negotiations, solution decision-making, product development and other affairs are being carried out (such as a meeting room, or a specific room, etc.).

[0081] In this step, the external ambient sound can be collected first, and then the speech signal corresponding to the external ambient sound is rearranged to disrupt the speech signal corresponding to the external ambient sound, so that the semantic order of the speech signal changes, and a rearranged masked speech signal with chaotic semantics is generated.

[0082] Specifically, the external ambient sound can be divided into multiple speech segments, and then the rearrangement process is realized by rearranging these speech segments. Or, according to the time corresponding to the external ambient sound, the time sequence can be disrupted so that the speech signal corresponding to the external ambient sound is rearranged in time sequence to realize the rearrangement process. Or, the external ambient sound can be divided into multiple speech segments, and then after each speech segment is rearranged in time sequence, the multiple speech segments rearranged in time sequence are recombined in a specific order to realize the rearrangement process, etc.

[0083] S102. Perform sound masking processing on the target object based on the rearranged masked speech signal.

[0084] After obtaining the rearranged masked signal, the rearranged masked speech signal can be output so that the target object collects the rearranged masked speech signal when collecting sound, thereby realizing the function of masking the ambient sound with the rearranged masked speech signal. For example, the rearranged masked speech signal can be played to a specific location where the target object is stored through a speaker, a speaker array, etc., so that even if the target object turns on the sound collection function, it will not be able to extract the effective information in the ambient sound from the collected sound due to collecting the rearranged masked speech signal, thereby realizing the sound masking processing of the target object.

[0085] The method provided by the embodiments of the present application obtains a rearranged masked speech signal by rearranging the ambient sound of the target object, and performs acoustic masking processing on the target object based on the rearranged masked speech signal to interfere with the semantics of the sound collected by the target object, preventing the target object from extracting effective information from the ambient sound, thereby improving the effect of information confidentiality protection.

[0086] Next, a detailed introduction will be given on how to specifically rearrange the ambient sound of the target object in step S101 to obtain a rearranged masked speech signal. Figure 2 It is a schematic flowchart of another acoustic masking method provided by the embodiments of the present application. As Figure 2 shown, step S101 may specifically include:

[0087] S201. Obtain a beamforming signal corresponding to the ambient sound.

[0088] Among them, the beamforming signal is related to the sound source position and / or the number of sound sources of the ambient sound.

[0089] In this step, according to the ambient sound collected by a sensor (such as a microphone), the beamforming signal corresponding to the ambient sound can be calculated by using a beamforming algorithm. For how to specifically calculate the beamforming signal corresponding to the ambient sound through the beamforming algorithm, reference can be made to the prior art and will not be elaborated here.

[0090] Specifically, the arrival direction of the ambient sound can be estimated through a beamforming algorithm to obtain the sound source position, the number of sound sources, etc. of the ambient sound. For example, if the ambient sound is the sound from multiple directions collected by a microphone array arranged in a certain form, then multiple microphone signal inputs can be obtained based on the microphone array. According to the signal inputs of the microphone array, through the beamforming algorithm, the arrival direction of the sound source of the ambient sound in space can be estimated to obtain the number of sound sources and the sound source position in the environment. Then, based on each sound source, a beamforming signal corresponding to each sound source is generated, and the beamforming signal points to the arrival direction of the corresponding sound source.

[0091] S202. Rearrange the beamforming signal to obtain a rearranged masked speech signal.

[0092] A possible implementation manner is to divide each beamforming signal into multiple signal segments, and then obtain a rearranged masked speech signal by rearranging and combining these signal segments. For example, each signal segment can be numbered, and then after scrambling the number order (such as scrambling according to a preset rule or randomly scrambling, etc.), the multiple signal segments are recombined into a complete speech signal according to the scrambled number order, that is, the rearranged masked speech signal. Or, based on the time stamps of each signal segment, the time order can be scrambled for rearrangement.

[0093] In another possible implementation, each beamforming signal can be divided into multiple signal segments. For the time interval in which each signal segment is located, the time sequence within each signal segment is scrambled, and then the multiple signal segments with scrambled time sequences are rearranged and combined to obtain a rearranged masked speech signal.

[0094] Exemplarily, taking the rearrangement and combination of multiple signal segments with scrambled time sequences to obtain a rearranged masked speech signal as an example, this step can be implemented through the following sub-steps:

[0095] S2021: Perform time slicing on the beamforming signal to obtain at least two initial signal segments.

[0096] Among them, the beamforming signal can be time-sliced according to a preset time interval. The preset time interval can be set according to actual needs. For example, it can be 1 second, 0.5 second, 2 seconds, etc. This application does not limit this. Optionally, the beamforming signal can be divided into multiple initial signal segments with the same time slice duration using the same time interval, or the beamforming signal can be time-sliced using different time intervals.

[0097] Based on this preset time interval, starting from the starting moment of the beamforming signal, every time a preset time interval passes, the beamforming signal is time-sliced once, and finally at least two initial signal segments are generated.

[0098] S2022: Process the initial signal segments based on reversing the time order of the initial signal segments to obtain at least two target signal segments.

[0099] In this step, each initial signal segment can be reversed in time sequence. For example, if the original frequency corresponding to the first initial signal segment is from 0 second to 1 second, then based on reversing the time order, the frequency can be adjusted so that the frequency becomes the corresponding frequency from 1 second to 0 second.

[0100] In this way, all or part of the initial signal segments are processed in the same way to obtain at least two target signal segments. The frequency of the target signal segment is opposite to the frequency of the corresponding initial signal segment in time sequence (i.e., based on time sequence flipping).

[0101] S2023: Combine at least two target signal segments to obtain a rearranged masked speech signal.

[0102] Optionally, at least two target signal segments can be randomly combined to generate a complete signal, and this complete signal is the rearranged masked speech signal. Or, at least two target signal segments can be combined according to a preset rule to generate a complete signal, and this complete signal is the rearranged masked speech signal.

[0103] If based on a preset rule combination, for example, there are 4 target signal segments, namely target signal segment 1, target signal segment 2, target signal segment 3, and target signal segment 4, and the preset rule is 4-3-1-2, then in the rearranged masked speech signal obtained after splicing, it successively includes target signal segment 4, target signal segment 3, target signal segment 1, and target signal segment 2.

[0104] Optionally, the preset rule combination can also be the original order of the initial signal segments after beamforming signal segmentation, that is, in the rearranged masked speech signal obtained after splicing, it successively includes target signal segment 1, target signal segment 2, target signal segment 3, and target signal segment 4.

[0105] The method provided by the embodiments of this application obtains a beamforming signal corresponding to ambient sound, performs rearrangement processing on the beamforming signal to obtain a rearranged masked speech signal, so as to interfere with the semantics of the sound collected by the target object through the rearranged masked speech signal with scrambled semantics, preventing the target object from extracting valid information from the ambient sound, thereby improving the effect of information confidentiality protection.

[0106] Optionally, acoustic masking processing can be performed on the target object based on the rearranged masked speech signal and the target signal. The target signal includes: a pseudo-speech signal and / or an interference noise signal. The pseudo-speech signal is a signal that simulates speech according to the target corpus. The interference noise signal can be, for example, broadband interference noise. Exemplarily, white noise, pink noise, or babble noise, etc. can be selected as the interference noise signal. Since the broadband interference noise has the characteristics of a wide interference frequency band and stable interference effect, it can further improve the stability and masking effect of the acoustic masking processing.

[0107] Next, an introduction will be made with the target signal including a pseudo-speech signal.

[0108] A possible implementation manner is to directly obtain a preset pseudo-speech signal, and then perform acoustic masking processing on the target object based on the rearranged masked speech signal and the pseudo-speech signal.

[0109] Another possible implementation manner is to obtain a pseudo-speech signal corresponding to the target speech scenario based on the target speech scenario where the target object is located, and then perform acoustic masking processing on the target object based on the rearranged masked speech signal and the pseudo-speech signal. The target speech scenario can include, for example, procurement negotiation, solution decision-making, product development, etc. scenarios, or can include any other scenarios, etc. In the preset corpus, different target speech scenarios correspond to different pseudo-speech signals or corpus data.

[0110] When obtaining a pseudo-speech signal corresponding to a target speech scenario based on the target speech scenario where the target object is located, the target speech scenario can be determined, for example, according to the scenario corresponding to the ambient sound. For example, the target speech scenario can be obtained based on the user's input information, or the target speech scenario can be identified and determined according to the external ambient sound. The user can determine the target speech scenario, for example, by inputting the target speech scenario or selecting the target speech scenario from a preset list of candidate speech scenarios. In response to the user's operation, the target speech scenario selected or input by the user can be obtained.

[0111] Optionally, the pseudo-speech signal can be a pseudo-speech signal obtained from a corpus corresponding to the target speech scenario. The corpus is a pre-constructed collection containing a large number of pseudo-speech signals or corpus data, and these pseudo-speech signals or corpus data can cover various fields and various topics, such as news, novels, conversation records, etc. The pseudo-speech signal corresponding to the target speech scenario can be directly obtained from the corpus, or the corpus data corresponding to the target speech scenario can be obtained from the corpus, and then the corresponding pseudo-speech signal can be generated based on the corpus data.

[0112] Among them, the corpus data corresponding to the target scenario can be randomly selected from the corpus according to the preset number of corpus. Or, according to the preset number of target corpus, the first corpus number of corpus data corresponding to the target scenario can be selected from the preset corpus, etc.

[0113] In this implementation manner, if the corpus data is obtained from the corpus and the pseudo-speech signal is generated according to the corpus data, the text-to-speech technology can be used to simulate the human pronunciation method according to the semantic and grammatical information of the text information in the target corpus to generate the pseudo-speech signal. Optionally, during the generation process, parameters such as the timbre, intonation, and speech rate of the speech can also be adjusted to increase the diversity and authenticity of the pseudo-speech signal.

[0114] The method provided in the embodiments of the present application obtains a pseudo-speech signal corresponding to the target speech scenario based on the target speech scenario where the target object is located, and then performs acoustic masking processing on the target object based on the rearranged masked speech signal and the pseudo-speech signal, so as to further increase the pseudo-speech signal that can confuse the effective information included in the external ambient sound on the basis of outputting the rearranged masked speech signal, improve the difficulty for the target object to extract effective information from the external ambient sound, and thus enhance the effect of information confidentiality protection.

[0115] Figure 3 For the flowchart of another acoustic masking method provided in the embodiments of the present application, as Figure 3 shown, the method may specifically include:

[0116] S301. Obtain a target masking signal based on the superimposed signal of the rearranged masked speech signal and the target signal.

[0117] In this step, the rearranged masked speech signal and the target signal can be added to achieve the superposition based on the rearranged masked speech signal and the target signal, and obtain the target masking signal. Specifically, the time stamps of the rearranged masked speech signal and the target signal can be directly aligned for addition to achieve superposition, or they can be added without aligning the time stamps (but it is necessary to ensure that there is a superimposed area between the rearranged masked speech signal and the target signal).

[0118] Optionally, a target masking signal can be obtained based on the superimposed signal of the rearranged masked speech signal and the pseudo-speech signal; or a target masking signal can be obtained based on the superimposed signal of the rearranged masked speech signal and the interference noise signal; or a target masking signal can also be obtained based on the superimposed signal of the rearranged masked speech signal, the pseudo-speech signal, and the interference noise signal.

[0119] S302. Perform acoustic masking processing on the target object based on the target masking signal.

[0120] Output the target masking signal in the acoustic masking space where the target object is located, so that the target object can passively collect the target masking signal when collecting sound, and achieve the acoustic masking processing of the target object.

[0121] The method provided in the embodiments of the present application obtains a target masking signal including various signals for acoustic masking processing of the target object by superimposing at least one of the rearranged masked speech signal, the pseudo-speech signal, and the interference noise signal, and performs acoustic masking processing on the target object based on the target masking signal. This method can disrupt the semantics in the external environmental sound through the rearranged masked speech signal, further confuse the effective information included in the external environmental sound through the pseudo-speech signal, and improve the stability of the acoustic masking processing by adding a noise interference signal, thereby increasing the difficulty for the target object to extract effective information from the external environmental sound and enhancing the effect of information confidentiality protection.

[0122] Figure 4 It is a schematic flowchart of another acoustic masking method provided in the embodiments of the present application. As Figure 4 shown, the method may further include:

[0123] S401. Determine the masking intensity information of the target masking signal according to the external environmental sound.

[0124] Among them, the masking intensity information is used to determine or limit the masking intensity of the target masking signal. When the masking intensity of the target masking signal is stronger, the masking effect of the target masking signal is better, but the interference on the external environment of the acoustic masking space is stronger; when the masking intensity of the target masking signal is weaker, the masking effect of the target masking signal is worse, but the interference on the external environment of the acoustic masking space is weaker.

[0125] Therefore, the masking intensity information of the target masking signal can be determined according to the external environmental sound to ensure that, while ensuring the masking effect of the target masking signal, the interference on the external environment is reduced as much as possible.

[0126] Specifically, for example, the masking intensity information of the target masking signal related to the external environmental sound can be determined based on one or more of factors such as the energy information of the external environmental sound, the sound source type (for example, different masking intensity information is used for different types of sound sources), etc. Optionally, the passive sound insulation isolation degree of the target object can also be further considered to determine the masking intensity information of the target masking signal.

[0127] Optionally, the masking intensity information can include, for example, a masking intensity upper limit, and / or, a masking state.

[0128] Among them, the masking intensity upper limit is used to limit the sound pressure level upper limit of the target masking signal, and this masking intensity upper limit can be determined according to the energy information of the external environmental sound and / or the sound source type to ensure that the sound pressure level of the target masking signal is less than or equal to the masking intensity upper limit, so that the target masking signal has less interference on the external environment while ensuring the masking effect.

[0129] Among them, the masking state is used to configure various strategies of the target masking signal to adaptively adjust the strategy of the target masking signal according to the change of the external environmental sound, so as to realize that under different masking states, the content of the superimposed signal included in the target masking signal is adjusted with different masking intensity upper limits, and / or, so as to ensure that, while ensuring the masking effect of the target masking signal, the interference on the external environment is reduced as much as possible.

[0130] S402. Compress the superimposed signal according to the masking intensity information to obtain the target masking signal.

[0131] Compress the superimposed signal generated in the foregoing embodiment according to the masking intensity upper limit included in the masking intensity information, so that the sound pressure level of the superimposed signal is less than or equal to the masking intensity upper limit, thereby generating a target masking signal that reduces the interference on the external environment as much as possible while ensuring the masking effect.

[0132] Specifically, if the sound pressure level of the superimposed signal is greater than the upper limit of the masking intensity, the superimposed signal can be compressed to obtain the target masking signal. The sound pressure level of the target masking signal is less than or equal to the upper limit of the masking intensity. For example, if the upper limit of the masking intensity is 50 dB and the sound pressure level of the superimposed signal is 60 dB, the superimposed signal can be compressed (for example, an audio compression algorithm can be used, such as amplitude-based compression. Detect the amplitude of the superimposed signal at each moment. When the sound pressure level exceeds the upper limit of the masking intensity, reduce the signal amplitude proportionally. For example, reduce it by a certain coefficient so that the sound pressure level drops within the upper limit of the masking intensity, and compress the superimposed signal to obtain the target masking signal), so that the sound pressure level of the superimposed signal is lower than 50 dB, and a target masking signal with a sound pressure level lower than 50 dB can be obtained, while ensuring the masking effect, minimizing the interference to the external environment as much as possible.

[0133] The method provided by the embodiments of the present application determines the masking intensity information of the target masking signal according to the external environmental sound, compresses the superimposed signal according to the masking intensity information to obtain the target masking signal, adjusts the masking intensity of the target masking signal according to the actual situation of the external environmental sound, and minimizes the interference to the external environment of the target masking signal while ensuring the masking effect of the target masking signal. Therefore, not only the masking effect of the sound masking process is improved, but also the noise interference to the external environment is reduced, enhancing the user experience.

[0134] Next, taking the masking intensity information including the upper limit of the masking intensity as an example, a detailed introduction will be given on how to determine the masking intensity information of the target masking signal according to the environmental sound. Figure 5 It is a schematic flowchart of another sound masking method provided by the embodiments of the present application, as Figure 5 shown. Specifically, step S401 can include:

[0135] S501. Determine the upper limit of the masking intensity corresponding to the frequency band according to the energy information of the external environmental sound in at least one frequency band.

[0136] Among them, the frequency band refers to a continuous frequency interval within a certain frequency range. The external environmental sound can be decomposed into a combination of different frequency components for the convenience of analyzing and processing the sound signal. The energy information of the external environmental sound reflects the intensity of the external environmental sound in a certain frequency band.

[0137] In this step, the spectrum analysis of the external environmental sound can be performed. For example, through the fast Fourier transform, the audio signal in the time domain is converted into a signal in the frequency domain to obtain the amplitude information of the environmental sound at each frequency. Then, according to the pre-divided frequency bands, the total energy within each frequency band is calculated. For example, the entire audible frequency range can be divided into multiple 1 / 3 octave frequency bands, and the signal power within each frequency band is calculated respectively.

[0138] According to the energy information of each frequency band, the upper limit of the masking intensity corresponding to the frequency band can be determined by combining the auditory characteristics of the human ear and acoustic principles. For example, a mapping relationship between the energy information and the upper limit of the masking intensity can be established. For example, for a frequency band with higher energy, the upper limit of its masking intensity can be appropriately increased; for a frequency band with lower energy, the upper limit of its masking intensity can be correspondingly decreased. It is also possible to determine the appropriate upper limit of the masking intensity by experimentally testing the masking effect of the human ear on sounds of different frequency bands in different environments.

[0139] S502. Determine the upper limit of the masking intensity of the target masking signal according to the upper limit of the masking intensity corresponding to the frequency band.

[0140] A possible implementation method is to assign corresponding weights to each frequency band, multiply the upper limit of the masking intensity corresponding to each frequency band by its corresponding weight, and then sum the results of all frequency bands to obtain the preliminary upper limit of the masking intensity of the target masking signal. Among them, since the contributions of different frequency bands to the overall masking effect may be different, the weights corresponding to each frequency band can be considered according to factors such as the auditory characteristics of the human ear, the characteristics of the ambient sound, and the specific application scenario. For example, in an environment dominated by medium-frequency sounds, the weight of the medium-frequency band can be appropriately increased.

[0141] Another possible implementation method is to determine the upper limit of the masking intensity of the target masking signal on each frequency band according to the upper limit of the masking intensity corresponding to each frequency band.

[0142] In the implementation method of determining the upper limit of the masking intensity of the target masking signal on each frequency band according to the upper limit of the masking intensity corresponding to each frequency band, the upper limit of the masking intensity corresponding to the frequency band can be determined according to the energy information of the external ambient sound on at least one frequency band. Then, according to the upper limit of the masking intensity corresponding to the frequency band, the upper limit of the masking intensity of the target masking signal can be determined.

[0143] Specifically, the external ambient sound can be transformed from the time domain to the frequency domain through spectrum analysis methods such as fast Fourier transform, the entire audible frequency range can be divided into multiple frequency bands, and the total energy within each frequency band can be calculated to obtain the energy information of the external ambient sound on at least one frequency band. Then, based on the auditory characteristics of the human ear, acoustic principles, and relevant research results or experimental tests, a mapping relationship between the energy information and the upper limit of the masking intensity is established, so as to determine the upper limit of the masking intensity corresponding to each frequency band. According to the upper limit of the masking intensity corresponding to each frequency band, the upper limit of the masking intensity of the target masking signal on each frequency band is determined, and according to the upper limit of the masking intensity of the target masking signal on each frequency band, the upper limit of the masking intensity of the target masking signal is determined.

[0144] In this implementation, optionally, the upper limit of the masking intensity for each frequency band can be further determined in consideration of the passive sound insulation isolation degree of the target object. The passive sound insulation isolation degree of the target object refers to the passive sound insulation effect of the sound masking space where the target object is located. For example, taking the sound masking space as a closed box, the passive sound insulation effect of the box is the passive sound insulation isolation degree of the target object.

[0145] Since the passive sound insulation isolation degree of the target object can weaken the intensity of the target masking signal received outside the sound masking space to a certain extent (for example, when the passive sound insulation isolation degree is 30 dB and the target masking signal is 50 dB, then outside the sound masking space, after the target masking signal is weakened by the passive sound insulation isolation degree, its intensity becomes 20 dB). Therefore, when determining the upper limit of the masking intensity for each frequency band, the passive sound insulation isolation degree of the target object can be considered, so as to appropriately increase the upper limit of the masking intensity for each frequency band while ensuring that the external environment is not disturbed.

[0146] Exemplarily, for example, when the interference to the external environment is within an acceptable range when the target masking signal is 40 dB and the passive sound insulation isolation degree of the target object is 30 dB, then the upper limit of the masking intensity of the target masking signal can be determined to be 30 + 40 = 70 dB, that is, when the sound pressure level of the target masking signal is below 70 dB, the interference to the external environment is within an acceptable range.

[0147] Next, taking the masking intensity information masking state, the masking state is used to determine the content included in the target masking signal, and / or adjust the upper limit of the masking intensity as an example, a detailed introduction will be given on how to determine the masking intensity information of the target masking signal according to the ambient sound. Figure 6 It is a schematic flowchart of another sound masking method provided by the embodiment of the present application, as Figure 6 shown, the foregoing step S401 may specifically include:

[0148] S601. Determine whether there is a sound from a sound source belonging to the target type according to the spectrogram characteristics of the external ambient sound.

[0149] Among them, the sound source of the target type can be, for example, the sound source of human voices included in the external ambient sound, or it can also be other specific sound sources, and the present application does not limit this. Since usually information leakage is caused by the leakage of semantics included in human voices, when there are human voices in the external ambient sound, a stronger masking effect is required to avoid information leakage; while when there are no human voices in the external ambient sound, the risk of information leakage is usually small, and sound masking can be performed with a weaker masking effect. Therefore, the subsequent embodiments will be described by taking the sound source of human voices as an example.

[0150] In this step, short-time Fourier transform can be used to analyze the spectrogram features of the external environmental sound. For example, short-time Fourier transform can be used to convert the time-domain signal of the external environmental sound into the frequency domain to obtain a spectrogram, which can show the energy distribution of the sound over time and frequency. Then, construct or use an existing template library containing typical spectrogram features of human voices, such as fundamental frequency range, formant distribution, etc. Compare the spectrogram features of the current external environmental sound with the features in the template library, and methods such as correlation analysis can be used to calculate the similarity. If the similarity exceeds a preset threshold, it is determined that there is a sound from a human voice; if the similarity is lower than the threshold, it is determined that there is no such sound.

[0151] Alternatively, machine learning algorithms can be used to train a large number of human voice and non-human voice samples, enabling the model to automatically learn and identify whether there is a human voice in the external environmental sound, etc.

[0152] If there is a sound from a sound source belonging to the target type, that is, there is a human voice, it indicates that a stronger masking effect is required to avoid information leakage, and step S602 is executed; if there is no sound from a sound source belonging to the target type, that is, there is no human voice, it indicates that a stronger masking effect is not required, and step S603 is executed.

[0153] S602. Determine that the masking state of the target masking signal is the first masking state.

[0154] Among them, the content of the target masking signal corresponding to the first masking state is different from the content of the target masking signal corresponding to the second masking state in subsequent S603, and / or, the upper limit of the masking intensity corresponding to the first masking state is greater than the upper limit of the masking intensity corresponding to the second masking state. That is, the first masking state is a stronger masking state to provide a stronger masking effect to avoid information leakage when there is a human voice; the second masking state is a weaker masking state to provide a weaker masking effect to reduce the interference and impact on the external environment when there is no human voice.

[0155] Specifically, if the content of the target masking signal corresponding to the first masking state is different from the content of the target masking signal corresponding to the second masking state in subsequent S603, the content of the target masking signal corresponding to the first masking state can be determined according to a preset rule. For example, it can be set that the content of the target masking signal corresponding to the first masking state includes rearranged masked speech signals, pseudo-speech signals, and interference noise signals; and it can be set that the content of the target masking signal corresponding to the second masking state includes interference noise signals, so that the first masking state can provide a stronger masking effect to avoid information leakage, and the second masking state can provide a weaker masking effect to reduce the interference and impact on the external environment.

[0156] If the upper limit of the masking intensity corresponding to the first masking state is different from the upper limit of the masking intensity corresponding to the second masking state in subsequent S603, a higher upper limit of the masking intensity can be configured for the first masking state to enhance the masking effect to avoid information leakage, and a lower upper limit of the masking intensity can be configured for the second masking state to reduce the masking effect, thereby reducing the interference and impact on the external environment.

[0157] S603. Determine that the masking state of the target masking signal is the second masking state.

[0158] The method provided by the embodiments of the present application determines whether there is a sound from a sound source belonging to the target type according to the spectrogram characteristics of the external environmental sound. If there is a sound from a sound source belonging to the target type, it indicates that a stronger masking effect is required to avoid information leakage, and the masking state of the target masking signal is determined to be the first masking state with a stronger masking effect to reduce the risk of information leakage; if there is no sound from a sound source belonging to the target type, it indicates that a strong masking effect is not required to avoid information leakage, then the masking state of the target masking signal can be determined to be the second masking state with a weaker masking effect to reduce the interference and impact of the target masking signal on the external environment.

[0159] Figure 7 It is a schematic structural diagram of a sound masking system provided by the embodiments of the present application. As Figure 7 shown, the sound masking system includes: a microphone array, a speaker array, and a processing module.

[0160] Among them, the microphone array is located outside the sound masking space, and the speaker array, as well as the target object to be sound masked, is located inside the sound masking space.

[0161] The microphone array is used to acquire the environmental sound of the target object, and the environmental sound is the environmental sound outside the housing (i.e., the external environmental sound mentioned in the foregoing method embodiments). The processing module is used to obtain the rearranged masking speech signal based on the external environmental sound by the method of any one of the foregoing method embodiments, and play the rearranged masking speech signal through the speaker array to perform sound masking processing on the target object based on the rearranged masking speech signal.

[0162] Optionally, the processing module is used to obtain the superimposed signal, the masking intensity of the target masking signal, etc. based on the external environmental sound by the method of any one of the foregoing method embodiments, compress the superimposed signal based on the masking intensity to generate the target masking signal, and play the target masking signal through the speaker array to perform sound masking processing on the target object based on the target masking signal.

[0163] In a possible implementation, the sound masking system may further include a housing, inside which a sound masking space is formed. The target object can be placed in the sound masking space, and the speaker array plays the target masking signal to perform sound masking processing on the target object in the sound masking space based on the target masking signal.

[0164] Optionally, the side wall of the housing may be filled with sound insulation material, which can be any existing sound absorption / sound insulation material at present, and the present application does not limit this. Through this sound insulation material, the passive sound insulation isolation degree of the target object can be determined (that is, the passive sound insulation isolation degree of the target object is determined by the sound insulation effect of the sound insulation material). Or, the housing can be designed as a physical structure with sound insulation effect, and at this time the passive sound insulation isolation degree of the target object is determined by the sound insulation effect of the housing.

[0165] For ease of understanding, by way of example, Figure 8 is a schematic structural diagram of another sound masking system provided by an embodiment of the present application. As Figure 8 shown, the system includes: a passive sound insulation module, an active masking module, an environment perception module, and a masking generation module.

[0166] Among them, the passive sound insulation module includes an acoustic absorption structure unit, a sound absorption and insulation unit, and a system support structure unit. The system support structure unit is used to support the deployment of the active masking module, the environment perception module, and the masking generation module on the passive sound insulation module. The acoustic absorption structure unit is used to provide a physical structure with sound insulation effect. The sound absorption and insulation unit is composed of sound insulation material.

[0167] The active masking module includes a masking area (i.e., the sound masking area where the target object is placed), a speaker array unit (i.e., the aforementioned speaker array), and a driving filter unit. Among them, the driving filter unit is used to filter the target masking signal, and the speaker array unit outputs the filtered target masking signal to perform sound masking processing on the target object in the sound masking space based on the target masking signal.

[0168] The environment perception module includes a microphone array unit (i.e., the aforementioned microphone array), a scene analysis unit, and a beamforming unit. Among them, the functions of the scene analysis unit and the beamforming unit can be provided by the aforementioned processing module. The processing module realizes the function of the scene analysis unit by analyzing and determining the masking state corresponding to the external environmental sound and / or the upper limit of the masking intensity. The processing module processes the external environmental sound collected by the microphone array unit to generate the aforementioned beamforming signal.

[0169] The masking production module is used to generate a rearranged masked speech signal or a superimposed signal, which may include a dynamic range control unit, a speech interference unit, a fixed masking unit, and a pseudo-speech generation unit. The functions of the dynamic range control unit, the speech interference unit, the fixed masking unit, and the pseudo-speech generation unit can all be provided by the aforementioned processing module. The processing module can implement the function of the fixed masking unit by generating an interference noise signal, the function of the speech interference unit by generating a rearranged masked speech signal, and the function of the pseudo-speech generation unit by generating a pseudo-speech signal.

[0170] The modules included in this system and the units in each module can implement the corresponding sound masking method in the foregoing method embodiments. The implementation principles and technical effects are similar and will not be elaborated here.

[0171] Figure 9 It is a schematic structural diagram of a sound masking device provided by an embodiment of this application. As Figure 4 shown, the sound masking device may include: a processing module 11 and a control module 12.

[0172] The processing module 11 is configured to rearrange the external environmental sound in the sound masking space where the target object is located to obtain a rearranged masked speech signal, and the difference between the spectral characteristics of the rearranged masked speech signal and the external environmental sound is less than or equal to a preset threshold.

[0173] The control module 12 is configured to perform sound masking processing on the target object based on the rearranged masked speech signal.

[0174] Optionally, the processing module 11 is specifically configured to obtain a beamforming signal corresponding to the environmental sound, and perform a rearrangement process on the beamforming signal to obtain a rearranged masked speech signal. Among them, the beamforming signal is related to the sound source position and / or the number of sound sources of the environmental sound.

[0175] Optionally, the processing module 11 is specifically configured to perform time slicing on the beamforming signal to obtain at least two initial signal segments, process the initial signal segments based on reversing the time order of the initial signal segments to obtain at least two target signal segments, and combine the at least two target signal segments to obtain a rearranged masked speech signal.

[0176] Optionally, the control module 12 is specifically configured to perform sound masking processing on the target object based on the rearranged masked speech signal and the target signal. The target signal includes: a pseudo-speech signal and / or an interference noise signal.

[0177] Optionally, when the target signal includes a pseudo-speech signal, the processing module 11 is further configured to obtain a pseudo-speech signal corresponding to the target speech scene based on the target speech scene where the target object is located.

[0178] Optionally, the processing module 11 is specifically configured to obtain a pseudo voice signal corresponding to the target voice scenario from the corpus based on the target voice scenario in which the target object is located.

[0179] Optionally, the processing module 11 is specifically configured to obtain the target voice scenario based on the input information of the user.

[0180] Optionally, the processing module 11 is specifically configured to obtain a target masking signal based on the superimposed signal of the rearranged masked voice signal and the target signal. The control module 12 is specifically configured to perform acoustic masking processing on the target object based on the target masking signal.

[0181] Optionally, the processing module 11 is specifically configured to determine the masking intensity information of the target masking signal according to the external environmental sound. Compress the superimposed signal according to the masking intensity information to obtain the target masking signal.

[0182] Optionally, when the masking intensity information includes a masking intensity upper limit, the processing module 11 is specifically configured to determine the masking intensity upper limit corresponding to the frequency band according to the energy information of the external environmental sound in at least one frequency band. Determine the masking intensity upper limit of the target masking signal according to the masking intensity upper limit corresponding to the frequency band.

[0183] Optionally, the processing module 11 is specifically configured to determine the masking intensity upper limit of each frequency band according to the energy information of the external environmental sound in at least one frequency band and the passive sound insulation isolation degree of the target object.

[0184] Optionally, if the sound pressure level of the superimposed signal is greater than the masking intensity upper limit, the processing module 11 is specifically configured to compress the superimposed signal to obtain the target masking signal. The sound pressure level of the target masking signal is less than or equal to the masking intensity upper limit.

[0185] Optionally, when the masking intensity information includes a masking state, the masking state is used to determine the content included in the target masking signal, and / or adjust the masking intensity upper limit, the processing module 11 is specifically configured to determine whether there is a sound from a sound source belonging to the target type according to the spectrogram feature of the external environmental sound. If so, determine that the masking state of the target masking signal is the first masking state. If not, determine that the masking state of the target masking signal is the second masking state. Wherein, the content of the target masking signal corresponding to the second masking state is different from that corresponding to the first masking state, and / or the masking intensity upper limit corresponding to the second masking state is less than the masking intensity upper limit corresponding to the first masking state.

[0186] The acoustic masking device provided by the embodiment of the present application can execute the acoustic masking method in the above method embodiment, and its implementation principle and technical effect are similar, and will not be described in detail here.

[0187] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the above-mentioned method.

[0188] The present application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-mentioned method.

[0189] The above-mentioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk or an optical disc. The readable storage medium can be any available medium accessible by a general-purpose or special-purpose computer.

[0190] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in a device.

[0191] The division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.

[0192] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0193] In addition, the functional units in various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0194] If a function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs and other various media that can store program codes.

[0195] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When this program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: ROMs, RAMs, magnetic disks, or optical discs and other various media that can store program codes.

[0196] Finally, it should be noted that: After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily think of other implementation manners of the present invention. The present invention is intended to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include common general knowledge or conventional technical means in the technical field not disclosed in the present invention. It is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.

Claims

1. A sound masking method, characterized in that, The method includes: Rearranging the external environmental sound in the sound masking space where the target object is located to obtain a rearranged masked speech signal, and the difference between the rearranged masked speech signal and the spectral characteristics of the external environmental sound is less than or equal to a preset threshold; Determining the masking intensity information of the target masking signal according to the external environmental sound; Compressing the superimposed signal of the rearranged masked speech signal and the target signal according to the masking intensity information to obtain the target masking signal; Performing sound masking processing on the target object based on the target masking signal; The masking intensity information includes a masking intensity upper limit, and determining the masking intensity information of the target masking signal according to the external environmental sound includes: Determining the masking intensity upper limit of each frequency band according to the energy information of the external environmental sound in at least one frequency band and the passive sound insulation isolation degree of the target object; By assigning corresponding weights to each frequency band, multiplying the masking intensity upper limit corresponding to each frequency band by its corresponding weight, and summing the results of all frequency bands, the masking intensity upper limit of the target masking signal is obtained.

2. The method according to claim 1, characterized in that, Rearranging the environmental sound of the target object to obtain a rearranged masked speech signal includes: Obtaining a beamforming signal corresponding to the environmental sound, and the beamforming signal is related to the sound source position and / or the number of sound sources of the environmental sound; Performing a rearrangement process on the beamforming signal to obtain the rearranged masked speech signal.

3. The method according to claim 2, wherein Performing a rearrangement process on the beamforming signal to obtain the rearranged masked speech signal includes: Performing time slicing on the beamforming signal to obtain at least two initial signal segments; Processing the initial signal segments based on reversing the time order of the initial signal segments to obtain at least two target signal segments; Combining the at least two target signal segments to obtain the rearranged masked speech signal.

4. The method according to claim 1, wherein The target signal includes: a pseudo-speech signal and / or an interference noise signal.

5. The method according to claim 4, wherein The target signal includes a pseudo-speech signal, and the method further includes: Obtaining a pseudo-speech signal corresponding to the target speech scene based on the target speech scene where the target object is located.

6. The method according to claim 5, characterized in that, Obtaining a pseudo-speech signal corresponding to the target speech scene based on the target speech scene where the target object is located includes: Obtaining a pseudo-speech signal corresponding to the target speech scene from a corpus based on the target speech scene where the target object is located.

7. The method according to claim 5, characterized in that, The method further includes: Obtaining the target speech scene based on the user's input information.

8. The method according to claim 1, wherein Compressing the superimposed signal of the rearranged masked speech signal and the target signal according to the masking intensity information to obtain the target masking signal includes: If the sound pressure level of the superimposed signal is greater than the masking intensity upper limit, compressing the superimposed signal to obtain the target masking signal; the sound pressure level of the target masking signal is less than or equal to the masking intensity upper limit.

9. The method according to claim 1, wherein The masking intensity information includes a masking state, and the masking state is used to determine the content included in the target masking signal and / or adjust the masking intensity upper limit; Determining the masking intensity of the target masking signal according to the external environmental sound includes: Determine whether there is a sound from a sound source belonging to the target type according to the spectrogram feature of the external environmental sound; If so, determine that the masking state of the target masking signal is the first masking state; If not, determine that the masking state of the target masking signal is the second masking state, and the content of the target masking signal corresponding to the second masking state is different from that corresponding to the first masking state, and / or the upper limit of the masking intensity corresponding to the second masking state is less than the upper limit of the masking intensity corresponding to the first masking state.

10. A sound masking system, characterized in that, The sound masking system includes: a microphone array, a speaker array, and a processing module; The microphone array is located outside the sound masking space, and the speaker array and the target object to be sound masked are located inside the sound masking space; The microphone array is used to acquire the environmental sound of the target object, and the environmental sound is the environmental sound outside the housing; The processing module is used to obtain the rearranged masking speech signal based on the environmental sound by the method according to any one of claims 1-9, and play the rearranged masking speech signal through the speaker array to perform sound masking processing on the target object based on the rearranged masking speech signal.

11. The system according to claim 10, wherein The sound masking system further includes a housing; the interior of the housing constitutes the sound masking space.

12. The system according to claim 11, characterized in that, The side wall of the housing is filled with sound insulation material.

13. A sound masking device, characterized in that, The device includes: A processing module that rearranges the external environmental sound of the sound masking space where the target object is located to obtain a rearranged masking speech signal, and the difference between the rearranged masking speech signal and the spectral feature of the external environmental sound is less than or equal to a preset threshold; A control module for performing sound masking processing on the target object based on the target masking signal; The processing module is specifically configured to determine the masking intensity information of the target masking signal according to the external environmental sound; and compress the superimposed signal of the rearranged masking speech signal and the target signal according to the masking intensity information to obtain the target masking signal; The processing module is further specifically configured to determine the upper limit of the masking intensity of each frequency band according to the energy information of the external environmental sound in at least one frequency band and the passive sound insulation isolation degree of the target object; multiply the upper limit of the masking intensity corresponding to each frequency band by its corresponding weight by assigning corresponding weights to each frequency band, and sum the results of all frequency bands to obtain the upper limit of the masking intensity of the target masking signal.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement the method according to any one of claims 1-9.

15. A computer program product, characterized in that, Including a computer program, which when executed by a processor implements the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Target voice privacy protection method and system

    CN102543066A

  • Method and device for improving voice privacy of mute cabin

    CN115910018A

  • Generating method for shielding signals used for protecting chinese speech privacy

    WO2016138605A1