Sound masking method, system and device, storage medium and program product

By rearranging the ambient sound of the target object, a rearranged masked voice signal similar to the external ambient sound spectrum characteristics are generated, and a sound masking process is performed on the target object, which solves the problem of poor information confidentiality protection in the prior art, and achieves more efficient information confidentiality protection.

CN119993188AActive Publication Date: 2025-05-13BYD CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510480181.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-05-13
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

The prior art has poor effect in information confidentiality protection and cannot effectively prevent information leakage caused by the collection of environmental sounds through electronic devices.

Method used

By rearranging the ambient sound of the target object, a rearranged masked voice signal is obtained. The spectrum characteristics difference between the signal and the external ambient sound is less than or equal to the preset threshold value, and the target object is sound masked based on the signal.

Benefits of technology

Effectively interfere with the semantics of the sound collected by the target object, preventing effective information from being extracted from the ambient sound, thereby improving the effect of information confidentiality protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993188A_ABST
    Figure CN119993188A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a sound masking method, system and device, a storage medium and a program product. The method comprises the steps that external environment sound of a sound masking space where a target object is located is rearranged, a rearranged masking voice signal is obtained, and the difference between the frequency spectrum characteristics of the rearranged masking voice signal and the frequency spectrum characteristics of the external environment sound is smaller than or equal to a preset threshold value. And performing sound masking processing on the target object based on the rearranged masking voice signal. The method is used for achieving the effect of improving information confidentiality protection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of acoustic technology, and in particular to a sound masking method, system, device, storage medium and program product. Background Art

[0002] With the rapid development of science and technology, especially the rapid development and widespread application of large-scale integrated circuits and mobile electronic devices, the risk of information leakage is increasing. For example, the use of electronic devices such as smart phones and voice recorders can improve work efficiency, but it also increases the possibility of information leakage. Therefore, how to achieve information confidentiality protection in specific places is becoming more and more important. Traditional conference anti-leakage systems mainly use passive sound insulation to protect against sound. By adding sound-absorbing materials and wall sound insulation design, the sound transmission path is cut off to achieve conference anti-leakage. However, the current means of information confidentiality protection still have the problem of poor effect.

[0003] Therefore, how to improve the effectiveness of information confidentiality protection is an urgent problem to be solved. Summary of the invention

[0004] The sound masking method, system, device, storage medium and program product provided in the embodiments of the present application are used to achieve the effect of improving information confidentiality protection.

[0005] In a first aspect, an embodiment of the present application provides a sound masking method, comprising:

[0006] Rearranging the external environmental sound of the sound masking space where the target object is located to obtain a rearranged masking speech signal, wherein the difference between the spectral characteristics of the rearranged masking speech signal and the external environmental sound is less than or equal to a preset threshold;

[0007] Based on the rearranged masking speech signal, sound masking processing is performed on the target object.

[0008] Optionally, the rearrangement of the ambient sound of the target object to obtain the rearranged masking speech signal includes:

[0009] Acquire a beamforming signal corresponding to the ambient sound, where the beamforming signal is related to a sound source position and / or the number of sound sources of the ambient sound;

[0010] The beamforming signal is rearranged to obtain the rearranged masked speech signal.

[0011] Optionally, the rearrangement processing of the beamforming signal to obtain the rearranged masked speech signal includes:

[0012] Time slicing the beamforming signal to obtain at least two initial signal segments;

[0013] Processing the initial signal segments based on reversing the time order of the initial signal segments to obtain at least two target signal segments;

[0014] The at least two target signal segments are combined to obtain the rearranged masked speech signal.

[0015] Optionally, performing sound masking processing on the target object based on the rearranged masking speech signal includes:

[0016] Based on the rearranged masking speech signal and the target signal, the target object is subjected to sound masking processing; the target signal includes: a pseudo speech signal and / or an interfering noise signal.

[0017] Optionally, the target signal includes a pseudo speech signal, and the method further includes:

[0018] Based on the target speech scene in which the target object is located, a pseudo speech signal corresponding to the target speech scene is acquired.

[0019] Optionally, the acquiring, based on a target speech scene in which the target object is located, a pseudo speech signal corresponding to the target speech scene includes:

[0020] Based on the target speech scene in which the target object is located, a pseudo speech signal corresponding to the target speech scene is acquired from a corpus.

[0021] Optionally, the method further includes:

[0022] Based on the user's input information, the target voice scene is acquired.

[0023] Optionally, performing sound masking processing on the target object based on the rearranged masking speech signal and the target signal includes:

[0024] Acquire a target masking signal based on a superposition signal of the rearranged masking speech signal and the target signal;

[0025] Based on the target masking signal, sound masking processing is performed on the target object.

[0026] Optionally, acquiring a target masking signal based on a superimposed signal of the rearranged masking speech signal and the target signal includes:

[0027] Determining masking strength information of the target masking signal according to the external environmental sound;

[0028] The superimposed signal is compressed according to the masking strength information to obtain the target masking signal.

[0029] Optionally, the masking strength information includes a masking strength upper limit, and determining the masking strength information of the target masking signal according to the external ambient sound includes:

[0030] Determining, according to energy information of the external ambient sound in at least one frequency band, an upper limit of masking intensity corresponding to the frequency band;

[0031] The upper limit of the masking strength of the target masking signal is determined according to the upper limit of the masking strength corresponding to the frequency band.

[0032] Optionally, determining the upper limit of the masking intensity corresponding to the frequency band according to the energy information of the external ambient sound in at least one frequency band includes:

[0033] The upper limit of the masking intensity of each frequency band is determined according to the energy information of the external ambient sound in at least one frequency band and the passive sound insulation isolation of the target object.

[0034] Optionally, compressing the superimposed signal according to the masking strength information to obtain the target masking signal includes:

[0035] If the sound pressure level of the superimposed signal is greater than the upper limit of the masking strength, the superimposed signal is compressed to obtain the target masking signal; the sound pressure level of the target masking signal is less than or equal to the upper limit of the masking strength.

[0036] Optionally, the masking strength information includes a masking state, and the masking state is used to determine the content included in the target masking signal and / or to adjust the masking strength upper limit;

[0037] The step of determining the masking strength of the target masking signal according to the external ambient sound comprises:

[0038] Determining whether there is sound from a sound source belonging to the target type according to the spectral features of the external environmental sound;

[0039] If so, determining that the masking state of the target masking signal is a first masking state;

[0040] If not present, the masking state of the target masking signal is determined to be the second masking state, the second masking state is different from the content of the target masking signal corresponding to the first masking state, and / or the masking strength upper limit corresponding to the second masking state is less than the masking strength upper limit corresponding to the first masking state.

[0041] In a second aspect, an embodiment of the present application provides a sound masking system, the sound masking system comprising: a microphone array, a speaker array, and a processing module;

[0042] The microphone array is located outside the sound masking space, and the speaker array and the target object to be sound masked are located inside the sound masking space;

[0043] The microphone array is used to acquire the ambient sound of the target object, where the ambient sound is the ambient sound outside the shell;

[0044] The processing module is used to obtain the rearranged masking speech signal based on the ambient sound by the method as described in any one of the first aspects, and play the rearranged masking speech signal through the speaker array, so as to perform sound masking processing on the target object based on the rearranged masking speech signal.

[0045] In a third aspect, an embodiment of the present application provides a sound masking device, the device comprising:

[0046] A processing module is used to rearrange the external environmental sound of the sound masking space where the target object is located to obtain a rearranged masking speech signal, wherein the difference between the spectral characteristics of the rearranged masking speech signal and the external environmental sound is less than or equal to a preset threshold;

[0047] A control module is used to perform sound masking processing on the target object based on the rearranged masking speech signal.

[0048] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the first aspect above and / or various possible implementations of the first aspect.

[0049] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the above first aspect and / or various possible implementation methods of the first aspect.

[0050] The sound masking method, system, device, storage medium and program product provided in the embodiments of the present application obtain a rearranged masking voice signal by rearranging the ambient sound of the target object, and perform sound masking processing on the target object based on the rearranged masking voice signal to interfere with the semantics of the sound collected by the target object, thereby preventing the target object from extracting effective information from the ambient sound, thereby improving the effect of information confidentiality protection. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0052] Figure 1 A schematic diagram of a flow chart of a sound masking method provided in an embodiment of the present application;

[0053] Figure 2 A schematic flow chart of another sound masking method provided in an embodiment of the present application;

[0054] Figure 3 A schematic flow chart of another sound masking method provided in an embodiment of the present application;

[0055] Figure 4 A schematic flow chart of another sound masking method provided in an embodiment of the present application;

[0056] Figure 5 A schematic flow chart of another sound masking method provided in an embodiment of the present application;

[0057] Figure 6 A schematic flow chart of another sound masking method provided in an embodiment of the present application;

[0058] Figure 7 A structural schematic diagram of a sound masking system provided in an embodiment of the present application;

[0059] Figure 8 A schematic diagram of the structure of another sound masking system provided in an embodiment of the present application;

[0060] Fig. 9 A schematic diagram of the structure of a sound masking device provided in an embodiment of the present application.

[0061] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION

[0062] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0063] At present, traditional conference anti-leakage systems mainly use passive sound insulation to protect against sound. By installing sound-absorbing materials and wall sound insulation design, the sound transmission path is cut off to achieve conference anti-leakage. However, this method has high requirements for room sound insulation performance, the equipment is large in size, complex in design, lacks flexibility, and cannot protect against electronic equipment in the conference room from eavesdropping. Some emerging technologies currently install noise-generating devices in ventilation ducts to mask leaked voices through noise. However, the noise emitted by such solutions lacks specificity and has low interference efficiency. At the same time, it only protects the outside of the conference room and cannot prevent participants from recording through mobile phones and other electronic devices.

[0064] In the prior art, electronic devices are protected from theft mainly through the following methods:

[0065] Method 1: Protect electronic devices from theft by setting up a storage box for storing communication equipment. The storage box includes an aluminum alloy box body and a box cover, and a horizontal or vertical partition is set in the box body. The inner surface of the box body and the box cover and the partition are provided with a metal shielding net and covered with a cushion. Sliders are installed at both ends of the partition, and a slideway is set on the edge of the inner surface of the box to adjust the distance between the partitions. In this way, electromagnetic shielding is used to affect the communication signal of the communication equipment to improve the protection effect against theft of electronic equipment.

[0066] However, this method does not take into account the problem of sound masking, that is, it cannot prevent information leakage caused by recording through communication equipment.

[0067] Method 2: The ambient sound is captured by the microphone subsystem, the signal spectrum is analyzed by the signal processing subsystem to form a specific directional masking sound, and the masking sound is emitted through the speaker subsystem.

[0068] However, although this method takes the problem of sound masking into consideration, it can only generate masking sounds in a specific direction according to the ambient sound. When playing the masking sound, if the masking sound is small, the masking effect will be poor; if the masking sound is large, it will cause noise interference to the people in the scene, affecting the normal communication of the people in the scene.

[0069] Method 3: Using multi-microphone sound source localization technology, after locating the indoor coordinates of the speaker, the sound energy at the coordinates of each vibration terminal is calculated based on the measured sound source host microphone volume and coordinates, the coordinates of each vibration terminal (or speaker), and the sound energy attenuation law, so as to more accurately set the energy of the interference signal, meet the signal-to-noise ratio required by the system, and obtain the best anti-eavesdropping effect with minimal noise interference.

[0070] However, although this method meets the signal-to-noise ratio required by the system and reduces the noise interference caused by people in the scene, it only determines the interference signal based on the sound energy and cannot interfere with the semantics of the collected sound. That is, the interference signal cannot fully and effectively destroy the eavesdropper's extraction of voice information, resulting in the problem of low interference effect and interference efficiency in this method.

[0071] Therefore, how to protect electronic devices from theft and improve the effectiveness of information confidentiality protection is an urgent problem to be solved.

[0072] In view of this, the present application provides a sound masking method, which collects the ambient sound of a target object (such as a smart phone), rearranges the ambient sound of the target object, and obtains a rearranged masking voice signal that can disrupt the semantics in the ambient sound, and the difference in the spectral characteristics of the rearranged masking voice signal and the ambient sound is small. By playing the rearranged masking voice signal, the target object is subjected to sound masking processing to interfere with the semantics of the sound collected by the target object, preventing the target object from extracting effective information from the ambient sound, thereby improving the effect of information confidentiality protection.

[0073] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below through specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0074] Figure 1 A flow chart of a sound masking method provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, the method includes:

[0075] S101. Rearrange the external environmental sound of the sound masking space where the target object is located to obtain a rearranged masking speech signal.

[0076] The difference between the spectral characteristics of the rearranged masking speech signal and the ambient sound is less than or equal to a preset threshold. Since the rearranged masking speech signal is directly rearranged according to the ambient sound, it has a spectral characteristic similar to that of the ambient sound, that is, the difference between the spectral characteristics of the rearranged masking speech signal and the ambient sound is less than or equal to a preset threshold.

[0077] Since the human auditory system has a certain degree of adaptability to the ambient sound, when a sound with a certain spectral characteristic persists in the environment, the hearing will gradually adapt to this sound pattern, making it less likely to attract special attention. Therefore, if the spectral characteristics of the rearranged masking speech signal and the ambient sound are slightly different, the rearranged masking speech signal will be more easily integrated into the existing ambient sound background, reducing the impact on people in the environment.

[0078] In addition, the spectral characteristics of the rearranged masking speech signal and the ambient sound are slightly different, indicating that the rearranged masking speech signal can match the ambient sound well in each frequency component, thereby more comprehensively covering each frequency range in the ambient sound and effectively masking various sounds in the environment.

[0079] The target object may be, for example, a mobile phone, tablet computer, smart wearable device, voice recorder or other electronic device with recording function, which may collect ambient sound and cause information leakage. The sound masking space may be, for example, a specific location where the target object is stored, for example, the target object is stored in a specific closed structure or semi-closed structure (for example, a box that can perform sound masking on the target object).

[0080] The external environmental sound of the sound masking space is the sound existing in the environment near the sound masking space, which is an environment where information confidentiality protection is required, for example, it may be a scene where procurement negotiations, program decisions, product development and other matters are being carried out (such as a conference room, or a specific room, etc.).

[0081] In this step, the external environmental sound may be collected first, and then the speech signal corresponding to the external environmental sound may be rearranged to disrupt the speech signal corresponding to the external environmental sound, so that the semantic order of the speech signal changes, thereby generating a semantically confused rearranged masked speech signal.

[0082] Specifically, the external environmental sound can be divided into multiple voice segments, and then the reordering process can be achieved by rearranging these voice segments. Alternatively, the time sequence can be disrupted according to the time corresponding to the external environmental sound, so that the voice signal corresponding to the external environmental sound is rearranged in time sequence to achieve reordering. Alternatively, the external environmental sound can be divided into multiple voice segments, and then each voice segment is rearranged in time sequence, and then the multiple voice segments rearranged in time sequence are recombined in a specific order to achieve reordering, etc.

[0083] S102: Perform sound masking processing on the target object based on the rearranged masking speech signal.

[0084] After obtaining the rearranged masking signal, the rearranged masking voice signal can be output so that the target object can collect the rearranged masking voice signal when collecting sound, thereby realizing the function of masking the ambient sound by the rearranged masking voice signal. For example, the rearranged masking voice signal can be played to a specific location where the target object is stored through a speaker, a speaker array, etc., so that even if the target object turns on the sound collection function, it will not be able to extract effective information in the ambient sound from the collected sound because the rearranged masking voice signal is collected, thereby realizing the sound masking processing of the target object.

[0085] The method provided in the embodiment of the present application obtains a rearranged masking voice signal by rearranging the ambient sound of the target object, and performs sound masking processing on the target object based on the rearranged masking voice signal to interfere with the semantics of the sound collected by the target object, thereby preventing the target object from extracting valid information from the ambient sound, thereby improving the effect of information confidentiality protection.

[0086] Next, how to rearrange the ambient sound of the target object to obtain the rearranged masking speech signal in the aforementioned step S101 is described in detail. Figure 2 A flow chart of another sound masking method provided in an embodiment of the present application is shown as follows: Figure 2 As shown, the aforementioned step S101 may specifically include:

[0087] S201: Obtain a beamforming signal corresponding to ambient sound.

[0088] The beamforming signal is related to the sound source position and / or the number of sound sources of the ambient sound.

[0089] In this step, a beamforming signal corresponding to the ambient sound can be calculated using a beamforming algorithm based on the ambient sound collected by a sensor (such as a microphone). How to calculate the beamforming signal corresponding to the ambient sound using a beamforming algorithm can refer to the prior art and will not be described in detail here.

[0090] Specifically, the direction of the incoming wave of the ambient sound can be estimated through the beamforming algorithm to obtain the location of the sound source of the ambient sound, the number of sound sources, etc. For example, the ambient sound is the sound from multiple directions collected by a microphone array arranged in a certain form, then the multi-channel microphone signal input can be obtained based on the microphone array, and the direction of the incoming wave of the sound source of the ambient sound in the space can be estimated based on the signal input of the microphone array through the beamforming algorithm, and the number and location of the sound sources in the environment can be obtained. Then, based on each sound source, a beamforming signal corresponding to each sound source is generated, and the beamforming signal points to the direction of the incoming wave of the corresponding sound source.

[0091] S202: Rearrange the beamformed signal to obtain a rearranged masked speech signal.

[0092] In a possible implementation, each beamforming signal may be divided into a plurality of signal segments, and then these signal segments may be rearranged and combined to obtain a rearranged masked speech signal. For example, each signal segment may be numbered, and then the numbering sequence may be disrupted (for example, according to a preset rule, or randomly), and then the plurality of signal segments may be recombined into a complete speech signal according to the disrupted numbering sequence, that is, the rearranged masked speech signal. Alternatively, the time sequence may be disrupted and rearranged based on the timestamp of each signal segment.

[0093] In another possible implementation, each beamforming signal may be divided into multiple signal segments. For the time interval of each signal segment, the time sequence within each signal segment is disrupted, and then the multiple signal segments with disrupted time sequence are rearranged and combined to obtain a rearranged masked speech signal.

[0094] Exemplarily, taking the example of rearranging and combining a plurality of signal segments after the time sequence is disrupted to obtain a rearranged masked speech signal, this step can be implemented by the following sub-steps:

[0095] S2021. Time slice the beamforming signal to obtain at least two initial signal segments.

[0096] The beamforming signal may be time sliced ​​according to a preset time interval, and the preset time interval may be set according to actual needs, for example, 1 second, 0.5 second, 2 seconds, etc., and the present application does not impose any limitation on this. Optionally, the beamforming signal may be divided into a plurality of initial signal segments with the same time slice duration at the same time interval, or the beamforming signal may be time sliced ​​at different time intervals.

[0097] Based on the preset time interval, starting from the start time of the beamforming signal, the beamforming signal is time sliced ​​once every preset time interval, and finally at least two initial signal segments are generated.

[0098] S2022: Process the initial signal segments based on reversing the time sequence of the initial signal segments to obtain at least two target signal segments.

[0099] In this step, each initial signal segment may be reversed in time sequence. For example, if the original frequency corresponding to the first initial signal segment is from 0 seconds to 1 second, the frequency may be adjusted based on the reversed time sequence so that the frequency becomes the corresponding frequency from 1 second to 0 seconds.

[0100] In this way, all or part of the initial signal segments are processed in the same way to obtain at least two target signal segments, wherein the frequency of the target signal segment is opposite to the frequency of the corresponding initial signal segment in terms of timing (ie, based on timing inversion).

[0101] S2023. Combine at least two target signal segments to obtain a rearranged masked speech signal.

[0102] Optionally, at least two target signal segments may be randomly combined to generate a complete signal, which is the rearranged masked speech signal. Alternatively, at least two target signal segments may be combined according to a preset rule to generate a complete signal, which is the rearranged masked speech signal.

[0103] If based on the preset rule combination, for example, there are 4 target signal segments, namely target signal segment 1, target signal segment 2, target signal segment 3, and target signal segment 4, and the preset rule is 4-3-1-2, then the rearranged masked speech signal obtained after splicing includes target signal segment 4, target signal segment 3, target signal segment 1, and target signal segment 2, respectively.

[0104] Optionally, the preset rule combination may also be the original order of the initial signal segments after beamforming signal segmentation, that is, the rearranged masked speech signal obtained after splicing includes target signal segment 1, target signal segment 2, target signal segment 3, and target signal segment 4 in sequence.

[0105] The method provided in the embodiment of the present application obtains a beamforming signal corresponding to the ambient sound, rearranges the beamforming signal, and obtains a rearranged masking speech signal, so as to interfere with the semantics of the sound collected by the target object through the rearranged masking speech signal after semantic disruption, thereby preventing the target object from extracting effective information from the ambient sound, thereby improving the effect of information confidentiality protection.

[0106] Optionally, sound masking processing can be performed on the target object based on the rearrangement of the masking speech signal and the target signal. The target signal includes: a pseudo speech signal and / or an interference noise signal. The pseudo speech signal is a signal of simulated speech generated according to the target corpus. The interference noise signal can be, for example, broadband interference noise. For example, white noise, pink noise or bable noise can be selected as the interference noise signal. Since broadband interference noise has the characteristics of wide interference frequency band and stable interference effect, it can further improve the stability and masking effect of the sound masking processing.

[0107] The following is an introduction using the target signal including a pseudo speech signal.

[0108] A possible implementation manner is to directly obtain a preset pseudo speech signal, and then perform sound masking processing on the target object based on the rearranged masking speech signal and the pseudo speech signal.

[0109] Another possible implementation method is to obtain a pseudo voice signal corresponding to the target voice scene based on the target voice scene in which the target object is located, and then perform sound masking processing on the target object based on the rearrangement of the masking voice signal and the pseudo voice signal. The target voice scene may include, for example, scenes such as procurement negotiation, program decision-making, product development, or any other scene. Different target voice scenes in the preset corpus correspond to different pseudo voice signals or corpus data.

[0110] When a pseudo voice signal corresponding to a target voice scene is obtained based on a target voice scene in which a target object is located, the target voice scene may be determined, for example, according to a scene corresponding to the ambient sound. For example, the target voice scene may be obtained based on user input information, or the target voice scene may be determined based on external ambient sound recognition. For example, a user may determine the target voice scene by inputting a target voice scene, or by selecting a target voice scene from a preset list of candidate voice scenes. In response to the user's operation, the target voice scene selected or input by the user may be obtained.

[0111] Optionally, the pseudo voice signal may be a pseudo voice signal corresponding to the target voice scene obtained from a corpus. The corpus is a pre-constructed collection of a large number of pseudo voice signals or corpus data, which may cover various fields and various topics, such as news, novels, conversation records, etc. The pseudo voice signal corresponding to the target voice scene may be directly obtained from the corpus, or the corpus data corresponding to the target voice scene may be obtained from the corpus, and then the corresponding pseudo voice signal is generated based on the corpus data.

[0112] Among them, the corpus data corresponding to the target scenario can be randomly selected from the corpus according to the preset corpus quantity. Alternatively, the first corpus quantity corpus data corresponding to the target scenario can be selected from the preset corpus according to the preset target corpus quantity.

[0113] In this implementation, if corpus data is obtained from a corpus and a pseudo voice signal is generated based on the corpus data, the text-to-speech technology can be used to simulate the human pronunciation method based on the semantics, grammar and other information of the text information in the target corpus to generate a pseudo voice signal. Optionally, during the generation process, the timbre, intonation, speech speed and other parameters of the voice can also be adjusted to increase the diversity and authenticity of the pseudo voice signal.

[0114] The method provided in the embodiment of the present application obtains a pseudo voice signal corresponding to the target voice scene based on the target voice scene in which the target object is located, and then performs sound masking processing on the target object based on the rearranged masking voice signal and the pseudo voice signal, so as to further increase the pseudo voice signal that can confuse the effective information included in the external ambient sound on the basis of outputting the rearranged masking voice signal, thereby increasing the difficulty for the target object to extract effective information from the external ambient sound, thereby improving the effect of information confidentiality protection.

[0115] Figure 3 A flow chart of another sound masking method provided in an embodiment of the present application is shown as follows: Figure 3 As shown, the method may specifically include:

[0116] S301. Obtain a target masking signal based on a superimposed signal of a rearranged masking speech signal and a target signal.

[0117] In this step, the rearranged masking speech signal and the target signal may be added to achieve superposition based on the rearranged masking speech signal and the target signal to obtain the target masking signal. Specifically, the rearranged masking speech signal and the target signal may be directly superimposed by aligning the timestamps for addition, or by not aligning the timestamps for addition (but it is necessary to ensure that there is a superposition area between the rearranged masking speech signal and the target signal).

[0118] Optionally, the target masking signal can be obtained based on the superposition signal of the rearranged masking speech signal and the pseudo speech signal; or the target masking signal can be obtained based on the superposition signal of the rearranged masking speech signal and the interference noise signal; or the target masking signal can be obtained based on the superposition signal of the rearranged masking speech signal, the pseudo speech signal, and the interference noise signal.

[0119] S302: Perform sound masking processing on the target object based on the target masking signal.

[0120] A target masking signal is output in the sound masking space where the target object is located, so that the target object can passively collect the target masking signal when collecting sound, thereby realizing sound masking processing for the target object.

[0121] The method provided in the embodiment of the present application obtains a target masking signal including multiple target masking signals for performing sound masking processing on the target object by superimposing the rearranged masking voice signal with at least one of a pseudo voice signal and an interference noise signal, and performs sound masking processing on the target object based on the target masking signal. The method can disrupt the semantics in the external environmental sound by rearranging the masking voice signal, further confuse the effective information included in the external environmental sound by the pseudo voice signal, and improve the stability of the sound masking processing by adding the noise interference signal, thereby increasing the difficulty for the target object to extract effective information from the external environmental sound and improving the effect of information confidentiality protection.

[0122] Figure 4 A flow chart of another sound masking method provided in an embodiment of the present application is shown as follows: Figure 4 As shown, the method may also include:

[0123] S401: Determine masking strength information of a target masking signal according to external ambient sound.

[0124] The masking strength information is used to determine or limit the masking strength of the target masking signal. When the masking strength of the target masking signal is stronger, the masking effect of the target masking signal is better, but the interference to the external environment of the sound masking space is stronger; when the masking strength of the target masking signal is weaker, the masking effect of the target masking signal is worse, but the interference to the external environment of the sound masking space is weaker.

[0125] Therefore, the masking strength information of the target masking signal can be determined according to the external environmental sound, so as to reduce the interference to the external environment as much as possible while ensuring the masking effect of the target masking signal.

[0126] Specifically, for example, the masking strength information of the target masking signal related to the external ambient sound can be determined based on one or more factors such as the energy information of the external ambient sound, the type of the sound source (for example, different masking strength information is used for different types of sound sources), etc. Optionally, the passive sound insulation isolation of the target object can be further considered to determine the masking strength information of the target masking signal.

[0127] Optionally, the masking strength information may include, for example, a masking strength upper limit and / or a masking state.

[0128] Among them, the masking strength upper limit is used to limit the upper limit of the sound pressure level of the target masking signal. The masking strength upper limit can be determined according to the energy information and / or sound source type of the external ambient sound to ensure that the sound pressure level of the target masking signal is less than or equal to the masking strength upper limit, so that the target masking signal can generate less interference to the external environment while ensuring the masking effect.

[0129] Among them, the masking state is used to configure multiple strategies of the target masking signal to adaptively adjust the strategy of the target masking signal according to the changes in the external environmental sound, so as to achieve different masking strength upper limits and / or adjustments to the content of the superimposed signal in the target masking signal under different masking states, so as to minimize the interference to the external environment while ensuring the masking effect of the target masking signal.

[0130] S402: compress the superimposed signal according to the masking strength information to obtain a target masking signal.

[0131] According to the masking strength upper limit included in the masking strength information, the superimposed signal generated in the above-mentioned embodiment is compressed so that the sound pressure level of the superimposed signal is less than or equal to the masking strength upper limit, thereby generating a target masking signal that minimizes the interference to the external environment while ensuring the masking effect.

[0132] Specifically, if the sound pressure level of the superimposed signal is greater than the upper limit of the masking intensity, the superimposed signal can be compressed to obtain the target masking signal. Among them, the sound pressure level of the target masking signal is less than or equal to the upper limit of the masking intensity. For example, if the upper limit of the masking intensity is 50db and the sound pressure level of the superimposed signal is 60db, the superimposed signal can be compressed (for example, an audio compression algorithm can be used, such as amplitude-based compression. Detect the amplitude of the superimposed signal at each moment, and when the sound pressure level exceeds the upper limit of the masking intensity, reduce the signal amplitude proportionally. For example, reduce it by a certain coefficient so that the sound pressure level drops to the upper limit of the masking intensity, and compress the superimposed signal to obtain the target masking signal), so that the sound pressure level of the superimposed signal is lower than 50db, so as to obtain a target masking signal with a sound pressure level lower than 50db, and minimize the interference to the external environment while ensuring the masking effect.

[0133] The method provided in the embodiment of the present application determines the masking strength information of the target masking signal according to the external ambient sound, compresses the superimposed signal according to the masking strength information, and obtains the target masking signal, so as to adjust the masking strength of the target masking signal according to the actual situation of the external ambient sound, and minimizes the target masking signal that interferes with the external environment while ensuring the masking effect of the target masking signal, thereby not only improving the masking effect of the sound masking processing, but also reducing the noise interference to the external environment, thereby improving the user experience.

[0134] In the following, taking the masking strength information including the masking strength upper limit as an example, how to determine the masking strength information of the target masking signal according to the ambient sound is described in detail. Figure 5 A flow chart of another sound masking method provided in an embodiment of the present application is shown as follows: Figure 5 As shown, the aforementioned step S401 may specifically include:

[0135] S501: Determine an upper limit of masking intensity corresponding to the frequency band according to energy information of the external ambient sound in at least one frequency band.

[0136] Among them, the frequency band refers to a continuous frequency interval within a certain frequency range. The external environmental sound can be decomposed into a combination of different frequency components to facilitate the analysis and processing of the sound signal. The energy information of the external environmental sound reflects the intensity of the external environmental sound in a certain frequency band.

[0137] In this step, the external ambient sound can be subjected to spectrum analysis, for example, by converting the time domain audio signal into a frequency domain signal through fast Fourier transform, and obtaining the amplitude information of the ambient sound at each frequency. Then, according to the pre-divided frequency bands, the total energy in each frequency band is calculated. For example, the entire audible frequency range can be divided into multiple 1 / 3 octave frequency bands, and the signal power in each frequency band is calculated respectively.

[0138] According to the energy information of each frequency band, the upper limit of the masking intensity corresponding to the frequency band can be determined in combination with the auditory characteristics of the human ear and the acoustic principle. For example, a mapping relationship between energy information and the upper limit of masking intensity can be established. For example, for frequency bands with higher energy, the upper limit of masking intensity can be appropriately increased; for frequency bands with lower energy, the upper limit of masking intensity can be correspondingly reduced. It is also possible to test the masking effect of the human ear on sounds of different frequency bands in different environments through experimental methods, so as to determine the appropriate upper limit of masking intensity.

[0139] S502: Determine the upper limit of the masking strength of the target masking signal according to the upper limit of the masking strength corresponding to the frequency band.

[0140] One possible implementation method is to assign corresponding weights to each frequency band, multiply the upper limit of the masking strength corresponding to each frequency band by its corresponding weight, and then sum the results of all frequency bands to obtain the preliminary upper limit of the masking strength of the target masking signal. Since different frequency bands may contribute differently to the overall masking effect, the weight corresponding to each frequency band can be considered based on factors such as the auditory characteristics of the human ear, the characteristics of the ambient sound, and the specific application scenario. For example, in an environment dominated by mid-frequency sounds, the weight of the mid-frequency band can be appropriately increased.

[0141] Another possible implementation manner may be to determine the upper limit of the masking strength of the target masking signal in each frequency band according to the upper limit of the masking strength corresponding to each frequency band.

[0142] In an implementation method of determining the upper limit of the masking strength of the target masking signal in each frequency band according to the upper limit of the masking strength corresponding to each frequency band, the upper limit of the masking strength corresponding to the frequency band can be determined according to the energy information of the external ambient sound in at least one frequency band. Then, the upper limit of the masking strength of the target masking signal is determined according to the upper limit of the masking strength corresponding to the frequency band.

[0143] Specifically, the external environmental sound can be converted from the time domain to the frequency domain through spectrum analysis methods such as fast Fourier transform, the entire audible frequency range can be divided into multiple frequency bands, and the sum of the energy in each frequency band can be calculated to obtain the energy information of the external environmental sound in at least one frequency band. Then, based on the auditory characteristics of the human ear, acoustic principles, and related research results or experimental tests, a mapping relationship between energy information and the upper limit of masking intensity is established to determine the upper limit of masking intensity corresponding to each frequency band. According to the upper limit of masking intensity corresponding to each frequency band, the upper limit of masking intensity of the target masking signal in each frequency band is determined, and according to the upper limit of masking intensity of the target masking signal in each frequency band, the upper limit of masking intensity of the target masking signal is determined.

[0144] In this implementation, optionally, the passive sound insulation isolation of the target object may be further considered to determine the upper limit of the masking intensity of each frequency band. The passive sound insulation isolation of the target object refers to the passive sound insulation effect of the sound masking space where the target object is located. For example, taking the sound masking space as a closed box, the passive sound insulation effect of the box is the passive sound insulation isolation of the target object.

[0145] Due to the passive sound insulation isolation of the target object, the intensity of the target masking signal received outside the sound masking space can be weakened to a certain extent (for example, when the passive sound insulation isolation is 30db and the target masking signal is 50db, then outside the sound masking space, the target masking signal is weakened by the passive sound insulation isolation and its intensity becomes 20db). Therefore, when determining the upper limit of the masking intensity of each frequency band, the passive sound insulation isolation of the target object can be considered, so as to appropriately increase the upper limit of the masking intensity of each frequency band while ensuring that the external environment is not disturbed.

[0146] For example, when the target masking signal is 40db, the interference to the external environment is within an acceptable range, and the passive sound insulation isolation of the target object is 30db, then it can be determined that the upper limit of the masking intensity of the target masking signal is 30+40=70db, that is, when the sound pressure level of the target masking signal is below 70db, the interference to the external environment is within an acceptable range.

[0147] Next, taking the masking state of the masking strength information, where the masking state is used to determine the content included in the target masking signal and / or adjust the upper limit of the masking strength as an example, how to determine the masking strength information of the target masking signal according to the ambient sound is described in detail. Figure 6 A flow chart of another sound masking method provided in an embodiment of the present application is shown as follows: Figure 6 As shown, the aforementioned step S401 may specifically include:

[0148] S601: Determine whether there is sound from a sound source of the target type according to the spectral features of the external environmental sound.

[0149] Among them, the sound source of the target type can be, for example, the sound source of the human voice included in the external environmental sound, or it can also be other specific sound sources, and the present application does not limit this. Since information leakage is usually caused by the leakage of semantics included in the human voice, when there is a human voice in the external environmental sound, a strong masking effect is required to avoid information leakage; and when there is no human voice in the external environmental sound, there is usually a small risk of information leakage, and sound masking can be performed with a weaker masking effect. Therefore, the subsequent embodiments are introduced using the sound source of the human voice as an example.

[0150] In this step, the short-time Fourier transform can be used to analyze the spectral features of the external environmental sound. For example, the short-time Fourier transform can be used to convert the time domain signal of the external environmental sound into the frequency domain to obtain a spectrogram, which can show the energy distribution of the sound over time and frequency. Then, construct or use an existing template library containing typical human voice spectral features, which include fundamental frequency range, resonance peak distribution, etc. The spectral features of the current external environmental sound are compared with the features in the template library, and the similarity can be calculated using correlation analysis and other methods. If the similarity exceeds a preset threshold, it is determined that there is a sound from the human voice; if the similarity is lower than the threshold, it is determined that there is no sound.

[0151] Alternatively, a machine learning algorithm can be used to train a large number of human voice and non-human voice samples, allowing the model to automatically learn and identify whether there is a human voice in the external environmental sound.

[0152] If there is sound from a sound source belonging to the target type, that is, there is a human voice, the representation requires a strong masking effect to avoid information leakage, and step S602 is executed; if there is no sound from a sound source belonging to the target type, that is, there is no human voice, the representation does not require a strong masking effect, and step S603 is executed.

[0153] S602: Determine that the masking state of the target masking signal is a first masking state.

[0154] The content of the target masking signal corresponding to the first masking state is different from the content of the target masking signal corresponding to the second masking state in the subsequent S603, and / or the upper limit of the masking strength corresponding to the first masking state is greater than the upper limit of the masking strength corresponding to the second masking state. That is, the first masking state is a stronger masking state, so as to provide a stronger masking effect when there is a human voice to avoid information leakage; the second masking state is a weaker masking state, so as to provide a weaker masking effect when there is no human voice to reduce interference and influence on the external environment.

[0155] Specifically, if the content of the target masking signal corresponding to the first masking state is different from the content of the target masking signal corresponding to the second masking state in the subsequent S603, it is possible to determine what the content of the target masking signal corresponding to the first masking state includes according to a preset rule. For example, the content of the target masking signal corresponding to the first masking state can be set to include a rearranged masking voice signal, a pseudo voice signal, and an interference noise signal; and the content of the target masking signal corresponding to the second masking state can be set to include an interference noise signal, so that the first masking state can provide a stronger masking effect to avoid information leakage, and the second masking state can provide a weaker masking effect to reduce interference and impact on the external environment.

[0156] If the masking strength upper limit corresponding to the first masking state is different from the masking strength upper limit corresponding to the second masking state in the subsequent S603, a higher masking strength upper limit can be configured for the first masking state to enhance the masking effect to avoid information leakage, and a lower masking strength upper limit can be configured for the second masking state to reduce the masking effect, thereby reducing interference and impact on the external environment.

[0157] S603: Determine that the masking state of the target masking signal is a second masking state.

[0158] The method provided in the embodiment of the present application determines whether there is sound from a sound source belonging to the target type based on the spectral features of the external ambient sound. If there is sound from the sound source belonging to the target type, it indicates that a stronger masking effect is needed to avoid information leakage, and the masking state of the target masking signal is determined to be a first masking state with a stronger masking effect to reduce the risk of information leakage. If there is no sound from the sound source belonging to the target type, it indicates that a stronger masking effect is not needed to avoid information leakage, and the masking state of the target masking signal can be determined to be a second masking state with a weaker masking effect to reduce the interference and influence of the target masking signal on the external environment.

[0159] Figure 7 This is a schematic diagram of the structure of a sound masking system provided in an embodiment of the present application. Figure 7 As shown, the sound masking system includes: a microphone array, a speaker array, and a processing module.

[0160] The microphone array is located outside the sound masking space, and the speaker array and the target object to be sound masked are located inside the sound masking space.

[0161] The microphone array is used to obtain the ambient sound of the target object, and the ambient sound is the ambient sound outside the shell (i.e., the external ambient sound in the aforementioned method embodiment). The processing module is used to obtain the rearranged masking voice signal based on the external ambient sound by any method in the aforementioned method embodiment, and play the rearranged masking voice signal through the speaker array to perform sound masking processing on the target object based on the rearranged masking voice signal.

[0162] Optionally, the processing module is used to obtain the masking strength of the superimposed signal and the target masking signal based on the external ambient sound by any of the methods in the aforementioned method embodiments, and compress the superimposed signal based on the masking strength to generate a target masking signal, and play the target masking signal through a speaker array to perform sound masking processing on the target object based on the target masking signal.

[0163] In a possible implementation, the sound masking system may further include a shell, the interior of the shell forming a sound masking space, the target object may be placed in the sound masking space, and the speaker array plays a target masking signal to perform sound masking processing on the target object in the sound masking space based on the target masking signal.

[0164] Optionally, the side wall of the shell can be filled with a sound insulation material, and the sound insulation material can be any currently available sound absorbing / insulating material, which is not limited in this application. Through the sound insulation material, the passive sound insulation isolation of the target object can be determined (that is, the passive sound insulation isolation of the target object is determined by the sound insulation effect of the sound insulation material). Alternatively, the shell can be designed as a physical structure with a sound insulation effect, in which case the passive sound insulation isolation of the target object is determined by the sound insulation effect of the shell.

[0165] For ease of understanding, illustrative purposes only, Figure 8 This is a schematic diagram of the structure of another sound masking system provided in an embodiment of the present application. Figure 8 As shown, the system includes: a passive sound insulation module, an active masking module, an environment perception module, and a masking generation module.

[0166] The passive sound insulation module includes a sound absorption structure unit, a sound absorption and insulation unit, and a system support structure unit. The system support structure unit is used to support the deployment of the active masking module, the environment perception module, and the masking generation module on the passive sound insulation module. The sound absorption structure unit is used to provide a physical structure with sound insulation effect. The sound absorption and insulation unit is composed of sound insulation materials.

[0167] The active masking module includes a masking area (i.e., the sound masking area where the target object is placed), a speaker array unit (i.e., the aforementioned speaker array), and a drive filter unit. The drive filter unit is used to filter the target masking signal, and the speaker array unit outputs the target masking signal after filtering, so as to perform sound masking on the target object in the sound masking space based on the target masking signal.

[0168] The environment perception module includes a microphone array unit (i.e., the aforementioned microphone array), a scene analysis unit, and a beamforming unit. Among them, the functions of the scene analysis unit and the beamforming unit can be provided by the aforementioned processing module, and the processing module realizes the role of the scene analysis unit by analyzing and determining the masking state and / or the upper limit of the masking strength corresponding to the external environmental sound. The processing module generates the aforementioned beamforming signal by processing the external environmental sound collected by the microphone array unit.

[0169] The masking production module is used to generate a rearranged masking voice signal or a superimposed signal, which may include a dynamic range control unit, a voice interference unit, a fixed masking unit, and a pseudo voice generation unit. The functions of the dynamic range control unit, the voice interference unit, the fixed masking unit, and the pseudo voice generation unit may all be provided by the aforementioned processing module. The processing module may realize the function of the fixed masking unit by generating an interference noise signal, realize the function of the voice interference unit by generating a rearranged masking voice signal, and realize the function of the pseudo voice generation unit by generating a pseudo voice signal.

[0170] The modules included in the system and the units in each module can implement the corresponding sound masking method in the aforementioned method embodiment. The implementation principles and technical effects are similar and will not be repeated here.

[0171] Fig. 9 Schematic diagram of the structure of a sound masking device provided in an embodiment of the present application. Figure 4 As shown, the sound masking device may include: a processing module 11 and a control module 12 .

[0172] The processing module 11 is used to rearrange the external environmental sound of the sound masking space where the target object is located to obtain a rearranged masking speech signal, and the difference between the spectral characteristics of the rearranged masking speech signal and the external environmental sound is less than or equal to a preset threshold.

[0173] The control module 12 is used to perform sound masking processing on the target object based on the rearranged masking speech signal.

[0174] Optionally, the processing module 11 is specifically configured to obtain a beamforming signal corresponding to the ambient sound, and to rearrange the beamforming signal to obtain a rearranged masking speech signal, wherein the beamforming signal is related to the sound source position and / or the number of sound sources of the ambient sound.

[0175] Optionally, the processing module 11 is specifically configured to time slice the beamforming signal to obtain at least two initial signal segments, process the initial signal segments based on reversing the time order of the initial signal segments to obtain at least two target signal segments, and combine the at least two target signal segments to obtain a rearranged masked speech signal.

[0176] Optionally, the control module 12 is specifically configured to perform sound masking processing on the target object based on the rearrangement of the masking speech signal and the target signal. The target signal includes: a pseudo speech signal and / or an interfering noise signal.

[0177] Optionally, when the target signal includes a pseudo speech signal, the processing module 11 is further configured to acquire a pseudo speech signal corresponding to the target speech scene based on the target speech scene in which the target object is located.

[0178] Optionally, the processing module 11 is specifically configured to obtain a pseudo speech signal corresponding to the target speech scene from a corpus based on the target speech scene in which the target object is located.

[0179] Optionally, the processing module 11 is specifically configured to obtain a target voice scene based on user input information.

[0180] Optionally, the processing module 11 is specifically configured to obtain a target masking signal based on the superposition signal of the rearranged masking speech signal and the target signal. The control module 12 is specifically configured to perform sound masking processing on the target object based on the target masking signal.

[0181] Optionally, the processing module 11 is specifically configured to determine masking strength information of the target masking signal according to the external ambient sound, and compress the superimposed signal according to the masking strength information to obtain the target masking signal.

[0182] Optionally, when the masking strength information includes a masking strength upper limit, the processing module 11 is specifically configured to determine a masking strength upper limit corresponding to a frequency band according to energy information of the external ambient sound in at least one frequency band, and determine a masking strength upper limit of the target masking signal according to the masking strength upper limit corresponding to the frequency band.

[0183] Optionally, the processing module 11 is specifically configured to determine an upper limit of the masking intensity of each frequency band according to energy information of the external ambient sound in at least one frequency band and the passive sound insulation isolation of the target object.

[0184] Optionally, the processing module 11 is specifically configured to compress the superimposed signal to obtain a target masking signal if the sound pressure level of the superimposed signal is greater than the upper limit of the masking strength. The sound pressure level of the target masking signal is less than or equal to the upper limit of the masking strength.

[0185] Optionally, when the masking strength information includes a masking state, and the masking state is used to determine the content included in the target masking signal, and / or when adjusting the upper limit of the masking strength, the processing module 11 is specifically used to determine whether there is a sound from a sound source belonging to the target type based on the spectral characteristics of the external ambient sound. If so, the masking state of the target masking signal is determined to be the first masking state. If not, the masking state of the target masking signal is determined to be the second masking state. The second masking state is different from the content of the target masking signal corresponding to the first masking state, and / or the upper limit of the masking strength corresponding to the second masking state is less than the upper limit of the masking strength corresponding to the first masking state.

[0186] The sound masking device provided in the embodiment of the present application can execute the sound masking method in the above method embodiment, and its implementation principle and technical effect are similar, which will not be repeated here.

[0187] The present application also provides a computer program product, including a computer program, which implements the above method when executed by a processor.

[0188] The present application also provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the above method is implemented.

[0189] The above-mentioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The readable storage medium can be any available medium that can be accessed by a general or special-purpose computer.

[0190] An exemplary readable storage medium is coupled to a processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (Application Specific Integrated Circuits, referred to as: ASIC). Of course, the processor and the readable storage medium can also exist in the device as discrete components.

[0191] The division of units is only a logical function division, and there may be other divisions in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.

[0192] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0193] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0194] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0195] Those skilled in the art can understand that all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the aforementioned storage medium includes: ROM, RAM, disk or optical disk and other media that can store program codes.

[0196] Finally, it should be noted that those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses or adaptations of the present invention, which follow the general principles of the present invention and include common knowledge or customary technical means in the art not disclosed by the present invention, are not limited to the precise structure described above and shown in the drawings, and may be modified and changed in various ways without departing from the scope thereof. The scope of the present invention is limited only by the appended claims.

Claims

1. A sound masking method, characterized in that: The method comprises: Rearranging the external environmental sound of the sound masking space where the target object is located to obtain a rearranged masking speech signal, wherein the difference between the spectral characteristics of the rearranged masking speech signal and the external environmental sound is less than or equal to a preset threshold; Based on the rearranged masking speech signal, sound masking processing is performed on the target object.

2. The method according to claim 1, characterized in that The step of rearranging the ambient sound of the target object to obtain a rearranged masking speech signal includes: Acquire a beamforming signal corresponding to the ambient sound, where the beamforming signal is related to a sound source position and / or the number of sound sources of the ambient sound; The beamforming signal is rearranged to obtain the rearranged masked speech signal.

3. The method according to claim 2, characterized in that The step of rearranging the beamforming signal to obtain the rearranged masked speech signal includes: Time slicing the beamforming signal to obtain at least two initial signal segments; Processing the initial signal segments based on reversing the time order of the initial signal segments to obtain at least two target signal segments; The at least two target signal segments are combined to obtain the rearranged masked speech signal.

4. The method according to claim 1, characterized in that The performing sound masking processing on the target object based on the rearranged masking speech signal comprises: Based on the rearranged masking speech signal and the target signal, the target object is subjected to sound masking processing; the target signal includes: a pseudo speech signal and / or an interfering noise signal.

5. The method according to claim 4, characterized in that The target signal includes a pseudo speech signal, and the method further includes: Based on the target speech scene in which the target object is located, a pseudo speech signal corresponding to the target speech scene is acquired.

6. The method according to claim 5, characterized in that The step of acquiring a pseudo speech signal corresponding to the target speech scene based on the target speech scene in which the target object is located includes: Based on the target speech scene in which the target object is located, a pseudo speech signal corresponding to the target speech scene is acquired from a corpus.

7. The method according to claim 5, characterized in that The method further comprises: Based on the user's input information, the target voice scene is acquired.

8. The method according to claim 4, characterized in that The performing sound masking processing on the target object based on the rearranged masking speech signal and the target signal comprises: Acquire a target masking signal based on a superposition signal of the rearranged masking speech signal and the target signal; Based on the target masking signal, sound masking processing is performed on the target object.

9. The method according to claim 8, characterized in that The step of obtaining a target masking signal based on a superposition signal of the rearranged masking speech signal and the target signal comprises: Determining masking strength information of the target masking signal according to the external environmental sound; The superimposed signal is compressed according to the masking strength information to obtain the target masking signal.

10. The method according to claim 9, characterized in that The masking strength information includes a masking strength upper limit, and determining the masking strength information of the target masking signal according to the external ambient sound includes: Determining, according to energy information of the external ambient sound in at least one frequency band, an upper limit of masking intensity corresponding to the frequency band; The upper limit of the masking strength of the target masking signal is determined according to the upper limit of the masking strength corresponding to the frequency band.

11. The method according to claim 10, characterized in that The determining, according to energy information of the external ambient sound in at least one frequency band, an upper limit of the masking intensity corresponding to the frequency band comprises: The upper limit of the masking intensity of each frequency band is determined according to the energy information of the external ambient sound in at least one frequency band and the passive sound insulation isolation of the target object.

12. The method according to claim 10, characterized in that The compressing the superimposed signal according to the masking strength information to obtain the target masking signal includes: If the sound pressure level of the superimposed signal is greater than the upper limit of the masking strength, the superimposed signal is compressed to obtain the target masking signal; the sound pressure level of the target masking signal is less than or equal to the upper limit of the masking strength.

13. The method according to claim 11, characterized in that The masking strength information includes a masking state, and the masking state is used to determine the content included in the target masking signal and / or to adjust the upper limit of the masking strength; The step of determining the masking strength of the target masking signal according to the external ambient sound comprises: Determining whether there is sound from a sound source belonging to a target type according to the spectral features of the external environmental sound; If so, determining that the masking state of the target masking signal is a first masking state; If not present, the masking state of the target masking signal is determined to be the second masking state, the second masking state is different from the content of the target masking signal corresponding to the first masking state, and / or the masking strength upper limit corresponding to the second masking state is less than the masking strength upper limit corresponding to the first masking state.

14. A sound masking system, characterized in that: The sound masking system comprises: a microphone array, a speaker array, and a processing module; The microphone array is located outside the sound masking space, and the speaker array and the target object to be sound masked are located inside the sound masking space; The microphone array is used to acquire the ambient sound of the target object, where the ambient sound is the ambient sound outside the housing; The processing module is used to obtain the rearranged masking speech signal based on the ambient sound by the method according to any one of claims 1 to 13, and play the rearranged masking speech signal through the speaker array to perform sound masking processing on the target object based on the rearranged masking speech signal.

15. The system according to claim 14, characterized in that The sound masking system further comprises a shell; the interior of the shell constitutes the sound masking space.

16. The system according to claim 15, characterized in that The side wall of the shell is filled with sound insulation material.

17. A sound masking device, characterized in that: The device comprises: A processing module is used to rearrange the external environmental sound of the sound masking space where the target object is located to obtain a rearranged masking speech signal, wherein the difference between the spectral characteristics of the rearranged masking speech signal and the external environmental sound is less than or equal to a preset threshold; A control module is used to perform sound masking processing on the target object based on the rearranged masking speech signal.

18. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 13 when executed by a processor.

19. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 13 when being executed by a processor.

Citation Information

Patent Citations

  • Target voice privacy protection method and system

    CN102543066A

  • Sound masking signal generating method and system

    CN103886858A

  • A noise masking device and a method for masking noise

    CN113302681A

  • Method and device for improving voice privacy of mute cabin

    CN115910018A

  • System for providing a reduction of audiable noise perception for a human user

    US20090074199A1