Personal electronic device enhancing call privacy

TWI939047BActive Publication Date: 2026-09-11XMEMS LABS INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
TW114121463
Authority / Receiving Office
TW · TW
Patent Type
Patents
Current Assignee / Owner
Priority Date
2025-06-05
Filing Date
2025-06-09
Publication Date
2026-09-11
Estimated Expiration
2045-06-08

AI Technical Summary

Technical Problem

In enclosed public spaces, mobile phone users face significant risks of unintentional eavesdropping, leading to potential information theft, fraud, social embarrassment, and security breaches due to inadequate call privacy measures.

Method used

A personal electronic device with a primary sound-generating device and auxiliary sound-generating devices that emit masking sounds or countersounds to reduce speech intelligibility of bystanders, utilizing audio beamforming to direct acoustic energy away from unintended listeners.

Benefits of technology

Enhances call privacy by reducing bystanders' ability to understand conversations, minimizing the risk of information leakage and maintaining user confidentiality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure TWG2TB001910615_001
    Figure TWG2TB001910615_001
  • Figure TWG2TB001910615_002
    Figure TWG2TB001910615_002
  • Figure TWG2TB001910615_003
    Figure TWG2TB001910615_003
Patent Text Reader

Abstract

A personal electronic device includes a primary sound-emitting device for emitting a predetermined sound to a predetermined user; and an auxiliary sound-emitting device for emitting a masking sound or a countersound to reduce the speech intelligibility of a bystander located near the predetermined user of the personal electronic device.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application refers to a personal electronic device, and more particularly to a personal electronic device that can enhance call privacy. [Previous Technology]

[0002] Unless otherwise stated herein, the methods described in this section are not prior art within the scope of this application, nor are they considered prior art by virtue of their inclusion in this section.

[0003] For mobile phone users, call privacy is of paramount importance, especially in enclosed public spaces such as elevators. In such environments, the risk of unintentional eavesdropping increases significantly. Close proximity to others means that private conversations—whether concerning financial details, health issues, or sensitive work discussions—are easily overheard.

[0004] A lack of privacy can have serious consequences. Stolen information can be used for identity theft, targeted fraud, or corporate espionage. Accidental leaks of confidential information can also lead to social embarrassment, career consequences, and even security breaches. Furthermore, the awareness that being eavesdropped on can severely impact freedom of speech makes people hesitant to discuss important issues publicly. Finally, robust call privacy measures are crucial to ensuring the security and confidentiality of mobile calls, regardless of the actual circumstances.

[0005] Therefore, it is necessary to strengthen call privacy. [Summary of the Invention]

[0006] Therefore, the main purpose of this application is to propose a personal electronic device to improve the deficiencies of prior art.

[0007] One embodiment of this application provides a personal electronic device. The personal electronic device includes a primary sound-generating device for generating a predetermined sound for a predetermined user; and an auxiliary sound-generating device for emitting a masking sound or a countersound to reduce the speech intelligibility of a bystander located near the predetermined user of the personal electronic device.

[0008] One embodiment of this application provides a personal electronic device. The personal electronic device includes a plurality of auxiliary sound-generating devices for performing an audio beamforming operation and forming at least one audio beam; wherein the at least one audio beam is used to eliminate or minimize acoustic energy in an angular direction toward a bystander, so as to reduce the speech intelligibility of the bystander.

Implementation Method

[0010] Figure 1 illustrates a schematic diagram of the front (1(a)), side (1(b)), and back (1(c)) of a personal electronic device 10 according to an embodiment of this application. In the embodiment of Figure 1, the personal electronic device 10 is a telephone. The personal electronic device 10 includes a primary sound-generating device SPDmain and an auxiliary sound-generating device SPDaux. In one embodiment, the primary or auxiliary sound-generating device can produce sound by generating a plurality of air pulses. In other words, the primary or auxiliary sound-generating device can be implemented by an air-pulse generating (APG) device (such as the air-pulse generating devices proposed in U.S. Patent Application Nos. 18 / 321,759, 18 / 829,245, etc.), which has the advantage of small size, and is not limited thereto.

[0011] The main sound-generating device SPDmain is generally referred to as the receiver of a telephone. It is usually located on the front of the telephone and is used to generate the other party's voice during a call. The other party's voice is identified as a predetermined voice of a predetermined user, wherein the predetermined user in this application usually refers to the telephone holder / user, and vice versa.

[0012] The auxiliary sound-emitting device SPDaux may be disposed on the bottom or side of the personal electronic device (as shown in Figures 1(a) and 1(b)) or on the back of the personal electronic device (as shown in Figure 1(c)) to emit sound to the periphery of the personal electronic device 10. The auxiliary sound-emitting device SPDaux is configured to emit sound to disrupt / reduce the speech intelligibility of an observer. The sound emitted by the auxiliary sound-emitting device SPDaux may be or may include an anti-sound (corresponding to a predetermined voice of a predetermined user), a masking sound (described in detail below), or a combination of both.

[0013] In this application, a bystander refers to a non-intended listener near the intended user of a personal electronic device (such as 10).

[0014] It is worth noting that, in the embodiment shown in Figure 1, the personal electronic device 10 includes a plurality of auxiliary sound-emitting devices, but is not limited thereto. The number, location, and / or arrangement of the auxiliary sound-emitting devices can be designed according to actual needs. As long as the personal electronic device includes at least one auxiliary sound-emitting device, and the auxiliary sound-emitting device is configured to emit masking or counter-sound to disrupt / reduce the speech intelligibility of bystanders located near the user of the personal electronic device, it should fall within the scope of this invention.

[0015] It should also be noted that personal electronic devices are not limited to telephones. Personal electronic devices can be personal computers, tablets, and smart wearable devices, such as smartwatches, smart bracelets, and smart glasses. As long as the personal electronic device can make voice calls, it should fall within the scope of this invention.

[0016] Figure 2 illustrates a schematic diagram of a personal electronic device 20 according to an embodiment of this application. The personal electronic device 20 demonstrates an embodiment that emits a reverse sound, corresponding to a predetermined sound for a predetermined user. In addition to the main sound-emitting device SPDmain and the auxiliary sound-emitting device SPDaux, the personal electronic device 20 also includes a polarity reversal circuit 22 and a gain adjustment circuit 24.

[0017] The main sound-emitting device SPDmain can emit a predetermined sound p(t) according to a voice call signal Svc, wherein the voice call signal Svc can be obtained from a modem (used to bridge the personal electronic device and the communication network system) of the personal electronic device 20 (not shown in Figure 2).

[0018] On the other hand, the polarity reversal circuit 22 and the gain adjustment circuit 24 can work together to make the auxiliary sound-generating device SPDaux produce an anti-sound -a∙p(t), where the negative sign "-" is contributed by the polarity reversal circuit 22, and the gain factor or amplitude "a" is contributed by the gain adjustment circuit 24.

[0019] The purpose of anti-sound -a∙p(t) is to ideally cancel out the predetermined sound p(t) emitted from the main sound-emitting device SPDmain, or at least to disruptively interfere with the predetermined sound p(t) to reduce the acoustic energy perceived by an observer (corresponding to the predetermined sound p(t) or the voice call signal Svc). Furthermore, the purpose of anti-sound -a∙p(t) is to reduce the acoustic energy perceived by an observer to below a certain level, making it impossible for the observer to understand the speech.

[0020] The amplitude "a" can be determined according to the actual situation. For example, the amplitude "a" of the reverse sound can be estimated by taking into account propagation attenuation, 1 / r path loss, and / or channel obstruction between the auxiliary sound-generating device SPDaux and the ear canal opening of the telephone user. In one embodiment, a (proximity) sensor 26 can be provided to detect whether the telephone is attached to the ear of the telephone user. If the (proximity) sensor 26 detects a positive value (for path loss and obstruction), the gain adjustment block / circuit 24 can determine a = a1; or if the (proximity) sensor 26 detects a negative value (for path loss only), the gain adjustment block / circuit 24 can determine a = a2.

[0021] In addition to eliminating sound, sound masking can also be used to reduce speech intelligibility.

[0022] Sound masking and sound cancellation are different concepts. Sound cancellation involves generating an anti-sound (e.g., -a∙p(t)) with the opposite polarity to the canceled sound (e.g., p(t)), thus reducing the acoustic effect of the sound and making it more difficult to distinguish. On the other hand, sound masking can be, for example, adding a specially tuned sound to match the frequency of human speech, thereby reducing speech intelligibility. Furthermore, sound masking or auditory masking (in the field of psychoacoustics) represents the concept of being unable to perceive a sound due to the presence of another sound.

[0023] In order to reduce speech intelligibility, the auxiliary sound-generating device SPDaux can generate masking noise or interference noise as a masking sound.

[0024] Figure 3 illustrates a schematic diagram of a personal electronic device 30 according to an embodiment of this application. Unlike the auxiliary sound-generating device SPDaux in personal electronic device 20, which generates a countersound, the auxiliary sound-generating device SPDaux in personal electronic device 30 emits a masking tone ms(t) based on a masking tone signal MS. In one embodiment, the masking tone ms(t) is masking noise. In another embodiment, the masking noise is modulated noise used to match the frequency of human speech. In the embodiment shown in Figure 3, personal electronic device 30 includes a masking tone generator 32, which includes a filter 320 for generating a modulated noise TN as the masking tone signal MS based on broadband noise or white noise N.

[0025] It is worth noting that human speech contains vowels and consonants. Consonants typically have higher frequencies than vowels, but their acoustic energy is usually lower than that of vowels. However, distinguishing between vowels plays a crucial role in speech intelligibility. For example, words like "top," "pop," and "bob" share the same vowel but different consonants, thus conveying different meanings.

[0026] Human hearing is very sensitive to the spectrum of consonants (usually 2-4 kHz), which can be verified by the hearing threshold diagram shown in Figure 4 (in a quiet room). As can be seen from the figure, the hearing threshold is lower in the 2-4 kHz frequency band, and some or many consonant frequencies are located in this frequency band.

[0027] To illustrate in more detail, Figure 4 shows the relationship between sound pressure level (SPL) and frequency, illustrating the hearing threshold Hth0 in a quiet room. The hearing threshold is the lowest sound level of a pure tone that a normal human ear with normal hearing can hear in a specific environment.

[0028] The hearing threshold Hth0 is relatively low in the range of 800 Hz to 6.3 kHz, and human hearing is more sensitive in this frequency band. In other words, a 2.5 kHz monotone with a sound pressure level of about 40 dB can be clearly heard in a quiet room. However, if noise of the shape of spectrum SM1 or SM2 as shown in Figure 4 is present, the hearing threshold will increase in response to spectrum SM1 / SM2. For example, the hearing threshold Hth1 / Hth2 will be higher than 40 dB near 2.5 kHz. In other words, if noise on spectrum SM1 / SM2 is present, even if a 40 dB 2.5 kHz monotone is present, human hearing will not be able to hear it.

[0029] Inspired by Figure 4, the purpose of masking tone ms(t) or modulated noise TN is to raise the human hearing threshold within a specific frequency band, making it impossible to distinguish or interpret the telephone user's voice (at least the consonants in the telephone user's voice). This is called spectral masking or simultaneous masking. The masking tone ms(t) or modulated noise TN may have a spectrum similar to SM1 or SM2, which means that the masking tone ms(t) may contain band-limited noise, whose noise energy is concentrated in a noise frequency band (similar to SM1), and this noise frequency band may cover the spectrum of human voice, or cover the spectrum of consonants in human voice. Alternatively, the masking tone ms(t) may contain a plurality of narrow-band tones located at a plurality of masking frequency tones (similar to SM2), and the plurality of masking frequency tones are scattered over a sub-frequency band (such as 2 kHz to 4 kHz) or a speech frequency band (such as 250 Hz to 8 kHz).

[0030] In this application, some intelligibility indicators can be used to evaluate or quantify speech intelligibility, such as the Speech Intelligibility Index (SII), the Speech Transmission Index (STI), the Common Intelligibility Scale (CIS), etc., but not limited to these.

[0031] In one embodiment, the auxiliary sound-generating device SPDaux can simultaneously emit an anti-sound and a masking sound. For example, Figure 5 illustrates a schematic diagram of a personal electronic device 40 according to an embodiment of this application. The personal electronic device 40 includes a masking sound generator 42, which can be considered as an integration of devices 20 and 30, the operational details of which will not be described here.

[0032] It is worth noting that in the personal electronic device 20, the anti-sound system is used to eliminate the sound emitted by the main sound-emitting device SPDmain, but is not limited thereto. The auxiliary sound-emitting device SPDaux can also generate anti-sound to eliminate the voice of the telephone user.

[0033] Figure 6 illustrates a scenario of a telephone user's voice V and an anti-voice U emitted by an auxiliary sound-generating device SPDaux. The anti-voice U is used to cancel or eliminate the voice V. The voice V can first be captured by a sound sensing device SSD (such as a microphone), and the personal electronic device can perform signal processing on the captured voice and generate the anti-voice U accordingly.

[0034] It should be noted that the SSD in Figure 6 will sense the convergence of the inverted sound and the speech sound, represented by U+V. Therefore, it is necessary to extract the speech sound V from the convergence sound U+V.

[0035] Figure 7 illustrates a schematic diagram of a personal electronic device 50 according to an embodiment of this application. The personal electronic device 50 includes a speech extraction circuit 54 and a speech cancellation circuit 52. Based on the aggregate tone U+V, the speech extraction circuit 54 can capture a speech signal Vd corresponding to the speech sound V, wherein the speech signal Vd can be or represents the speech sound V in digital or electronic form / format. The purpose of the speech cancellation circuit 52 is generally to minimize the acoustic energy of the aggregate tone U+V, which can be achieved by generating an inverse signal Ud (through the speech cancellation circuit 52) ​​or an inverse sound U (through the auxiliary sound-generating device SPDaux) to cancel the speech sound V or reduce the acoustic energy of the speech sound V. The speech cancellation circuit 52 can employ an adaptive cancellation algorithm. In addition, adaptive prediction operations known in the field of adaptive signal processing can also be performed in / by the speech cancellation circuit 52 to compensate for the delay or phase lag of the inverse sound U relative to the speech sound V.

[0036] Figure 8 illustrates a schematic diagram of a speech extraction circuit 64 according to an embodiment of this application. The speech extraction circuit 64 can be used to implement the speech extraction circuit 54, and the speech extraction circuit 64 includes a one-channel simulator 640' and a subtractor 642.

[0037] The speech extraction circuit 64 can receive an inverted signal Ud, which (assuming it is a digital signal) can pass through a practical equivalent channel 640. The practical equivalent channel 640 includes a digital-to-analog converter (D / A), an auxiliary sound-generating device SPDaux, a sound channel from the auxiliary sound-generating device SPDaux to the sound sensing device SSD, the sound sensing device SSD, an analog-to-digital converter (A / D), or a combination of these devices and channels. The practical equivalent channel 640 has a transfer function S. The output of the analog-to-digital converter can be mathematically expressed as Vd + S∙Ud, and the output signal represented by Vd + S∙Ud can be regarded as an aggregated signal.

[0038] On the other hand, the channel simulator 640' can be designed to have a transfer function S' to approximate or simulate the actual equivalent channel 640 or transfer function S, such that S-S'→0 or |S-S'|→0, where |∙| represents some norm or energy-related measure of the input parameters, and "→" means "approximate". The channel simulator 640' can receive the inverse signal Ud and output an output signal, which can be mathematically represented as S'∙Ud. The output signal represented by S'∙Ud can be regarded as an analog inverse signal corresponding to the inverse sound sensed by the sound sensing device SSD.

[0039] Subtractor 642 can subtract the analog inverse signal S'∙Ud from the signal Vd+ S∙Ud, and the result of the subtraction is Vd+ (S-S')∙Ud. Since S-S'→0, the result of the subtraction will approach Vd (i.e., Vd+ (S-S')∙Ud≈ Vd). Therefore, the speech signal Vd can be extracted from the aggregate tone U+V.

[0040] In one embodiment, the channel simulator 640' having the transfer function S' can be implemented by an infinite impulse response (IIR) digital filter whose coefficients can be obtained through software simulation tools (such as MATLAB's "System Identification" function), but is not limited thereto.

[0041] In short, in order to reduce the speech intelligibility of bystanders, the auxiliary sound-generating device SPDaux can generate anti-sound to eliminate the sound emitted by the main sound-generating device SPDmain or the user's (such as a telephone user's) voice. The auxiliary sound-generating device SPDaux can also generate masking sounds (such as noise with a specific shaped / modulated spectrum).

[0042] In addition to the spectrum masking described in Figures 3 and 4 and related paragraphs, temporal masking can also be used to protect or maintain call privacy. Temporal masking refers to increasing the auditory threshold before (also known as pre-masking) and / or after (also known as post-masking) the masking tone.

[0043] One embodiment of time masking involves generating artificial reverberation of human voice and broadcasting the artificial reverberation to the surrounding environment of the personal electronic device via an auxiliary sound-emitting device SPDaux. This artificial reverberation severely interferes with the speech recognition function in the brain of an observer, making it almost impossible for the observer to recognize what the (telephone) user is saying. Therefore, the observer's speech intelligibility will be greatly reduced.

[0044] Figure 9 illustrates a schematic diagram of a masking tone generator 72 according to an embodiment of this application. The masking tone generator 72 includes a reverberation generator 720, which in turn includes a plurality of filters c1,…,cK. In one embodiment, the filter ck can be a comb filter (e.g., k = 1,…,K), wherein the comb filter ck has an impulse response, which can be represented as hk(τk, αk). The impulse response hk(τk, αk) is parameterized by a delay factor τk and an attenuation / gain factor αk, as shown in Figure 9. More specifically, the impulse response h(τ,α) can be h(τ,α) = α ∙ δ(t-τ) + α 2 ∙ δ(t- 2τ) + α 3 ∙ δ(t- 3τ) + α 4 ∙ δ(t- 4τ) +…… (Equation 1), which can be of finite or infinite length, with the index k omitted for simplification. The (comb) filter ck produces a reverberant component rk, and the reverberant components r 1,…,rK can be combined into a reverberant signal y, where y can be expressed as y = w 1∙r 1 + … + wK∙rK, where wk represents a weighting factor. The source / input signal x can come from the telephone user's own voice, the voice of another person, or a combination of the telephone user's voice and the voice of another person.

[0045] The auxiliary sound-generating device SPDaux can generate a reverberant sound U' based on the reverberation signal y. The reverberant sound U' interferes with the bystander's brain's speech recognition of the telephone user's voice V, making it almost impossible for the bystander to recognize / decipher the telephone user's voice, thus reducing the bystander's speech clarity.

[0046] From a certain perspective, the masking sound emitted by the auxiliary sound-emitting device SPDaux contains reverberant sound U'.

[0047] Figure 10 illustrates a schematic diagram of a personal electronic device 80 according to an embodiment of this application. The personal electronic device 80 includes a speech extraction circuit 84 and a masking tone generator 82. The speech extraction circuit 84 may be implemented by the speech extraction circuit 64. The masking tone generator 82 includes a reverberation generator 820, which may be implemented by the reverberation generator 720 shown in Figure 9, or have a similar structure to the reverberation generator 720. The reverberation generator 820 generates a reverberation signal Rd, which allows the auxiliary sound-generating device SPDaux to generate a reverberation tone U' to reduce the speech intelligibility of an observer. In the personal electronic device 80, the source / input signal of the reverberation generator 820 is the voice of the telephone user. Therefore, an observer will perceive a series of reverberations from the telephone user himself, which is less annoying than the narrow-band masking noise shown in Figure 4, but due to the presence of the reverberation tone U', the observer still finds it difficult to discern the content of the telephone user's speech.

[0048] Optionally, the reverb generator 820 may receive a reverb control signal 822. In one embodiment, the reverb control signal 822 may control the volume of the reverb tone U' so that it is loud enough to disrupt the speech intelligibility of an observer without being too annoying.

[0049] In addition, to reduce the unpleasant experience for onlookers when generating reverberation, the reverberation generator can parse speech into vowel segments (or vowel phonemes) and consonant segments (or consonant phonemes), wherein the vowel segments and consonant segments correspond to different delay times and / or different repetition counts, and remix the vowel segments and consonant segments with different delay times and / or different repetition counts. Furthermore, each phoneme segment may include a rising portion and selectively include a falling portion, which can reduce the annoyance to onlookers when perceiving reverberation.

[0050] Figure 11 illustrates a schematic diagram of a reverb generator 920 according to an embodiment of this application. The reverb generator 920 includes a resolving element 903, delay elements 904 and 905, and a mixing element 906. Speech SS can be converted into a speech signal 902 by a conversion element 901, wherein the conversion element 901 may include a sound sensing device (e.g., a microphone) and / or an analog-to-digital converter. The resolving element 903 can resolve the speech signal 902 into consonant segments CNS and vowel segments VWL. In one embodiment, the resolving element 903 may include a two-way frequency divider filter centered at a center frequency fc, wherein fc may be 900 to 1,200 Hz but is not limited thereto, in order to facilitate the resolution of vowels and consonants. Delay elements 904 and 905 can apply time delays according to time delay factors Td_C (for consonant segments) and Td_V (for vowel segments), respectively. The delayed vowel / consonant segments are then remixed by mixing element 906, which may include a suitable mixer or adder for the mixing operation and generate a masking / reverb signal 907. Based on the masking / reverb signal 907, a masking tone 909 can be generated by a conversion element 908, which may include an auxiliary sound-generating device SPDaux and / or a digital-to-analog converter. Finally, the masking / reverb tone 909 and the speech SS (via a sound path 910) are mixed near the observer's ear to form an overall sound OS, thereby reducing the observer's perception of the speech SS's intelligibility.

[0051] In one embodiment, delay elements 904 and 905 may be implemented through storage devices such as first-in, first-out (FIFO) queues or buffers, and the time delay factors Td_C and Td_V are only related to the index of the FIFO buffer (similar to the address of a memory).

[0052] Figure 12(a) illustrates a timing diagram of a single segment according to an embodiment of the present application, and Figure 12(b) illustrates a timing diagram of multiple segments according to an embodiment of the present application. The vertical axis of Figure 12 is about the sound intensity (e.g., sound pressure level SPL) corresponding to a specific segment. In Figure 12(a), S may represent C (for consonants) or V (for vowels).

[0053] Data stored in the first-in-first-out (FIFO) storage device can be retrieved using addresses C#x and V#y, where x and y range from 0 to the length of the FIFO storage device minus 1. For example, if the length of the FIFO storage device is 4096, then the effective range of x and y is 0 to 4095.

[0054] It is worth noting that the data read at address C#x (V#y) corresponds to the current state of the first-in-first-out (FIFO) storage device 904 (905). That is, whenever new data enters the storage device, the data read is also updated simultaneously. Due to the FIFO characteristic, the data read at address C#x (V#y) will correspond to the CNS (VWL) generated x (y) cycles ago. For example, C#0 (V#0) will obtain the CNS (VWL) of the current cycle without delay, C#1 (V#1) will obtain the CNS (VWL) of the previous cycle (meaning a delay of 1 cycle), C#m (V#m) will obtain the CNS (VWL) of m cycles ago (meaning a delay of m cycles), and so on.

[0055] Each segment (or a segment) has time parameters: rise time tr, fall time tf, start time ts, end time te, and total time length tL. This means that each segment (or a segment) may selectively have a rising part and a falling part.

[0056] As can be seen from Figure 12(a), each phoneme segment may contain a rising portion and selectively a falling portion in order to minimize / prevent pops / clicks. The timing parameters may remain constant or may be adjusted periodically or non-periodically as needed.

[0057] Please note that in Figure 12(b), each VWL is followed by 1 to 2 CNS segments with different delays. The principle is that by moving the CNS (e.g., "s", "z", "f", "v", "th", "sh" in English) and remixing it with the VWL of different words, the human brain's speech comprehension process is severely interfered with, resulting in a serious decrease in speech clarity.

[0058] For example, in time period tx, segment C#a4 begins to decline, segments V#a2, C#a3, V#a5, C#a6, and V#a7 are at their maximum intensity, while segment C#a8 is nearing the end of its rise. In time period ty, segment C#a9 is in a half-decreasing state, while segments V#a5, V#a7, V#a10, C#a11, and C#a12 are at their maximum intensity (C#aN and V#aN are the first-in-first-out storage addresses for a specific period N, which can be based on a random number generator or a heuristic algorithm).

[0059] In addition, tw shown in Figure 12(b) represents the waiting time between subsequent segments, which can be determined according to actual needs.

[0060] In one embodiment, a segment (preferably a consonant segment) may appear multiple times (not shown in Figure 12). For example, the consonant segment C#a6 may appear (or be repeated) 5 or 6 times in the reverberation signal 907.

[0061] Since the focus of the above principle is on the adjustment of the relationship / position of CNS relative to VWL, there are more CNS segments than VWL segments in Figure 12(b). The purpose is to further confuse the observer's brain in understanding speech by randomly placing multiple copies of the constant around the vowel.

[0062] In addition to sound reflection or masking, due to the small size of the sound-emitting device, multiple auxiliary sound-emitting devices (SPDaux) can be installed on a single personal electronic device, thus enabling sound direction control (similar to beamforming) in the direction of communication. The personal electronic device can identify the position or specific angle of the bystander and form an audio beam to cancel or minimize the acoustic energy directed towards the bystander at that specific angle, thereby reducing the speech intelligibility of the bystander.

[0063] Figure 13 illustrates a schematic diagram of a personal electronic device A0 according to an embodiment of this application. The personal electronic device A0 includes a plurality of auxiliary sound-generating devices SPDaux and a direction controller A02. The auxiliary sound-generating devices SPDaux may be implemented by or include an air pulse generating device. The direction controller A02 can be used to generate a weight vector containing a plurality of weights for the plurality of auxiliary sound-generating devices SPDaux to perform audio beamforming or to form an audio beam A01. The beamforming algorithm may be referred to as an electromagnetic beamforming algorithm, which should be well known in the art and will not be described in detail herein. In one embodiment, forming the audio beam A01 can eliminate the acoustic energy of an observer A03 relative to the personal electronic device A0 at a specific angular direction, thereby reducing the observer's speech intelligibility.

[0064] In addition, the personal electronic device A0 may also include a plurality of sound sensing devices SSD, which can form a microphone array to identify the angle and direction of an observer relative to the personal electronic device.

[0065] It is worth noting that Figure 13 is only used to illustrate a personal electronic device with multiple auxiliary sound-generating devices SPDaux and multiple sound-sensing devices SSD. The arrangement of SPDaux and SSD can be designed according to actual needs and is not limited thereto.

[0066] In short, this application utilizes an auxiliary sound-generating device to produce countersound, masking sounds, or reverberation, or to perform audio beamforming, to reduce the speech intelligibility of bystanders. The above description is merely a preferred embodiment of the present invention; all equivalent variations and modifications made within the scope of the claims of this invention should be considered within the scope of this invention. [Simplified Explanation of the Diagram]

[0009] Figure 1 is a schematic diagram of a personal electronic device according to an embodiment of this application. Figure 2 is a schematic diagram of a personal electronic device according to an embodiment of this application. Figure 3 is a schematic diagram of a personal electronic device according to an embodiment of this application. Figure 4 illustrates the auditory threshold and spectrum of masking sound / noise. Figure 5 is a schematic diagram of a personal electronic device according to an embodiment of this application. Figure 6 illustrates a scene of a voice sound and an anti-sound emitted by an auxiliary sound-generating device. Figure 7 is a schematic diagram of a personal electronic device according to an embodiment of this application. Figure 8 is a schematic diagram of a voice extraction circuit according to an embodiment of this application. Figure 9 is a schematic diagram of a masking sound generator according to an embodiment of this application. Figure 10 is a schematic diagram of a personal electronic device according to an embodiment of this application. Figure 11 is a schematic diagram of a reverberation generator according to an embodiment of this application. Figure 12 illustrates a timing diagram of consonant or vowel segments according to an embodiment of this application. Figure 13 illustrates a schematic diagram of a personal electronic device according to an embodiment of this application.

Claims

1. A personal electronic device comprising: a primary sound-emitting device for emitting a predetermined sound to a predetermined user; and an auxiliary sound-emitting device for emitting a masking sound or a countersound to reduce the speech intelligibility of a bystander located near the predetermined user of the personal electronic device; wherein... The auxiliary sound-generating device includes an air pulse generating device; wherein the air pulse generating device generates a plurality of air pulses at an ultrasonic pulse rate to emit the masking sound or the anti-sound.

2. The personal electronic device as described in claim 1, wherein, The auxiliary sound-emitting device is located on the bottom or side of the personal electronic device.

3. The personal electronic device as described in claim 1, wherein, The auxiliary sound-emitting device is located on the back of the personal electronic device.

4. The personal electronic device as described in claim 1, wherein, The auxiliary sound-emitting device emits the masking sound or the countersound towards the vicinity of the personal electronic device.

5. The personal electronic device as described in claim 1, wherein, The auxiliary sound-emitting device emits the masking sound, which contains a band-limited noise.

6. The personal electronic device as described in claim 5, wherein, The noise energy of this band-limited noise is concentrated in a noise frequency band; wherein, the noise frequency band covers the sub-frequency spectrum of human speech.

7. The personal electronic device as described in claim 1, wherein, The auxiliary sound-emitting device emits the masking sound, which contains a plurality of narrow-band tones located at a plurality of masking frequency tones; wherein the plurality of masking frequency tones are distributed in a sub-audio band or a speech band.

8. The personal electronic device as described in claim 1, wherein, The auxiliary sound-emitting device emits the masking sound, which includes the reverberation of the user's voice.

9. The personal electronic device as claimed in claim 1, comprising: a masking sound generator for generating the masking sound.

10. The personal electronic device as described in claim 9, wherein, The masking sound generator includes a filter for generating a band-limited noise as a part of the masking sound; wherein the band-limited noise is concentrated in a noise frequency band; wherein the noise frequency band covers the sub-frequency spectrum of human speech.

11. The personal electronic device as described in claim 9, wherein, The masking tone generator includes a reverb generator for generating a reverb tone in speech; wherein the masking tone includes the reverb tone.

12. The personal electronic device as described in claim 11, wherein, The reverberation generator includes a plurality of filters for generating a plurality of reverberation components; wherein the plurality of reverberation components combine to form a reverberation signal, which the auxiliary sound-generating device generates the reverberation sound according to the reverberation signal.

13. The personal electronic device as claimed in claim 11, wherein the reverberation generator comprises: a resolving element for receiving a speech signal and resolving the speech signal into a plurality of consonant segments and a plurality of vowel segments; a first delay element for applying a first time delay to the plurality of consonant segments; a second delay element for applying a second time delay to the plurality of vowel segments; and a mixing element for mixing the delayed consonant segments and the delayed vowel segments to form a reverberation signal.

14. The personal electronic device as described in claim 13, wherein, The plurality of consonant segments and one of the plurality of vowel segments contain an ascending part and a descending part.

15. The personal electronic device as described in claim 11, wherein, The reverb generator receives a reverb control signal to control the volume of the reverb sound.

16. The personal electronic device as claimed in claim 1, comprising: a voice cancellation circuit for generating a countersignal for the auxiliary sound-emitting device to emit the countersignal.

17. The personal electronic device as described in claim 16, wherein, The speech cancellation circuit performs an adaptive prediction operation to generate the inverse signal.

18. The personal electronic device as claimed in claim 1, comprising: a speech extraction circuit coupled between a sound sensing device and the speech cancellation circuit, for extracting a speech signal corresponding to a spoken voice based on a converged tone sensed by the sound sensing device.

19. The personal electronic device as described in claim 18, wherein, The speech extraction circuit includes a channel simulator and a subtractor; wherein the channel simulator is used to generate an analog inverse signal, and the subtractor is used to subtract the analog inverse signal from an aggregated signal.

20. The personal electronic device as claimed in claim 1, comprising: a sensor for detecting whether the personal electronic device is attached to a user; wherein, The amplitude of the countersound is determined based on the detection results of one of the sensors.

21. A personal electronic device comprising: a plurality of auxiliary sound-generating devices for performing an audio beamforming operation and forming at least one audio beam; wherein, The at least one audio beamforming system is used to eliminate or minimize acoustic energy in an angular direction toward a bystander in order to reduce the bystander's speech intelligibility.

22. The personal electronic device as claimed in claim 21, further comprising: a direction controller for generating a weight vector for the plurality of auxiliary sound-generating devices to form the at least one audio beam.

23. The personal electronic device as claimed in claim 21, further comprising: a plurality of sound sensing devices for identifying an angular orientation of the bystander relative to the personal electronic device.

24. The personal electronic device as described in claim 21, wherein, One of the plurality of auxiliary sound-generating devices includes a gas pulse generating device.

25. A personal electronic device comprising: a primary sound-emitting device for emitting a predetermined sound to a predetermined user; and an auxiliary sound-emitting device for emitting a masking sound or a countersound to reduce the speech intelligibility of a bystander located near the predetermined user of the personal electronic device; wherein... The masking tone generator includes a reverb generator for generating a reverb tone in speech; wherein the masking tone includes the reverb tone.

26. The personal electronic device as described in claim 25, wherein, The reverberation generator includes a plurality of filters for generating a plurality of reverberation components; wherein the plurality of reverberation components combine to form a reverberation signal, which the auxiliary sound-generating device generates the reverberation sound according to the reverberation signal.

27. The personal electronic device as claimed in claim 25, wherein the reverberation generator comprises: a resolving element for receiving a speech signal and resolving the speech signal into a plurality of consonant segments and a plurality of vowel segments; a first delay element for applying a first time delay to the plurality of consonant segments; a second delay element for applying a second time delay to the plurality of vowel segments; and a mixing element for mixing the delayed consonant segments and the delayed vowel segments to form a reverberation signal.

28. The personal electronic device as described in claim 27, wherein, The plurality of consonant segments and one of the plurality of vowel segments contain an ascending part and a descending part.

29. The personal electronic device as described in claim 25, wherein, The reverb generator receives a reverb control signal to control the volume of the reverb sound.

Citation Information

Patent Citations

  • Sound masking method and device and terminal equipment

    CN116684514A

  • Fuel rod of sound device

    TWM400639U

  • Speech privacy system and / or associated method

    US20180268836A1

  • Apparatus and method for mounting a sound masking device in a hotel room

    US20210248989A1