Method for generating a spatial voice signal and a device thereof, method for receiving a spatial voice signal and a device thereof
The method generates and receives spatial voice signals by processing audio from multiple microphones to create separate channels with time delays, addressing the lack of spatial audio in teleconferences and enhancing the listening experience.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-25
- Publication Date
- 2026-04-02
AI Technical Summary
Current devices and applications do not provide spatial voice information during teleconferences or phone calls, limiting the spatial audio experience for far-end users.
A method for generating spatial voice signals using a voice pickup device that processes audio signals from multiple microphones to create separate channels with time delays based on speaker locations, and transforming these signals into a mono channel for transmission, while a voice reception device transforms the received mono channel into separate channels for spatial audio experience.
Enables spatial voice experience even in mono channel transmission scenarios like teleconferences by effectively mapping and transforming audio signals to recreate spatial audio for listeners.
Smart Images

Figure CN2024120960_02042026_PF_FP_ABST
Abstract
Description
METHOD FOR GENERATING A SPATIAL VOICE SINGAL AND A DEVICE THEREOF, METHOD FOR RECEIVING A SPATIAL VOICE SIGNAL AND A DEVICE THEREOFTECHNICAL FIELD
[0001] The disclosure relates to the technical field of spatial voice, in particular to a method for generating a spatial voice signal and a device thereof, and a method for receiving a spatial voice signal and a device thereof.BACKGROUND
[0002] There has been a lot of exploration and development in realizing spatial audio across various environments. Users expect a spatial audio listening experience with spatial audio for activities such as video games and movies. To meet these requirements, spatial audio function is implanted in a wide variety of TVs, soundbars, and home theater systems and significantly improves the sound quality and the user’s experience. As the technology of spatial audio becomes more prevalent, satisfactory listening experience is provided to users. Consequently, there is a wide demand for spatial audio to be expanded to more scenarios, such as recording voice and making calls.
[0003] However, unlike TVs, soundbars, and home theater systems, which have been carefully implanted a spatial audio function, in situations such as making a phone call or teleconference, spatial information is normally not available to far end users. Take teleconference scenario as an example, current devices and applications normally support a mono calling, such that through the current devices and applications, far-end side of the teleconference cannot receive the spatial voice of the near-end side of the teleconference. Further, the one side of the teleconference cannot notice if the talker in other side of the teleconference changes position, so no spatial information or experience could be provided.
[0004] Therefore, it’s necessary to find a way to realize spatial voice, especially for making calls or teleconference.SUMMARY
[0005] According to an aspect of the present disclosure, a method for generating a spatial voice signal performed by a voice pickup device. The method comprises: obtaining, by a plurality of microphones associated with the voice pickup device, a plurality of audio signals; determining a number of voice sources based on the plurality of audio signals; extracting, for each of the voice sources, a voice signal based on the plurality of audio signals, so as to generate a plurality of voice signals corresponding to the number of voice sources; mapping the plurality of voice signals into a first synthesis signal corresponding to a first channel of the spatial voice signal and a second synthesis signal corresponding to a second channel of the spatial voice signal, wherein for each of the voice sources, a time delay exists between the first synthesis signal and the second synthesis signal, and the time delay is associated with a relative location between the voice source and the plurality of microphones; transforming the first synthesis signal and the second synthesis signal into a mono channel signal associated with the spatial voice signal.
[0006] According to another aspect of the present disclosure, a method for receiving a spatial voice signal performed by a voice reception device. The method comprising: receiving a mono channel signal associated with the spatial voice signal through a mono channel, wherein the mono channel signal comprises a first frequency section and a second frequency section; and transforming the first and second frequency sections of the mono channel signal into a first channel signal corresponding to a first channel of the spatial voice signal and a second channel signal corresponding to a second channel of the spatial voice signal, respectively.
[0007] According to another aspect of the present disclosure, a voice pickup device for implementing spatial voice is provided. The voice pickup device may comprise: amemory; and a processor, configured to perform the method for generating a spatial voice.
[0008] According to another aspect of the present disclosure, avoice reception device for implementing spatial voice is provided. The voice reception device may comprise: a memory; areceiver, configured to receive a voice signal; and a processor, configured to perform the method for receiving a spatial voice.BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly explain the embodiments of the present disclosure or the technical solution in the prior art, the drawings needed to be used in the description of the embodiments of the present disclosure or the prior art will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments recorded in the present disclosure, and other drawings can be obtained according to these drawings of the embodiments of the present disclosure for those skilled in the art.
[0010] Fig. 1 illustrates a schematic diagram of a method for generating a first synthesis signal and a second synthesis signal from a plurality of audio signals by a voice pickup device according to at least one embodiment of the present disclosure.
[0011] Fig. 2 illustrates a schematic diagram of an arrangement of a plurality of microphones in a microphone array.
[0012] Fig. 3A illustrates a schematic diagram of transforming a first synthesis signal and a second synthesis into a mono channel signal corresponding to a spatial voice signal by a voice pickup device according to at least one embodiment of the present disclosure.
[0013] Fig. 3B illustrates a schematic diagram of transforming a mono signal associated with a voice signal into a first channel signal corresponding to a first channel of a spatial voice signal and a second channel signal corresponding to a second channel of the spatial voice signal by a voice reception device according to at least one embodiment of the present disclosure.
[0014] Fig. 4 illustrates a flowchart of a method for generating a spatial voice signal performed a voice pickup device according to at least one embodiment of the present disclosure.
[0015] Fig. 5 illustrates a flowchart of a method for receiving a spatial voice signal performed a voice reception device according to at least one embodiment of the present disclosure.
[0016] Fig. 6 shows an example configuration of a voice pickup device according to at least one embodiment of the present disclosure.
[0017] Fig. 7 shows an example configuration of a voice reception device according to at least one embodiment of the present disclosure.DETAILED DESCRIPTION
[0018] In order to provide a clearer and more complete description of the purpose, technical solution, and advantages of the present disclosure, the following description, in conjunction with the accompanying drawings, will provide a clear and comprehensive understanding of the technical solution in the present disclosure. It should be noted that the described embodiments are only a part of the embodiments disclosed herein, and not the entire embodiments. All other embodiments that ordinary skilled persons in the art can obtain without exercising inventive labor based on the embodiments disclosed herein are within the scope of the present disclosure.
[0019] The terms “first, ” “second, ” “third, ” “fourth, ” etc. (if present) used in the specification and claims, as well as in the accompanying drawings, are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the use of such data can be interchangeable in appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than that shown or described here.
[0020] It should be understood that the numbering of the processes in various embodiments of the present disclosure does not imply a specific order of execution. The execution order of the processes should be determined based on their functionality and inherent logic, and should not impose any limitations on the implementation process of the embodiments of the present disclosure.
[0021] It should be understood that the terms “comprising” and “having” and their variations intend to cover non-exclusive inclusion, such as a process, method, system, product, or apparatus that includes a series of steps or units, not necessarily limited to those explicitly listed, but may include other steps or units that are inherently present in these processes, methods, products, or apparatus.
[0022] It should be understood that the term “multiple” means two or more. The term “and / or” is merely a description of the associated relationship between related objects, indicating that there can be three possible relationships. For example, “Aand / or B” can mean: A exists alone, A and B exist simultaneously, or B exists alone. The character “ / ” generally indicates an “or” relationship between the preceding and following related objects. “Including A, B, and C, ” “including A, B, C” means that A, B, and C are all included, and “including A, B, or C” means that one of A, B, and C is included. “Including A, B and / or C” means that any one or two or all three of A, B, and C are included.
[0023] It should be understood that “corresponding to B with A, ” “corresponding to A with B, ” “A corresponds to B, ” or “B corresponds to A” means that B is associated with A and can be determined based on A. Determining B based on A does not mean that B can only be determined based on A, but can also be determined based on A and / or other information. The matching of A and B means that the similarity between A and B is greater than or equal to a predetermined threshold.
[0024] Depending on the context, the term “if” used herein can be interpreted as “when” or “in response to determining” or “in response to detecting. ”
[0025] The following specific embodiments will provide a detailed description of the technical solution of the present disclosure. These specific embodiments can be combined with each other, and certain concepts or processes may not be reiterated in some embodiments if they are the same or similar. In order to provide a clearer understanding of the purpose, technical solution, and advantages of the present disclosure, the following description will be provided in conjunction with the accompanying drawings.
[0026] Fig. 1 illustrates a schematic diagram of a method for generating a first synthesis signal and a second synthesis signal from a plurality of audio signals by a voice pickup device according to at least one embodiment of the present disclosure.
[0027] Referring to Fig. 1, a scenario where two speakers are speaking at the same time is illustrated. However, depend on different scenarios, the number of speakers may vary. As shown in Fig. 1, a microphone array may scan the whole azimuthal space and obtain a plurality of audio signals from the whole azimuthal space. The microphone array may include a plurality of microphones. In the scenario of Fig, 1, the microphone array may include three microphones. However, depend on different scenarios, the number of microphones included in the microphone array can be three or more and the microphone array will be described in detail with reference to Fig. 2 later. The plurality of audio signals obtained by the microphone array include voices of the two speakers. To generate a spatial voice signal, the plurality of audio signals obtained by the microphone array should be processed into a first synthesis signal corresponding to a first channel of the spatial voice signal and a second synthesis signal corresponding to a second channel of the spatial voice signal. Although Fig. 1 shows that the audio signals are obtained by the microphone array, it should be understood that the audio signals can also be obtained by a plurality of independent microphones instead of the form the microphone array.
[0028] The processing of the plurality of audio signals may include determining / estimating the number of voice signals, extracting voice signals corresponding to different speakers, enhancing the extracted voice signals and mapping the enhanced voice signals into the first synthesis signal and the second synthesis signal.
[0029] Upon receiving the plurality of audio signals, a pre-processing will be performed on the plurality of audio signals. In an example, the pre-processing may include pre-emphasis, framing, windowing. However, the present disclosure is not limited to the mentioned pre-processing and other pre-processing can be performed as needed.
[0030] In an example, voiceprint recognition will be used to determine the number of voice sources based on the pre-processed audio signals. The vocal organs used by people when speaking -each person varies greatly in size and shape, so any two people's voiceprints will be different. Therefore, voiceprint recognition is convenient to determine the number of voice sources from the audio signals. However, the present disclosure is not limited thereto, and any other suitable methods may be used to determine the number of voice sources. Also, when the number of speakers is less than the number of microphones, the number of large eigen values of the covariance matrix is usually used as the number of source signals (speakers) . Then for each voice source, a corresponding voice signal is extracted. In an example of two speakers, the two speakers may correspond to two voice sources and accordingly two voice signals will be extracted.
[0031] In an example, beamforming may be used to enhancing the plurality of voice signals. For example, the beamforming may be based on Angle of Arrival (AOA) of each of the plurality of voice signals.
[0032] As shown in Fig. 1, when there are two speakers and two voice signals corresponding to the two speakers are extracted, AOA will be estimated for each voice signal. The estimated AOA of voice signal of speaker 1 may be denoted as θ1, and the estimated AOA of voice signal of speaker 2 may be denoted as θ2. In an example, the θ1 and θ2 may be estimated by temporal-spectral analysis. For example, when estimating the AOA θ1 of the voice signal of speaker 1, assuming that the microphone array includes three microphones, time differences between the voice signal of speaker 1 reaching different microphones are calculated. Based on the time differences, the θ1 of the voice signal of speaker 1 may be estimated. In a same manner, the θ2 of the voice signal of speaker 2 may also be estimated. However, the method for estimating AOA for each voice signal is not limited to temporal-spectral analysis, and any suitable algorithms may be adopted for the estimation. For example, the high-resolution method, such as MUltiple SIgnal Classification (MUSIC) could be considered.
[0033] The plurality of voice signals may be enhanced using beamforming. In the scenario shown in Fig. 1, when two speakers are speaking and two voice signals are extracted with respect to the two speakers, for the voice signal with an AOA θ1, the enhanced voice signal may be represented as y (θ1) =wH (θ1) X1 , where wH (θ1) is the conjugate transposition of w (θ1) , which represents the set of weights for beamforming regarding the AOA θ1, X1 is a vector of the voice signal received on the microphone array. In an example, when the microphone array includes three microphones, X1 may be a 3×1 vector and wH (θ1) may be a 1×3 vector. The wH (θ1) is based on θ1 and may be designed to enhance certain sections of the voice signal associated with a first microphone of the three microphones and suppress remaining sections of the voice signal associated with a second and a third microphone of the three microphones. Similarly, for the voice signal with an AOA θ2, the enhanced voice signal may be represented as y (θ2) =wH (θ2) X2 , wH (θ2) is the transposition of w (θ2) , which represents the set of weights for beamforming regarding the AOA θ2 , X2 is a vector of the voice signal. In an example, when the microphone array includes three microphones, X2 may be a 3×1 vector and wH (θ2) may be a 1×3 vector. The wH (θ2) is based on θ2 and may be designed to enhance certain portions of the voice signal associated with the second microphone of the three microphones and suppress remaining portions of the voice signal associated with a the first and the third microphone of the three microphones. By performing beamforming on the extracted voice signals to generated enhanced voice signals, the quality of the voice signals is improved.
[0034] Then the enhanced voice signals are mapped into the first synthesis signal and the second synthesis signal. In the scenario shown in Fig. 1, when two speakers are speaking, and two enhanced signalsy (θ1) and y (θ2) are generated. The first synthesis signal may be represented as yL and calculated as that is, the first synthesis signal is calculated as y (θ1) +y (θ2) . The second synthesis signal may be represented as yR and calculated as that is, the second synthesis signal is calculated as wherein d is the distance between left ear and right ear of a user, or the distance between left ear cup and right ear cup of the voice reception device, λ is wavelength for each frequency bin.
[0035] As a result, the plurality of voice signals are mapped into the first synthesis signal yL and the second synthesis signal yR. The first synthesis signal yL corresponds to a first channel of the spatial voice and the second synthesis signal yR corresponds to a second channel of the spatial voice. Through the above mapping, there is a corresponding time delay for each voice between the first synthesis signal yL and the second synthesis signal yR. When the first synthesis signal yL and the second synthesis signal yR are heard by a listener via his / her left ear and right ear respectively, the time delay between the two synthesis signals will make the listener experience the spatial voice.
[0036] However, the current devices only support mono channel transmission. For example, mono-phone call system or teleconference system does not support dual channel transmission, therefore, the listener is still unable to feel the spatial voice experience during making a call or teleconference. A background is that the current call system or network have supported high resolution call, which provides much wide bandwidth than human voice. To realize the spatial voice, after mapping the plurality of voice signals into two synthesis signals with corresponding time delays, transforming the two synthesis signals into a mono channel signal associated with the spatial voice signal is further needed and the method for generating the mono channel signal will be described with reference to Fig. 3A later.
[0037] Fig. 2 illustrates a schematic diagram of an arrangement of a plurality of microphones in a microphone array.
[0038] Referring to Fig. 2, a microphone array is shown. In an example, the microphone array is an additive microphone array. In another example, the microphone array is a differential microphone array.
[0039] As shown in Fig. 2, the microphone array includes three microphones, however, the number of microphones included in the microphone array can be pre-configured as needed. In an example, to generate the spatial voice, at least three microphones are included in the microphone array. In an example, the number of microphones included in the microphone array may depend on the number of speakers. When there are four speakers, a microphone array including at least four microphones is needed. The number of microphones included in the microphone array is at least three and always no less than the number of speakers.
[0040] Fig. 3A illustrates a schematic diagram of transforming a first synthesis signal and a second synthesis into a mono channel signal corresponding to a spatial voice signal by a voice pickup device according to an embodiment of the present disclosure.
[0041] Referring to Fig. 3A, the first synthesis signal yL is inputted to a low-pass filter and Fourier transformed to generate a filtered first synthesis signal yL_LP. Similarly, the second synthesis signal yR is inputted to the low-pass filter and Fourier transformed to generate a filtered second synthesis signal yR_LP. In an example, the low-pass filter may have a cutoff frequency of 8kHz, since human voice is mainly below 8 kHz and signals above 8 kHz contain little information. However, the cutoff frequency of the low-pass filter may be set differently depending on different scenarios.
[0042] The filtered first synthesis signal yL_LP may be further modulated to generate a modulated first synthesis signal yL_LP_MOD . For example, when both the filtered first synthesis signal yL_LP and the filtered second synthesis signal yR_LP have a central frequency of 0kHz and a bandwidth of 16kHz, the 0kHz to 8kHz section of the filtered first synthesis signal yL_LP is shifted to 8kHz to 16kHz, and the -8kHz to 0kHz section of the filtered first synthesis signal yL_LP is shifted to -16kHz to 8kHz. Therefore, the modulated first synthesis signal yL_LP_MOD and the filtered second synthesis signal yR_LP may not overlap in frequency. In another example, the filtered second voice signal yR_LP may be modulated.
[0043] After the modulation, the modulated first synthesis signal yL_LP_MOD and the filtered second synthesis signal yR_LP may be combined as the mono channel signal yTxcorresponding to the spatial voice signal for mono channel transmission.
[0044] Through the mapping and the transformation, the mono channel signal containing spatial information is obtained. Therefore, even for mono channel transmission, the listener may experience spatial voice.
[0045] Fig. 3B illustrates a schematic diagram of transforming a mono signal associated with a voice signal into a first channel signal corresponding to a first channel of a spatial voice signal and a second channel signal corresponding to a second channel of the spatial voice signal by a voice reception device according to at least one embodiment of the present disclosure.
[0046] Referring to Fig. 3B, the voice reception device may receive a mono channel signal yRx through mono channel transmission, theyRx corresponds to a spatial voice signal. The mono channel signal yRx comprises a first frequency section and a second frequency section.
[0047] The mono channel signal yRx is inputted into the high-pass filter to generate a filtered first signal yRx_HP. For example, the high-pass filter may have a cutoff frequency of 8kHz and the filtered first signal yRx_HP comprises the first frequency section. The mono channel signal yRx is also inputted into the low-pass filter to generate a second channel signal yRx_LP. For example, the low-pass filter may have a cutoff frequency of 8kHz and the second channel signal yRx_LP comprises the second frequency section.
[0048] The filtered first signal yRx_HP may be further demodulated to generate a first channel signal yRx_HP_MOD, the first channel signal yRx_HP_MOD is overlapped with the second channel signal yRx_LP in frequency.
[0049] By performing the above filtering and demodulation, the mono channel signal yRx is transformed into the first channel signal yRx_HP_MOD corresponding to the first channel of the spatial voice signal and the second channel signal yRx_LP corresponding to the second channel of the spatial voice signal. When the two synthesis signals heard by the left and right ear of the user respectively, the time differences for voice sound at different AOAs exists betweenyRx_HP_MOD andyRx_LP may make the user experience spatial voice. Before heard by the user, an inverse Fourier transform will be performed on the first channel signal yRx_HP_MoD and the second channel signal yRx_LP.
[0050] Fig. 4 illustrates a flowchart of a method for generating a spatial voice signal performed a voice pickup device according to at least one embodiment of the present disclosure.
[0051] At step 402, the voice pickup device may obtain a plurality of audio signals.
[0052] In an example, a plurality of microphones or a microphone array of the voice pickup device may obtain the plurality of audio signals. In another example, the voice pickup device may not comprise microphones or microphone, and a plurality of microphones or a microphone array associated with the voice pickup device may obtain the plurality of audio signals.
[0053] At step 404, the voice pickup device may determine a number of voice sources based on the plurality of audio signals.
[0054] The plurality of audio signals obtained in step 402 may contain all sounds in the azimuthal space, including voices from different people (different sound sources) . To determine the number of sound sources, voiceprint recognition will be used. However, the present disclosure is not limited thereto, and any other suitable methods for determining the number of sound sources may be used.
[0055] At step 406, the voice pickup device may extract a voice signal for each voice sources. In an example, when five sound sources are determined, then for each of the five sound sources, extract a voice signal.
[0056] At step 408, the voice pickup device may map the plurality of voice signals into the first synthesis signal and the second synthesis signal. In an example, for each of the plurality of the voice signals, generating a first signal corresponding to a first channel of the spatial voice signal from the voice signal based on beamforming, and generating a second signal corresponding to a second channel of the spatial voice signal by delaying the first signal a certain time, thus a time delay exists between the first signal and the second signal and the time delay is associated with a relative location between the voice source and the plurality of microphones or the microphone array. The first signals and the second signals for the plurality of voice signals are combined to generate the first synthesis signal corresponding to the first channel of the spatial voice signal and the second synthesis signal corresponding to the second channel of the spatial voice signal. The time delay between the two synthesis signals will make the listener experience the spatial voice when heard by the listener.
[0057] At step 410, the voice pickup device may transform the first synthesis signal and the second synthesis signal into a mono channel signal for mono channel transmission.
[0058] In an example, the first synthesis signal and the second synthesis signal are inputted into a low-pass filter to generate the filtered first synthesis signal and the filtered second synthesis signal. Then either the filtered first synthesis signal or the filtered second synthesis signal may be modulated to generate a modulated synthesis signal that is not overlapped with synthesis signal without modulation in frequency. In an example, the first synthesis signal is modulated as a modulated first synthesis signal and combine the modulated first synthesis signa and the filtered second synthesis signal to generate the mono channel signal. In another example, the filtered second synthesis signal is modulated as a modulated second synthesis signal and combine the modulated second synthesis signa and the filtered first synthesis signal to generate the mono channel signal.
[0059] Fig. 5 illustrates a flowchart of a method for receiving a spatial voice signal performed a voice reception device according to an embodiment of the present disclosure.
[0060] At step 502, the voice reception device may receive a mono channel signal associated with the spatial voice signal through a mono channel. The mono channel comprises a first frequency section and a second frequency section.
[0061] At step 504, the voice reception device may transform the mono channel signal into a first channel signal corresponding to a first channel of the spatial voice signal and a second channel signal corresponding to a second channel of the spatial voice signal, respectively. As described with reference to Fig. 3B, filtering and modulation will be applied to the mono channel signal to generate the first channel signal and the second channel signal.
[0062] In an example, the mono channel signal is inputted into a high-pass filter and a low-pass filter separately, to generate a first signal comprises the first frequency section and the second channel signal comprises the second frequency section. Then the first signal will be further modulated to generate the first channel signal, such that the first channel signal is overlapped with the second channel signal in frequency. The first channel signal and the second channel signal will also be inverse Fourier transformed.
[0063] Fig. 6 shows an example configuration of a voice pickup device according to at least one embodiment of the present disclosure.
[0064] As shown in Fig. 6, the voice pickup device 600 includes a memory 610, and a processor 620. The memory 610 is used to store non-transitory computer-readable instructions (e.g., one or more computer program modules) . In an example, the memory 610 is further used to store voice signals. The processor 620 is used to execute non-transitory computer-readable instructions, which, when executed by the processor 620, can execute the audio processing method described in Fig. 4. The memory 610 and the processor 620 may be interconnected through a bus system and / or other forms of connection mechanisms (not shown) . In an example, the voice pickup device 600 may include a microphone array (or a plurality of microphones) . The microphone array (or the plurality of microphones) is used to obtain audio signals. In an example, the voice pickup device may further include a transmitter. The transmitter is used to transmit the mono channel signal.
[0065] For example, the processor 620 may be a central processing unit (CPU) , graphics processing unit (GPU) , or other forms of processing units with data processing capability and / or program executing capability. For example, the central processing unit (CPU) may be X86 or ARM architecture. The processor 620 may be a general-purpose or special-purpose processor and may control other components in the voice pickup device 600 to perform desired functions.
[0066] For example, the memory 610 may include any combination of one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or nonvolatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache, etc. The non-volatile memory may include, for example, read-only memory (ROM) , hard disk, erasable programmable read-only memory (EPROM) , portable compact disk read-only memory (CD-ROM) , USB memory, flash memory, etc. One or more computer program modules may be stored on a computer-readable storage medium, and the processor 620 may execute one or more computer program modules to implement various functions of the voice pickup device 600. The computer-readable storage medium may also store various application programs and various data, as well as various data used and / or generated by the application programs, etc.
[0067] It should be noted that, in the at least one embodiments of the present disclosure, the specific functions and technical effects of the voice pickup device 600 can refer to the above description with regard to the communication method and will not be detailed here.
[0068] Fig. 7 shows an example configuration of a voice reception device according to at least one embodiment of the present disclosure.
[0069] As shown in Fig. 7, the voice reception device 700 includes a memory 710, a receiver 720 and a processor 730. The memory 710 is used to store non-transitory computer- readable instructions (e.g., one or more computer program modules) . In an example, the memory 710 is further used to store voice signals. The processor 730 is used to execute non-transitory computer-readable instructions, which, when executed by the processor 730, can execute the audio processing method described in Fig. 5. The memory 710 and the processor 730 may be interconnected through a bus system and / or other forms of connection mechanisms (not shown) . The receiver 720 is used to receive voice signal.
[0070] For example, the processor 730 may be a central processing unit (CPU) , graphics processing unit (GPU) , or other forms of processing units with data processing capability and / or program executing capability. For example, the central processing unit (CPU) may be X86 or ARM architecture. The processor 730 may be a general-purpose or special-purpose processor and may control other components in the voice reception device 700 to perform desired functions.
[0071] For example, the memory 710 may include any combination of one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or nonvolatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache, etc. The non-volatile memory may include, for example, read-only memory (ROM) , hard disk, erasable programmable read-only memory (EPROM) , portable compact disk read-only memory (CD-ROM) , USB memory, flash memory, etc. One or more computer program modules may be stored on a computer-readable storage medium, and the processor 730 may execute one or more computer program modules to implement various functions of the voice reception device 700. The computer-readable storage medium may also store various application programs and various data, as well as various data used and / or generated by the application programs, etc.
[0072] It should be noted that, in the embodiments of the present disclosure, the specific functions and technical effects of the voice reception device 700 can refer to the above description with regard to the communication method and will not be detailed here.
[0073] According to another aspect of the present disclosure, a computer-readable storage medium for storing a computer-readable program is provided, when being executed by a computer or a processor, the computer-readable program causes the computer or the processor to perform the sound reproduction method as described above.
[0074] Techniques of this disclosure may include additional aspects, such as any single aspect or any combination of aspects described below and / or in connection with one or more other processes described elsewhere herein.
[0075] In a first aspect, a method for generating a spatial voice signal performed by a voice pickup device is provided, which comprises: obtaining, by a plurality of microphones associated with the voice pickup device, a plurality of audio signals; determining a number of voice sources based on the plurality of audio signals; extracting, for each of the voice sources, a voice signal based on the plurality of audio signals, so as to generate a plurality of voice signals corresponding to the number of voice sources; mapping the plurality of voice signals into a first synthesis signal corresponding to a first channel of the spatial voice signal and a second synthesis signal corresponding to a second channel of the spatial voice signal, wherein for each of the voice sources, a time delay exists between the first synthesis signal and the second synthesis signal, and the time delay is associated with a relative location between the voice source and the plurality of microphones; transforming the first synthesis signal and the second synthesis signal into a mono channel signal associated with the spatial voice signal.
[0076] In a second aspect, alone or in combination with any other aspect, the method for generating a spatial voice signal further comprises: estimating an Angle of Arrival (AOA) of each of the plurality of voice signals, wherein the AOA of each voice signal indicates orientation of each voice source relative to the plurality of microphones; and enhancing each of the plurality of voice signals based on a corresponding AOA to generate a plurality of enhanced voice signals, wherein, for each of the plurality of voice signals, the relative location between the voice source and the plurality of microphones is based on the corresponding AOA.
[0077] In a third aspect, alone or in combination with any other aspect, estimating an Angle of Arrival (AOA) of each of the plurality of voice signals comprises: estimating the AOA of each of the plurality of voice signals based on temporal spectral analysis.
[0078] In a fourth aspect, alone or in combination with any other aspect, enhancing each of the plurality of voice signals based on the corresponding AOA comprises: for each of the plurality of voice signals, enhancing the voice signal based on a set of weights, wherein the set of weights is based on the corresponding AOA.
[0079] In a fifth aspect, alone or in combination with any other aspect, mapping the plurality of voice signals into a first synthesis signal and a second synthesis signal comprises: accumulating the plurality of enhanced voice signals to generate the first synthesis signal; delaying each of the plurality of enhanced voice signals based on the corresponding AOA, so as to generate a plurality of delayed voice signals; and accumulating the plurality of delayed voice signals to generate the second synthesis signal.
[0080] In a sixth aspect, alone or in combination with any other aspect, the plurality of microphones are comprised in a microphone array, and wherein obtaining a plurality of voice signals comprises: scanning a whole azimuthal space by the microphone array.
[0081] In a seven aspect, alone or in combination with any other aspect, the microphone array comprises at least three microphones.
[0082] In an eighth aspect, alone or in combination with any other aspect, transforming the first synthesis signal and the second synthesis signal into a mono channel signal associated with the spatial voice signal comprises: filtering the first synthesis signal and the second synthesis signal, via a low-pass filter, to generate a filtered first synthesis signal and a filtered second synthesis signal; modulating the filtered first synthesis signal to generate a modulated first synthesis signal, wherein the modulated first synthesis signal is not overlapped with the filtered second synthesis signal in frequency; and adding the modulated first synthesis signal and the filtered second synthesis signal to generate the mono channel signal.
[0083] In a ninth aspect, alone or in combination with any other aspect, the low-pass filter has a cutoff frequency of 8kHz.
[0084] In a tenth aspect, alone or in combination with any other aspect, the method for generating a spatial voice signal further comprises transmitting the mono channel signal through a mono channel.
[0085] In an eleventh aspect, alone or in combination with any other aspect, the time delay existing between the first synthesis signal and the second synthesis signal provides spatial voice experience, and wherein for each of the voice source, the time delay is determined based on an Angle of Arrival (AOA) of the voice signal corresponding to the voice source and distance between left ear and right ear of a user or distance between left ear cup and right earcup of a voice reception device.
[0086] In a twelfth aspect, alone or in combination with any other aspect, a method for receiving a spatial voice signal performed by a voice reception device is provided, which comprises: receiving a mono channel signal associated with the spatial voice signal through a mono channel, wherein the mono channel signal comprises a first frequency section and a second frequency section; and transforming the first frequency section and the second frequency section of the mono channel signal into a first channel signal corresponding to a first channel of the spatial voice signal and a second channel signal corresponding to a second channel of the spatial voice signal, respectively.
[0087] In a thirteenth aspect, alone or in combination with any other aspect, transforming the first frequency section and the second frequency section of the mono channel signal into a first channel signal and a second channel comprises: filtering the mono channel signal, via a high-pass filter, to generate a filtered first signal comprising the first frequency section; filtering the mono channel signal, via a low-pass filter, to generate the second channel signal comprising the second frequency section; and modulating the filtered first signal to generate the first channel signal, wherein the modulated first channel signal is overlapped with the second channel signal in frequency.
[0088] In a fourteenth aspect, alone or in combination with any other aspect, wherein the low-pass filter has a cutoff frequency of 8kHz, and wherein the high-pass filter has a cutoff frequency of 8kHz.
[0089] In a fifteenth aspect, alone or in combination with any other aspect, the method for receiving a spatial voice signal further comprises: outputting the first channel signal and the second channel signal to a user’s left ear and right ear respectively.
[0090] In a sixteenth aspect, alone or in combination with any other aspect, a time delay existing between the first channel signal and the second channel signal provides spatial voice experience.
[0091] In a seventeenth aspect, alone or in combination with any other aspect, a voice pickup device for generating a spatial voice signal is provided, which comprises: amemory; and a processor, configured to perform the method according to any of previous aspects.
[0092] In an eighteenth aspect, alone or in combination with any other aspect, a voice reception device for receiving a spatial voice signal is provided, which comprises: amemory; a receiver, configured to receive a mono channel signal; and a processor, configured to perform the method according to any of previous aspects.
[0093] Those of skill would appreciate that the logical blocks, circuits, and algorithm steps described here may be implemented as electronic hardware, computer software, or a combination. This interchangeability of hardware and software is shown by the illustrative components described functionally. Whether the functionality is implemented in hardware or software depends on the application and constraints. Experts may implement the functionality in various ways for each application, but those choices do not depart from the scope here. Experts also recognize the examples of components, methods, and interactions here are merely illustrative; the components, methods, or interactions may be combined or performed differently.
[0094] The illustrative logic, blocks, circuits, and processes described may be implemented as hardware, software, or a combination. This hardware and software interchangeability have been described generally in terms of functionality and illustrated in the components, blocks, circuits, and processes. Whether the functionality is implemented in hardware or software depends on the application and constraints.
[0095] It will be appreciated by a person skilled in the art, aspects of the present disclosure may be illustrated and described herein in any of a number of patentable classes or context including any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof. Accordingly, aspects of the present disclosure may be implemented entirely hardware, entirely software (including firmware, resident software, micro-code, etc. ) or combining software and hardware implementation that may all generally be referred to herein as a “data block” , “circuit” , “engine” , “unit” , “circuit” or “system” . Furthermore, aspects of the present disclosure may take the form of a computer program product embodied in one or more computer-readable media having computer-readable program code embodied thereon.
[0096] Certain terminology has been used to describe embodiments of the present disclosure. For example, the terms “first / second embodiment” , “one embodiment” , “an embodiment” , and / or “some embodiments” mean that a particular feature, structure or characteristic described in connection with the embodiment is included in at least one embodiment of the present disclosure. Therefore, it is emphasized and should be appreciated that two or more references to “an embodiment” or “one embodiment” or “an alternative embodiment” in various portions of this specification are not necessarily all referring to the same embodiment. Furthermore, the particular features, structures or characteristics may be combined as suitable in one or more embodiments of the present disclosure.
[0097] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having the meaning that is consistent with their meaning in the context of the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.
[0098] The above is illustration of the present disclosure and should not be construed as making limitation thereto. Although some exemplary embodiments of the present disclosure have been described, a person skilled in the art can easily understand that many modifications may be made to these exemplary embodiments without departing from the creative teaching and advantages of the present disclosure. Therefore, all such modifications are intended to be included within the scope of the present disclosure as defined by the appended claims. As will be appreciated, the above is to explain the present disclosure, it should not be constructed as limited to the specific embodiments disclosed, and modifications to the present disclosure and other embodiments are included in the scope of the attached claims. The present disclosure is defined by the claims and their equivalents.
Claims
1.A method for generating a spatial voice signal performed by a voice pickup device, the method comprising:obtaining, by a plurality of microphones associated with the voice pickup device, a plurality of audio signals;determining a number of voice sources based on the plurality of audio signals;extracting, for each of the voice sources, a voice signal based on the plurality of audio signals, so as to generate a plurality of voice signals corresponding to the number of voice sources;mapping the plurality of voice signals into a first synthesis signal corresponding to a first channel of the spatial voice signal and a second synthesis signal corresponding to a second channel of the spatial voice signal, wherein for each of the voice sources, a time delay exists between the first synthesis signal and the second synthesis signal, and the time delay is associated with a relative location between the voice source and the plurality of microphones;transforming the first synthesis signal and the second synthesis signal into a mono channel signal associated with the spatial voice signal.2.The method for generating a spatial voice signal according to claim 1, the method further comprising:estimating an Angle of Arrival (AOA) of each of the plurality of voice signals, wherein the AOA of each voice signal indicates orientation of each voice source relative to the plurality of microphones; andenhancing each of the plurality of voice signals based on a corresponding AOA to generate a plurality of enhanced voice signals,wherein, for each of the plurality of voice signals, the relative location between the voice source and the plurality of microphones is based on the corresponding AOA.3.The method for generating a spatial voice signal according to claim 2, wherein estimating an Angle of Arrival (AOA) of each of the plurality of voice signals comprises:estimating the AOA of each of the plurality of voice signals based on temporal spectral analysis.4.The method for generating a spatial voice signal according to claim 2, wherein enhancing each of the plurality of voice signals based on the corresponding AOA comprises:for each of the plurality of voice signals, enhancing the voice signal based on a set of weights, wherein the set of weights is based on the corresponding AOA.5.The method for generating a spatial voice signal according to claim 2, wherein mapping the plurality of voice signals into a first synthesis signal and a second synthesis signal comprises:accumulating the plurality of enhanced voice signals to generate the first synthesis signal;delaying each of the plurality of enhanced voice signals based on the corresponding AOA, so as to generate a plurality of delayed voice signals; andaccumulating the plurality of delayed voice signals to generate the second synthesis signal.6.The method for generating a spatial voice signal according to claim 1, wherein the plurality of microphones are comprised in a microphone array, and wherein obtaining a plurality of voice signals comprises:scanning a whole azimuthal space by the microphone array.7.The method for generating a spatial voice signal according to claim 6, wherein the microphone array comprises at least three microphones.8.The method for generating a spatial voice signal according to claim 1, wherein transforming the first synthesis signal and the second synthesis signal into a mono channel signal associated with the spatial voice signal comprises:filtering the first synthesis signal and the second synthesis signal, via a low-pass filter, to generate a filtered first synthesis signal and a filtered second synthesis signal;modulating the filtered first synthesis signal to generate a modulated first synthesis signal, wherein the modulated first synthesis signal is not overlapped with the filtered second synthesis signal in frequency; andadding the modulated first synthesis signal and the filtered second synthesis signal to generate the mono channel signal.9.The method for generating a spatial voice signal according to claim 8, wherein the low-pass filter has a cutoff frequency of 8kHz.10.The method for generating a spatial voice signal according to claim 1, wherein the method further comprising:transmitting the mono channel signal through a mono channel.11.The method for generating a spatial voice signal according to claim 1, wherein the time delay existing between the first synthesis signal and the second synthesis signal provides spatial voice experience, andwherein for each of the voice source, the time delay is determined based on an Angle of Arrival (AOA) of the voice signal corresponding to the voice source and distance between left ear and right ear of a user or distance between left ear cup and right earcup of a voice reception device.12.A method for receiving a spatial voice signal performed by a voice reception device, the method comprising:receiving a mono channel signal associated with the spatial voice signal through a mono channel, wherein the mono channel signal comprises a first frequency section and a second frequency section; andtransforming the first frequency section and the second frequency section of the mono channel signal into a first channel signal corresponding to a first channel of the spatial voice signal and a second channel signal corresponding to a second channel of the spatial voice signal, respectively.13.The method for receiving a spatial voice signal according to claim 12, wherein transforming the first frequency section and the second frequency section of the mono channel signal into a first channel signal and a second channel comprises:filtering the mono channel signal, via a high-pass filter, to generate a filtered first signal comprising the first frequency section;filtering the mono channel signal, via a low-pass filter, to generate the second channel signal comprising the second frequency section; andmodulating the filtered first signal to generate the first channel signal, wherein the modulated first channel signal is overlapped with the second channel signal in frequency.14.The method for receiving a spatial voice signal according to claim 13, wherein the low-pass filter has a cutoff frequency of 8kHz, and the high-pass filter has a cutoff frequency of 8kHz.15.The method for receiving a spatial voice signal according to claim 12, the method further comprising:outputting the first channel signal and the second channel signal to a user’s left ear and right ear respectively.16.The method for receiving a spatial voice signal according to claim 12, wherein a time delay existing between the first channel signal and the second channel signal provides spatial voice experience.17.A voice pickup device for generating a spatial voice signal, the voice pickup device comprising:a memory; anda processor, configured to perform the method according to any of claims 1-11.18.A voice reception device for receiving a spatial voice signal, the voice reception device comprising:a memory;a receiver, configured to receive a mono channel signal; anda processor, configured to perform the method according to any of claims 12-16.
Citation Information
Patent Citations
Implementation method of 3D audio
US20030202665A1
Multiplexing audio system and method
US20140314238A1
System and apparatus for tracking moving audio sources
WO2017129239A1