Voice signal processing method and device, electronic equipment and storage medium
Through the voice signal processing method based on acoustic features, different voice objects in multi-person calls are separated and positioned, and the audio-visual positioning output is used using head-related transfer functions, which solves the problem of missing directional sense of voice output in multi-person calls, and improves communication efficiency and listening experience.
Patent Information
- Application Number
- CN202311549391.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-20
- Publication Date
- 2025-05-20
AI Technical Summary
In multi-person online call scenarios, the prior art is difficult to provide a sense of direction and clear voice output, resulting in reduced auditory fatigue and communication efficiency.
Through the recognition method based on acoustic features, different vocal objects in the speech signal are separated and positioned, and the acoustic positioning output is performed using head-related transfer functions and azimuth positioning information.
The sound image orientation positioning of the voices of different voice objects is realized, which improves the user's listening experience and is suitable for scenes where the voice object cannot be distinguished in multi-person conversation scenarios.
Smart Images

Figure CN120020944A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of voice technologies, and in particular, to a method and apparatus for processing voice signals, an electronic device, and a storage medium. Background Art
[0002] Multi-person online calls are involved in scenarios such as online meetings, online games, and virtual reality (VR) games. Usually, each participant mostly uses a headset or a simple speaker and relatively simple audio hardware such as a single microphone for uplink and downlink calls. In this case, for each listener, the voices of others heard by them do not bring a sense of direction. Especially when others speak simultaneously, the sound images will overlap, resulting in auditory fatigue, chaotic listening experience, etc., affecting the communication efficiency. Summary of the Invention
[0003] The present disclosure provides a method and apparatus for processing voice signals, an electronic device, and a storage medium.
[0004] According to a first aspect of an embodiment of the present disclosure, there is provided a method for processing a voice signal, including:
[0005] When a voice signal is received, identifying a voice object in the voice signal based on acoustic features;
[0006] In response to different voice objects being included in the voice signal, performing sound image localization on the voices of the different voice objects and then outputting them; wherein, the sound image azimuths of the voices of different voice objects are different.
[0007] In some embodiments, the identifying a voice object in the voice signal based on acoustic features includes:
[0008] Performing voice separation on the voice signal based on acoustic features;
[0009] In response to different voices being separated in the voice signal, determining that different voice objects are included in the voice signal.
[0010] In some embodiments, the performing voice separation on the voice signal based on acoustic features includes:
[0011] Detecting a moment when the voice object changes in the voice signal based on acoustic features;
[0012] Based on the moment when the voice object changes, segmenting the voice signal into different voice segments;
[0013] Extracting the voiceprints of each voice segment and comparing the voiceprints of different voice segments;
[0014] Determine the voice segments with voiceprint differences less than the preset difference threshold as the same voice in the voice signal;
[0015] Determine the voice segments with voiceprint differences greater than or equal to the preset difference threshold as different voices in the voice signal.
[0016] In some embodiments, each sound-emitting object is correspondingly provided with an azimuth angle, and the azimuth angles of different sound-emitting objects are different; the outputting after sound image localization of the voices of different sound-emitting objects in response to the voice signal including different sound-emitting objects includes:
[0017] In response to the voice signal including different sound-emitting objects, perform sound image localization on the voices of each sound-emitting object based on the head-related transfer function and the azimuth angle corresponding to each sound-emitting object, and then output.
[0018] In some embodiments, the performing sound image localization on the voices of each sound-emitting object based on the head-related transfer function and the azimuth angle corresponding to each sound-emitting object, and then outputting, includes:
[0019] For the voice of each sound-emitting object, convert the voice into a frequency-domain signal;
[0020] Render the frequency-domain signal based on the azimuth angle corresponding to the sound-emitting object and the head-related transfer function to obtain a rendered frequency-domain signal;
[0021] Convert each rendered frequency-domain signal into a time-domain signal and output.
[0022] In some embodiments, the head-related transfer function includes transfer functions corresponding to different sound channels, and the rendering the frequency-domain signal based on the azimuth angle corresponding to the sound-emitting object and the head-related transfer function to obtain a rendered frequency-domain signal includes:
[0023] Render the frequency-domain signal based on the azimuth angle corresponding to the sound-emitting object and the transfer functions corresponding to each sound channel to obtain the frequency-domain signals rendered for each sound channel;
[0024] The converting each rendered frequency-domain signal into a time-domain signal and outputting includes:
[0025] Convert the frequency-domain signal rendered for each sound channel into a time-domain signal and output.
[0026] In some embodiments, the method further includes:
[0027] Receive a setting instruction for the azimuth angle of each sound-emitting object;
[0028] According to the setting instruction, set the azimuth angle corresponding to each sound-emitting object.
[0029] According to a second aspect of the embodiments of the present disclosure, there is provided a voice signal processing device, the device comprising:
[0030] An identification module, configured to, when receiving a voice signal, identify a voice object in the voice signal based on acoustic features;
[0031] An acoustic image localization module, configured to, in response to different voice objects being included in the voice signal, perform acoustic image localization on the voices of the different voice objects and then output; wherein, the acoustic image azimuths of the voices of different voice objects are different.
[0032] In some embodiments, the identification module is further configured to perform voice separation on the voice signal based on acoustic features; and in response to different voices being separated in the voice signal, determine that different voice objects are included in the voice signal.
[0033] In some embodiments, the identification module is further configured to detect a moment when the voice object in the voice signal changes based on acoustic features; segment the voice signal into different voice segments based on the moment when the voice object changes; extract the voiceprints of each voice segment, and compare the voiceprints of different voice segments; determine voice segments with a voiceprint difference less than a preset difference threshold as the same voice in the voice signal; and determine voice segments with a voiceprint difference greater than or equal to the preset difference threshold as different voices in the voice signal.
[0034] In some embodiments, an azimuth angle is correspondingly set for each voice object, and the azimuth angles of different voice objects are different; the acoustic image localization module is further configured to, in response to different voice objects being included in the voice signal, perform acoustic image localization on the voices of each voice object based on the head-related transfer function and the azimuth angle corresponding to each voice object and then output.
[0035] In some embodiments, the acoustic image localization module is further configured to, for the voice of each voice object, convert the voice into a frequency-domain signal; render the frequency-domain signal based on the azimuth angle corresponding to the voice object and the head-related transfer function to obtain a rendered frequency-domain signal; and convert each rendered frequency-domain signal into a time-domain signal and output.
[0036] In some embodiments, the head-related transfer function includes transfer functions corresponding to different sound channels, and the acoustic image localization module is further configured to render the frequency-domain signal based on the azimuth angle corresponding to the voice object and the transfer functions corresponding to each sound channel to obtain frequency-domain signals rendered for each sound channel; and convert the frequency-domain signal rendered for each sound channel into a time-domain signal and output.
[0037] In some embodiments, the device further comprises:
[0038] A receiving module, configured to receive a setting instruction for the azimuth angle of each sound-emitting object;
[0039] A setting module, configured to set the azimuth angle corresponding to each sound-emitting object according to the setting instruction.
[0040] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, and the electronic device includes:
[0041] A processor;
[0042] A memory for storing instructions executable by the processor;
[0043] Wherein, the processor is configured to execute the voice signal processing method as described in the first aspect above.
[0044] According to a fourth aspect of the embodiments of the present disclosure, a storage medium is provided, including:
[0045] When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is enabled to execute the voice signal processing method as described in the first aspect above.
[0046] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:
[0047] In the embodiments of the present disclosure, the electronic device identifies the sound-emitting objects in the voice signal based on acoustic features, and locates the voices of different sound-emitting objects at different sound image azimuths to improve the user's listening experience, without relying on a voice software to distinguish the sound-emitting objects for sound image positioning. It is applicable to scenarios where, for example, the IP of the conference software cannot be obtained to distinguish the voice signals from different sound-emitting objects, or scenarios where voice signals including multiple sound-emitting objects output from the same device are received. Thus, it can be seen that the solutions of the embodiments of the present disclosure have a wider applicability and higher intelligence.
[0048] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.
[0050] Figure 1 Is a flowchart of a voice signal processing method shown in an embodiment of the present disclosure.
[0051] Figure 2 Is an architecture diagram of a voice conference provided by an embodiment of the present disclosure.
[0052] Figure 3It is an example diagram for setting the horizontal azimuth angles of different sound-emitting objects in the embodiments of the present disclosure.
[0053] Figure 4 It is a schematic diagram of a voice signal processing by an audio-video teleconference device in the embodiments of the present disclosure.
[0054] Figure 5 It is an example diagram of a process for an electronic device to process a voice signal in the embodiments of the present disclosure.
[0055] Figure 6 It is Figure 5 an example diagram of a process for identifying a speaker in
[0056] Figure 7 It is Figure 5 an example diagram of a process for voice rendering in
[0057] Figure 8 It is an example diagram of a voice signal processing device in the embodiments of the present disclosure.
[0058] Figure 9 It is a block diagram of a device shown according to an exemplary embodiment. Detailed implementation manners
[0059] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0060] Figure 1 It is a flowchart of a voice signal processing method shown in the embodiments of the present disclosure. As Figure 1 shown, it includes the following steps:
[0061] S11. When a voice signal is received, identify the sound-emitting object in the voice signal based on acoustic features;
[0062] S12. In response to different sound-emitting objects being included in the voice signal, perform sound image localization on the voices of the different sound-emitting objects and then output them; wherein, the sound image azimuths of the voices of different sound-emitting objects are different.
[0063] The voice signal processing method according to an embodiment of the present disclosure can be applied to an electronic device having an audio output function. For example, the electronic device can be an audio and video teleconference device, a smart phone, a smart speaker, in-ear headphones, a head-mounted acoustic device such as a VR glasses or a head-mounted headphone, etc. The embodiments of the present disclosure are not limited thereto. The electronic device can receive a voice signal sent by a terminal device through a wireless or wired connection, and then play it through a speaker for the user to listen.
[0064] In step S11, when the electronic device receives a voice signal, it identifies the voice object in the voice signal based on the acoustic features. Wherein, the voice signal can be a single voice signal or multiple voice signals. Each voice signal can include one or more voice objects. In the case of including multiple voice objects, the multiple voice objects can speak simultaneously or at different times.
[0065] In the embodiments of the present disclosure, the electronic device identifies the voice object in the voice signal based on the acoustic features. Among them, the acoustic features include two types: time-domain features and frequency-domain features. For example, the time-domain features include the intensity of the sound, short-time energy, etc.; the frequency-domain features include the frequency, phase, power spectrum of the sound, etc. Since the acoustic features are audio feature parameters that can express individual information in the sound, the embodiments of the present disclosure can identify the voice object based on the acoustic features in the voice signal.
[0066] In step S12, when the electronic device identifies that the voice signal includes different voice objects, it outputs the voices of different voice objects after sound image localization, so that the sound image azimuths of the voices of different voice objects are different. It should be noted that in the embodiments of the present disclosure, when it is identified that the voice signal includes only one voice object, the voice of the one voice object can also be directly output after sound image localization.
[0067] In the embodiments of the present disclosure, when performing sound image localization, the voices of different voice objects can be output after a preset time interval; and / or the voices of different voice objects can be output at different preset sound intensities, etc. It can also be that for the voice of each voice object, relevant processing functions of sound image localization are adopted, such as Head Related Transfer Function (HRTF), Head Related Impulse Response (HRIR), etc. for processing and then output, so that the listener can feel a sense of hearing with a sense of picture; wherein, the sense of picture means that the listener can distinguish different objects from the output sound or can feel the spatial position of the sound.
[0068] In the related art, the voices of different objects are distinguished based on the Internet Protocol Address (IP) when the voice call software is enabled, and then the sound and image are localized, but it is impossible to process the mixed voices of multiple people in a single channel.
[0069] Figure 2 is an architecture diagram of a voice conference provided in an embodiment of the present disclosure, such as Figure 2 As shown, participant L can receive the voice output of the participant corresponding to each electronic device through the local electronic device. Since the conference software of the electronic devices transmitting different participants (L1, L2, L3 and L4) has different IPs, the local electronic device can distinguish different participants based on the IP and perform sound and image positioning on the voices of different participants and then output them. However, if one electronic device corresponds to multiple participants, the electronic device cannot achieve differential sound and image positioning for different participants after transmitting the voices of multiple participants to the local electronic device.
[0070] In this regard, in the embodiments of the present disclosure, the electronic device identifies the sound-emitting object in the voice signal based on acoustic features, and locates the voices of different sound-emitting objects in different sound and image directions to improve the user's listening experience, without relying on voice software to distinguish the sound-emitting objects for sound and image positioning. This is suitable for scenarios where, for example, it is impossible to obtain the conference software IP to distinguish the voice signals from different sound-emitting objects (i.e., the voice signal includes multiple voices), or scenarios where voice signals including multiple sound-emitting objects output by the same device are received. It can be seen that the solution of the embodiments of the present disclosure has wider applicability and higher intelligence.
[0071] Based on the above Figure 2 For example, in the solution of the embodiment of the present disclosure, the local electronic device processes the voice signals of other participants, automatically identifies the participant who is currently speaking through voice recognition technology, and performs sound and image positioning on the participant, so that the local participant L can perceive that different participants are in different positions when listening to the voices of different participants, and this process can be achieved without relying on the IP information of other participants that comes with the voice conference software.
[0072] In some embodiments, the step of identifying the sound object in the speech signal based on acoustic features includes:
[0073] Performing speech separation on the speech signal based on acoustic features;
[0074] In response to separating different voices from the voice signal, it is determined that the voice signal includes different sound objects.
[0075] In the embodiments of the present disclosure, for voice separation of a voice signal based on acoustic features, traditional methods such as Independent Component Analysis (ICA) and Sparse PCA can be used for voice separation. Deep learning methods can also be used. For example, the TasNet network can be used to perform voice separation on the voice signal in the time domain, or the voice signal can be transformed into the frequency domain or other transform domains, and deep clustering and other networks can be used for voice separation based on the obvious signal features of the voice signal in the transform domain. If different voices are separated from the voice signal, the electronic device determines that the voice signal includes different sound sources.
[0076] It should be noted that in the embodiments of the present disclosure, if the voice signal is a situation where multiple sound sources in a single signal emit sounds simultaneously or non-simultaneously, each sound source's voice can be separated from the multiple mixed sound sources through the aforementioned voice separation method. Of course, if the voice signal is multiple voice signals, the voice of each sound source can also be separated based on the voice separation method of the embodiments of the present disclosure. In addition, if only one voice is separated from the voice signal, it means that the voice signal includes only one sound source.
[0077] It can be understood that in the embodiments of the present disclosure, the method of using voice separation to determine the sound source in the voice signal is applicable to the situation where multiple sound sources in a single voice signal emit sounds simultaneously or non-simultaneously. The solution has a wide application range and good intelligence.
[0078] In some embodiments, the voice separation of the voice signal based on acoustic features includes:
[0079] Detecting the moment when the sound source in the voice signal changes based on acoustic features;
[0080] Based on the moment when the sound source changes, segmenting the voice signal into different voice segments;
[0081] Extracting the voiceprints of each voice segment and comparing the voiceprints of different voice segments;
[0082] Determining the voice segments with voiceprint differences less than a preset difference threshold as the same voice in the voice signal;
[0083] Determining the voice segments with voiceprint differences greater than or equal to the preset difference threshold as different voices in the voice signal.
[0084] In the embodiments of the present disclosure, the electronic device detects the moment when the voice object in the voice signal changes based on acoustic features. For example, the electronic device can input the voice signal into a speaker change detection model (SCD) to determine the moment when the voiceprint changes. In addition, the electronic device can also pre-segment the voice signal into different voice segments with a fixed length or an unfixed length, such as 0.5 seconds to 2 seconds, and then extract the voiceprints in adjacent time windows of adjacent pre-segmented voice segments, and determine the moment when the voice object changes by comparing the difference degree of the voiceprints.
[0085] When the number of voice objects in the voice signal is greater than 2, it is impossible to distinguish how many voice objects there are based on the moment when the voice objects change. Therefore, the electronic device further segments the voice signal into different voice segments based on the moment when the voice objects change.
[0086] In the embodiments of the present disclosure, after the electronic device re-segments to obtain different voice segments, it will extract the voiceprints of each voice segment and then perform voiceprint clustering. Since the voiceprint is an individual unique acoustic feature extracted through regular changes in the frequency domain, time domain, etc. of the sound, the voiceprints of the same voice object have relative stability and do not change with time or environment. Therefore, in the embodiments of the present disclosure, the voiceprints of different voice segments can be compared, so that the voice segments with a voiceprint difference less than a preset difference threshold are determined as the same voice; while the voice segments with a voiceprint difference greater than or equal to the preset difference threshold are determined as different voices in the voice signal. In the embodiments of the present disclosure, each voice after clustering can be identified by different identifiers, and the number of voices after clustering is the number of voice objects.
[0087] In the embodiments of the present disclosure, the electronic device extracts the voiceprint in the voice signal. For example, it can extract the Mel-scale Frequency Cepstral Coefficients (MFCC) feature or the Linear Predictive Coding feature, etc. in the voice signal, and the embodiments of the present disclosure do not make limitations.
[0088] It should be noted that in the embodiments of the present disclosure, when the electronic device extracts the voiceprints of each voice segment, it can extract the voiceprints of each entire voice segment, or extract the voiceprints of local voices in the voice segment to reduce power consumption and improve the speed of voice separation.
[0089] Taking a real-time voice signal as an example, at the beginning of a conversation, the electronic device continuously detects whether the voice source (the speaker) has changed. When it detects a change in the speaker, the electronic device can extract the current speaker's voice of a fixed length (i.e., the local voice in the voice segment) and perform voiceprint clustering on it with the voice of the speaker before the change is detected. If the clustering result identifies the speaker as an existing one, it is directly mapped to the corresponding speaker identifier; if the clustering results in a newly emerged speaker, a new next speaker identifier is assigned to them.
[0090] It can be understood that in the embodiments of the present disclosure, after the electronic device performs voice segmentation on the voice signal based on the voiceprint feature to obtain voice segments, it then extracts the voiceprint features of each voice segment for clustering analysis, thereby achieving voice separation. This is applicable to the situation where a single voice signal includes different voice sources and the voice sources do not speak simultaneously. Based on the voice separation method of the embodiments of the present disclosure, there is no need to rely on voice software to distinguish the voice sources, and the applicable range is wide.
[0091] In some embodiments, each voice source is correspondingly provided with an azimuth angle, and the azimuth angles of different voice sources are different; the step of performing sound image localization on the voices of the different voice sources and then outputting in response to the voice signal including different voice sources includes:
[0092] In response to the voice signal including different voice sources, the voices of each voice source are subjected to sound image localization based on the head-related transfer function and the azimuth angle corresponding to each voice source, and then output.
[0093] In the embodiments of the present disclosure, the voices of each voice source are subjected to sound image localization based on the azimuth angle preset for each voice source by using the head-related transfer function and then output. Since the head-related transfer function describes the transmission process of sound waves from the sound source to the two ears and is a function model established by simulating the human physiological structure (such as the head, auricle, and torso, etc.) and the human perception of sound, performing sound image localization based on the head-related transfer function enables the voices of different voice sources to provide a better listening experience for the listener after being output. Moreover, since the azimuth angles corresponding to each voice source are different, the sound images perceived by the listener after being localized based on the head-related transfer function will also be different.
[0094] The following formula is the expression of the head-related transfer function:
[0095]
[0096]
[0097] where H L is the transfer function corresponding to the left ear, and H Ris the transfer function corresponding to the right ear; the angular frequency f of the sound wave refers to the angular frequency when phonons propagate in a solid, liquid, or gas, and the magnitude of the angular frequency depends on the physical properties of the propagation medium of the phonons; r is the distance from the sound source to the center of the head. When r is greater than 1.2 meters, H L and H R are basically independent of r; θ is the horizontal azimuth angle of the sound source relative to the center of the head; is the elevation angle of the sound source relative to the center of the head; P 0 is the complex sound pressure at the position of the center of the head when the human head is absent; P L and P R are the complex sound pressures generated by the sound source in the left and right ears of the listener, respectively.
[0098] In the embodiments of the present disclosure, the azimuth angles set for each sound-emitting object may include the horizontal azimuth angle and / or the elevation angle of the sound source relative to the center of the head. Since the preset azimuth angles of different sound-emitting objects are different, after sound image localization using the head-related transfer function, the user can perceive and distinguish the voices of different sound-emitting objects. In addition, in the embodiments of the present disclosure, the azimuth angles set for each sound-emitting object may be the azimuth angles preset by the developer and allocated according to the identifiers of the voices of different sound-emitting objects after the electronic device identifies the number of sound-emitting objects, or may be the azimuth angles preset by the user of the electronic device. The embodiments of the present disclosure do not limit this.
[0099] Figure 3 is an example diagram of the setting of the horizontal azimuth angles of different sound-emitting objects in the embodiments of the present disclosure. As Figure 3 shown, the horizontal azimuth angles of different participants L1-L4 are different.
[0100] It should be noted that in the embodiments of the present disclosure, the electronic device may perform localization on a single channel (left ear or right ear) according to the azimuth angles of different sound-emitting objects using the head-related transfer function and then output, or perform cross-channel localization on the voices of different sound-emitting objects and then output. In addition, in the embodiments of the present disclosure, the head-related transfer function may be a function applicable to the listener himself / herself, or may be a general function.
[0101] In some embodiments, outputting the voices of each sound-emitting object after sound image localization based on the head-related transfer function and the azimuth angles corresponding to each sound-emitting object includes:
[0102] For the voice of each sound-emitting object, convert the voice into a frequency-domain signal;
[0103] Render the frequency-domain signal based on the azimuth angle corresponding to the sound-emitting object and the head-related transfer function to obtain a rendered frequency-domain signal;
[0104] Convert each rendered frequency-domain signal into a time-domain signal and output it.
[0105] The head-related transfer function processes speech in the frequency domain. In the embodiments of the present disclosure, for the speech of each sound-emitting object, the electronic device converts the time-domain speech into a frequency-domain signal. For example, the speech of each sound-emitting object can be converted into the frequency domain through Fourier transform. Subsequently, after substituting the azimuth angle corresponding to the sound-emitting object into the aforementioned formula (1) and / or formula (2), the electronic device can obtain the rendered frequency-domain signal, and then, for example, convert the frequency-domain signal into a time-domain signal through inverse Fourier transform and output it. Since the azimuth angles corresponding to different sound-emitting objects are different, the time-domain signals converted from the rendered frequency-domain signals will also be different, so that the listener can perceive different sound images of different sound-emitting objects.
[0106] In some embodiments, the head-related transfer function includes transfer functions corresponding to different channels. The rendering of the frequency-domain signal based on the azimuth angle corresponding to the sound-emitting object and the head-related transfer function to obtain the rendered frequency-domain signal includes:
[0107] Render the frequency-domain signal based on the azimuth angle corresponding to the sound-emitting object and the transfer functions corresponding to each channel to obtain the rendered frequency-domain signals of each channel;
[0108] The conversion of each rendered frequency-domain signal into a time-domain signal and output includes:
[0109] Convert the rendered frequency-domain signal of each channel into a time-domain signal and output it.
[0110] As described above, the electronic device can perform dual-channel localization and output for the speech of each sound-emitting object. In the embodiments of the present disclosure, the electronic device uses the above formulas (1) and (2) to perform multi-channel rendering on the speech of each sound-emitting object and then convert it into a time-domain signal and output it through the left and right channels of the electronic device, so that each sound-emitting object's speech can have a dual-channel stereo surround feeling after output. It can be understood that through multi-channel speech processing and output, the listener can enjoy the virtual surround sound perception on the basis of distinguishing the speech of different sound-emitting objects, so the quality of the speech output is relatively high.
[0111] In some embodiments, the method further includes:
[0112] Receive a setting instruction for the azimuth angle of each sound-emitting object;
[0113] According to the setting instruction, set the azimuth angle corresponding to each sound-emitting object.
[0114] As described above, the azimuth angles set for different sound-emitting objects can be preset by the user of the electronic device. In the embodiments of the present disclosure, the listener (such as the local speaker) of the electronic device can select an azimuth angle for each sound-emitting object in the software interface of the local electronic device, and can also set the azimuth angle of the sound-emitting object through a voice command. After the electronic device receives the setting commands for the azimuth angles of the respective sound-emitting objects, it can set the azimuth angles corresponding to the respective sound-emitting objects according to the setting commands.
[0115] It can be understood that in the embodiments of the present disclosure, the electronic device supports the custom setting of the azimuth angles of the respective sound-emitting objects, and has a relatively high intelligence.
[0116] Taking the electronic device in the embodiments of the present disclosure as an audio-video teleconference device as an example, Figure 4 is a schematic diagram of the processing of voice signals by an audio-video teleconference device in the embodiments of the present disclosure, including the sender of the voice signal (such as audio-video teleconference device A) and the receiver of the voice signal (such as audio-video teleconference device B). As Figure 4 shown, after the voice of a certain participant L1 is collected by the microphone A1 of the audio-video teleconference device A, the audio-video teleconference device A converts the collected analog voice signal into a digital signal through analog-to-digital conversion A2 and transmits it to the local audio-video teleconference device B through network transmission of A3. The local audio-video teleconference device B performs speaker recognition of B1 on the received voice signal, that is, recognizes the sound-emitting objects included in the voice signal based on acoustic features. After the local audio-video teleconference device B recognizes each speaker, it uses the HRTF function of B2 to render the voices of each speaker, and then converts the rendered digital signal into an analog signal through digital-to-analog conversion of B3 and outputs it through the speaker B4, so that the local participant L can perceive the voices of different sound-emitting objects.
[0117] Based on Figure 4 the principle, the embodiments of the present disclosure realize the auditory spatial sense of the local speaker during a multi-person call, increase the immersion and authenticity of the call. Since it does not need to rely on the system hardware to distinguish speakers based on IP, it has a low dependence on the hardware system. Only a single microphone is required at the upstream end, and only a common stereo headset or two speakers are required at the downstream end to achieve this. The system is simple and the cost is low.
[0118] Figure 5 is an example flowchart of the processing of voice signals by the electronic device in the embodiments of the present disclosure. As Figure 5 shown, it includes the following steps:
[0119] S21. Set the position of each participant relative to the local participant.
[0120] In an embodiment of the present disclosure, an electronic device may set the positions of each participant relative to the local participant based on a user's setting instruction for the azimuth angle, where the electronic device may be the above-mentioned audio and video teleconference device B (local electronic device).
[0121] S22. The local electronic device receives voice information.
[0122] In an embodiment of the present disclosure, the voice information received by the local electronic device is the voice signal described above in the embodiment of the present disclosure.
[0123] S23. Determine the current speaker by means of speech recognition.
[0124] In an embodiment of the present disclosure, the local electronic device determines the current speaker by means of speech recognition, that is, the local electronic device identifies the voice object in the voice signal based on acoustic features.
[0125] S24. Render the participant's voice with a specific HRTF according to the azimuth specified in advance for the current speaker.
[0126] In an embodiment of the present disclosure, the local electronic device renders the participant's voice with a specific HRTF according to the azimuth specified in advance for the current speaker, that is, for the voice of each voice object, the electronic device converts the voice into a frequency-domain signal, and then renders the frequency-domain signal based on the azimuth angle corresponding to the voice object and the head-related transfer function.
[0127] S25. Play the rendered audio to the local participant.
[0128] In an embodiment of the present disclosure, the electronic device plays the rendered audio to the local participant, that is, the electronic device converts each rendered frequency-domain signal into a time-domain signal and outputs it.
[0129] Figure 6 Yes Figure 5 is a flow example diagram for identifying a speaker in, as Figure 6 shown, including the following steps:
[0130] S32. It is detected that the speaker has changed.
[0131] In an embodiment of the present disclosure, an electronic device may detect a change in the speaker, for example, based on the aforementioned speaker conversion detection model, that is, detect the moment when the voice object in the voice signal changes its voice based on the acoustic features of the voice signal.
[0132] S33. Extract a fixed-length voice after the speaker change for voiceprint clustering.
[0133] S34. Obtain the current speaker identifier after the change.
[0134] In an embodiment of the present disclosure, the electronic device extracts fixed-length speech after a speaker change for voiceprint clustering, and obtains the current speaker identifier after the change, that is, the electronic device compares the voiceprint of the speech segment with the voiceprints of different speech segments, and determines the identifier of the currently compared speech segment according to the comparison result.
[0135] Figure 7 Yes Figure 5 is a flowchart example of voice rendering in, as Figure 7 shown, including the following steps:
[0136] S41. Single-channel speech.
[0137] S42. Spectrum of single-channel speech.
[0138] In an embodiment of the present disclosure, the electronic device can obtain single-channel speech based on the Figure 6 steps in, and the single-channel speech is the speech of one voice source. After the electronic device performs Fourier transform on the single-channel speech, the spectrum of the single-channel speech can be obtained.
[0139] S43. Spectrum of the rendered two-channel speech.
[0140] In an embodiment of the present disclosure, after the electronic device renders the spectrum of the single-channel speech using the aforementioned head-related transfer function, the frequency-domain signals after rendering for each channel can be obtained.
[0141] S44. Rendered two-channel speech.
[0142] In an embodiment of the present disclosure, the electronic device converts the frequency-domain signal after rendering for each channel into a time-domain signal using inverse Fourier transform to obtain the rendered two-channel speech.
[0143] It can be understood that in an embodiment of the present disclosure, the electronic device identifies the participants in the voice information, and locates the voices of different participants at different sound image positions to improve the user's listening experience, without relying on voice software to distinguish the participants for sound image positioning. The solution has a wider applicability and higher intelligence.
[0144] It should be noted that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order that constitutes any limitation on the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.
[0145] Figure 8 is an example diagram of a voice signal processing device in an embodiment of the present disclosure. Referring to Figure 8 , the device includes:
[0146] The recognition module 101 is configured to, when receiving a voice signal, recognize the voice object in the voice signal based on acoustic features;
[0147] The sound image localization module 102 is configured to, in response to different voice objects being included in the voice signal, perform sound image localization on the voices of the different voice objects and then output them; wherein, the sound image azimuths of the voices of different voice objects are different.
[0148] In some embodiments, the recognition module 101 is further configured to perform voice separation on the voice signal based on acoustic features; and in response to different voices being separated in the voice signal, determine that different voice objects are included in the voice signal.
[0149] In some embodiments, the recognition module 101 is further configured to detect the moment when the voice object in the voice signal changes based on acoustic features; segment the voice signal into different voice segments based on the moment when the voice object changes; extract the voiceprints of each voice segment and compare the voiceprints of different voice segments; determine the voice segments with voiceprint differences less than a preset difference threshold as the same voice in the voice signal; and determine the voice segments with voiceprint differences greater than or equal to the preset difference threshold as different voices in the voice signal.
[0150] In some embodiments, an azimuth angle is correspondingly set for each voice object, and the azimuth angles of different voice objects are different; the sound image localization module 102 is further configured to, in response to different voice objects being included in the voice signal, perform sound image localization on the voices of each voice object based on the head-related transfer function and the azimuth angle corresponding to each voice object and then output them.
[0151] In some embodiments, the sound image localization module 102 is further configured to, for the voice of each voice object, convert the voice into a frequency-domain signal; render the frequency-domain signal based on the azimuth angle corresponding to the voice object and the head-related transfer function to obtain a rendered frequency-domain signal; and convert each rendered frequency-domain signal into a time-domain signal and output it.
[0152] In some embodiments, the head-related transfer function includes transfer functions corresponding to different sound channels, and the sound image localization module 102 is further configured to render the frequency-domain signal based on the azimuth angle corresponding to the voice object and the transfer functions corresponding to each sound channel to obtain frequency-domain signals rendered for each sound channel; and convert the frequency-domain signal rendered for each sound channel into a time-domain signal and output it.
[0153] In some embodiments, the device further includes:
[0154] The receiving module 103 is configured to receive a setting instruction for the azimuth angle of each voice object;
[0155] A setting module 104, configured to set the azimuth angle corresponding to each sound-emitting object according to the setting instruction.
[0156] Regarding Figure 8 For the device in the illustrated embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment related to the method, and will not be elaborated herein.
[0157] Figure 9 It is a block diagram of a device shown according to an exemplary embodiment. For example, the device 800 (equipment) may be an audio-visual teleconference device, a smart phone, etc.
[0158] Referring to Figure 9 , the device 800 may include one or more of the following components: a processing component 802, a memory 804, a power supply component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0159] The processing component 802 generally controls the overall operation of the device 800, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0160] The memory 804 is configured to store various types of data to support the operation of the device 800. Examples of such data include instructions for any application or method operating on the device 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0161] The power supply component 806 provides power to various components of the device 800. The power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 800.
[0162] The multimedia component 808 includes a screen that provides an output interface between the device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operation mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0163] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive external audio signals when the device 800 is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 804 or sent via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.
[0164] The I / O interface 812 provides an interface between the processing component 802 and a peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power button, and a lock button.
[0165] The sensor component 814 includes one or more sensors for providing a status assessment of various aspects of the device 800. For example, the sensor component 814 can detect the on / off state of the device 800, the relative positioning of components, such as the display and the keypad of the device 800. The sensor component 814 can also detect a change in the position of the device 800 or a component of the device 800, the presence or absence of user contact with the device 800, the orientation or acceleration / deceleration of the device 800, and the temperature change of the device 800. The sensor component 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 814 can also include a light sensor, such as a CMOS or a CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 814 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0166] The communication component 816 is configured to facilitate communication between the device 800 and other devices in a wired or wireless manner. The device 800 may access a communication standard-based wireless network, such as Wi-Fi, 4G, or 5G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra-Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0167] In an exemplary embodiment, the device 800 may be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.
[0168] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as the memory 804 including instructions, and the above instructions can be executed by the processor 820 of the device 800 to complete the above method. For example, the non-transitory computer-readable storage medium may be a ROM, Random Access Memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0169] A non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to execute the foregoing voice signal processing method.
[0170] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and embodiments are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0171] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A speech signal processing method, characterized in that: The method comprises: When a speech signal is received, identifying a sounding object in the speech signal based on acoustic features; In response to the voice signal including different sound-generating objects, the voices of the different sound-generating objects are localized and then output; wherein the sound image orientations of the voices of the different sound-generating objects are different.
2. The method according to claim 1, characterized in that The identifying the sound object in the speech signal based on acoustic features includes: Performing speech separation on the speech signal based on acoustic features; In response to different voices being separated from the voice signal, it is determined that the voice signal includes different sound objects.
3. The method according to claim 2, characterized in that The performing speech separation on the speech signal based on acoustic features comprises: Detecting a moment when a sound object in the speech signal changes based on acoustic features; Segmenting the speech signal into different speech segments based on the moment when the sound object changes; Extract the voiceprint of each speech segment and compare the voiceprints of different speech segments; Determine the voice segments whose voiceprint differences are less than a preset difference threshold as the same voice in the voice signal; The speech segments whose voiceprint differences are greater than or equal to the preset difference threshold are determined as different speech in the speech signal.
4. The method according to claim 1, characterized in that: Each sound object is correspondingly provided with an azimuth angle, and different sound objects have different azimuth angles; in response to the voice signal including different sound objects, the voices of the different sound objects are localized and then output, including: In response to the speech signal including different sound objects, the speech of each sound object is localized based on the head-related transfer function and the azimuth angle corresponding to each sound object and then output.
5. The method according to claim 4, characterized in that The method of outputting the sound image of each sound object after performing sound image localization based on the head-related transfer function and the azimuth angle corresponding to each sound object comprises: For each sound-generating object's speech, convert the speech into a frequency domain signal; Rendering the frequency domain signal based on the azimuth angle corresponding to the sound-emitting object and the head-related transfer function to obtain a rendered frequency domain signal; Convert each rendered frequency domain signal into a time domain signal and output it.
6. The method according to claim 5, characterized in that The head-related transfer function includes transfer functions corresponding to different channels, and the frequency domain signal is rendered based on the azimuth angle corresponding to the sound-emitting object and the head-related transfer function to obtain the rendered frequency domain signal, including: Rendering the frequency domain signal based on the azimuth angle corresponding to the sound object and the transfer function corresponding to each channel to obtain the rendered frequency domain signal of each channel; The step of converting each rendered frequency domain signal into a time domain signal and outputting the signal comprises: The frequency domain signal rendered for each channel is converted into a time domain signal and output.
7. The method according to claim 4, characterized in that The method further comprises: Receiving instructions for setting the azimuth angle of each sound-emitting object; According to the setting instruction, the azimuth angle corresponding to each sound-emitting object is set.
8. A speech signal processing device, characterized in that: The device comprises: A recognition module, configured to recognize a sound object in the speech signal based on acoustic features when receiving the speech signal; The sound image localization module is configured to perform sound image localization on the voices of different sound objects in response to the voice signal including different sound objects, and then output the localized sound images; wherein the sound image orientations of the voices of different sound objects are different.
9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to execute the speech signal processing method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium, characterized in that: When the instructions in the storage medium are executed by a processor in an electronic device, the electronic device is enabled to execute the speech signal processing method as claimed in any one of claims 1 to 7.