Method, apparatus, device, medium and program product for processing call audio
By setting up multiple microphones and speakers on the terminal device, audio data containing spatial location information is collected and generated, solving the problem of lack of spatial location awareness in call audio in the prior art, and realizing the user's spatial location awareness experience in real-time calls.
Patent Information
- Application Number
- CN202210983451.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-16
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2042-08-16
AI Technical Summary
In existing real-time video call technology, the call audio only conveys information from the user's verbal expression and lacks spatial location awareness, causing the user to be unable to perceive changes in the other party's spatial location.
By setting up at least two microphones and speakers on the terminal device, spatial audio data containing the user's spatial location information is collected and generated in real time, and audio rendering is performed on the other terminal to restore the spatial audio scene.
It enables users to perceive changes in the other party's spatial location during real-time calls, enhancing their immersive experience.
Smart Images

Figure CN115550831B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of spatial audio technology, and in particular to a method, apparatus, device, medium, and program product for processing call audio. Background Technology
[0002] Today, real-time video calling technology is widely used in people's lives and work, such as in online healthcare, video conferencing, social entertainment, online education, and online finance.
[0003] In real-time video call scenarios, the processing of call audio involves one party's terminal acquiring the user's call audio through a microphone, encoding the sound signal, and then sending the encoded audio data to the other party's terminal. Upon receiving the audio data, the other party's terminal decodes the audio data to obtain the user's call audio and plays it through a speaker.
[0004] Call audio generally only conveys information about the user's verbal expression during the call. Summary of the Invention
[0005] This application provides a method, apparatus, device, medium, and program product for processing call audio. The technical solution is as follows:
[0006] According to one aspect of this application, a method for processing call audio is provided, the method being executed by a first terminal, wherein at least two microphones of the first terminal are disposed in different locations, the method comprising:
[0007] The user's voice is captured in real time through the at least two microphones to obtain at least two audio streams of the call.
[0008] Spatial audio data is generated based on the at least two audio channels of the call, and the spatial audio data refers to audio data containing the user's real-time location information in space;
[0009] The spatial audio data is sent to a second terminal, which is a user device in the same real-time call as the first terminal.
[0010] According to another aspect of this application, a method for processing call audio is provided, the method being executed by a second terminal, wherein at least two speakers of the second terminal are disposed in different positions, the method comprising:
[0011] The system receives spatial audio data sent by a first terminal. The spatial audio data includes the user's real-time location information in space and the user's call audio. The call audio is obtained by collecting the user's voice during a call. The first terminal and the second terminal are user devices in the same real-time call.
[0012] Based on the spatial audio data, at least two channels of spatial audio are generated corresponding to the at least two speakers;
[0013] The at least two channels of spatial audio are played through the at least two speakers.
[0014] According to another aspect of this application, a call audio processing apparatus is provided, the apparatus being disposed in a first terminal, wherein at least two microphones of the first terminal are disposed in different positions, the apparatus comprising:
[0015] The acquisition module is used to acquire the user's voice during a call in real time through the at least two microphones to obtain at least two audio streams of the call;
[0016] The generation module is used to generate spatial audio data based on the at least two call audios, wherein the spatial audio data refers to audio data containing the user's real-time location information in space;
[0017] The sending module is used to send the spatial audio data to a second terminal, which is a user device in the same real-time call as the first terminal.
[0018] According to another aspect of this application, a call audio processing apparatus is provided, the apparatus being disposed in a second terminal, wherein at least two speakers of the second terminal are disposed in different positions, the apparatus comprising:
[0019] The receiving module is used to receive spatial audio data sent by the first terminal. The spatial audio data includes the user's real-time location information in space and the user's call audio. The call audio is obtained by collecting the user's voice during the call. The first terminal and the second terminal are user devices in the same real-time call.
[0020] A generation module is used to generate at least two channels of spatial audio corresponding to the at least two speakers based on the spatial audio data;
[0021] A playback module for playing the at least two channels of spatial audio through the at least two speakers.
[0022] According to another aspect of this application, a terminal is provided, the terminal including a processor and a memory connected to the processor, the memory storing program instructions, and the processor executing the program instructions to implement the call audio processing method as provided in various aspects of this application.
[0023] According to another aspect of this application, a computer-readable storage medium is provided, wherein program instructions are stored therein, which, when executed by a processor, implement the call audio processing method as provided in various aspects of this application.
[0024] According to another aspect of this application, a computer program product (or computer program) is provided, the computer program product (or computer program) including computer instructions stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to perform the method provided in various alternative implementations of the above-described method for processing call audio.
[0025] According to another aspect of this application, a chip is provided, the chip including programmable logic circuitry and / or program instructions, which, when the chip is running, are used to implement the call audio processing methods provided in various aspects of this application.
[0026] The beneficial effects of the technical solutions provided in this application embodiment may include:
[0027] In the above-mentioned method for processing call audio, the first terminal in the real-time call collects the user's voice during the call through at least two microphones to obtain at least two audio streams. The at least two microphones are set in different known locations. Therefore, the first terminal can determine the user's real-time location information in space based on the at least two audio streams, thereby generating spatial audio data containing the user's real-time location information. This spatial audio data is then sent to the second terminal in the same real-time call. The second terminal reproduces the spatial audio field based on the spatial audio data, so that the user of the second terminal can perceive the user's spatial position relative to the first terminal. Attached Figure Description
[0028] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0029] Figure 1A schematic diagram of a communication system provided in an exemplary embodiment of this application is shown;
[0030] Figure 2 A schematic diagram of an audio call scenario provided by an exemplary embodiment of this application is shown;
[0031] Figure 3 A flowchart illustrating a method for processing call audio provided in an exemplary embodiment of this application is shown;
[0032] Figure 4 A flowchart illustrating a method for processing call audio provided in another exemplary embodiment of this application is shown;
[0033] Figure 5 A flowchart illustrating a method for processing call audio provided in another exemplary embodiment of this application is shown;
[0034] Figure 6 A flowchart illustrating a method for processing call audio provided in another exemplary embodiment of this application is shown;
[0035] Figure 7 A flowchart illustrating a method for processing call audio provided in another exemplary embodiment of this application is shown;
[0036] Figure 8 A flowchart illustrating a method for processing call audio provided in another exemplary embodiment of this application is shown;
[0037] Figure 9 This illustration shows a block diagram of a call audio processing apparatus provided in an exemplary embodiment of this application;
[0038] Figure 10 A block diagram of a call audio processing apparatus provided in another exemplary embodiment of this application is shown;
[0039] Figure 11 A schematic diagram of the structure of a computer device provided in an exemplary embodiment of this application is shown. Detailed Implementation
[0040] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0041] In the following description, when referring to the accompanying drawings, the same numbers in different drawings denote the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0042] In the description of this application, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. In the description of this application, it should be noted that, unless otherwise explicitly specified and limited, the terms "connected" and "linked" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances. Furthermore, in the description of this application, unless otherwise stated, "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.
[0043] Figure 1 A schematic diagram of a communication system provided in an exemplary embodiment of this application is shown. The communication system includes a first terminal 110 and a second terminal 120.
[0044] Both the first terminal 110 and the second terminal 120 described above have wired and / or wireless communication capabilities. Optionally, the first terminal 110 and the second terminal 120 can communicate using either a wireless network or a wired network. For example, the wired network can be a metropolitan area network, a local area network, a fiber optic network, etc.; the wireless network can be a mobile communication network or a Wireless Fidelity (WiFi) network.
[0045] Both the first terminal 110 and the second terminal 120 have real-time calling functionality. For example, this real-time calling functionality includes at least one of voice calling and audio / video calling. Exemplarily, both the first terminal 110 and the second terminal 120 have an operating system installed. The operating systems in the first terminal 110 and the second terminal 120 may be the same or different. The operating system may be Android OS, HarmonyOS, or iOS.
[0046] Both the first terminal 110 and the second terminal 120 have applications installed that support real-time calling. And / or, both the first terminal 110 and the second terminal 120 have applications installed that support the operation of mini-programs, which have real-time calling functionality. And / or, the operating systems of both the first terminal 110 and the second terminal 120 support the operation of quick apps, which have real-time calling functionality.
[0047] The applications in the first terminal 110 and the second terminal 120 can be the same version of the application or different versions of the application. The same version of the application refers to an application of the same release running on the same operating system; different versions of the application include applications of different releases running on different operating systems, and applications of different releases running on the same operating system.
[0048] For example, the first terminal 110 includes a microphone 111, a processor 112, a memory 113, a communication component 114, and a speaker 115; the memory 113 stores at least one instruction, which is executed by the processor 112 to control the microphone 111 to collect sound signals, control the communication component 114 to conduct wired or wireless communication with the communication component 124 in the second terminal 120, and control the speaker 115 to play sound signals.
[0049] The second terminal 120 includes a speaker 121, a processor 122, a memory 123, a communication component 124, and a microphone 125. The memory 123 stores at least one instruction, which is loaded and executed by the processor 122 to control wired or wireless communication between the communication component 124 and the communication component 114 in the first terminal 110, control the speaker 121 to play sound signals, and control the microphone 125 to collect sound signals.
[0050] A processor may include one or more processing cores. The processor uses various interfaces and lines to connect to various parts within the terminal, and performs various functions and processes data by running or executing instructions, programs, code sets or instruction sets stored in memory, and by calling data stored in memory.
[0051] Optionally, the processor can be implemented using at least one of the following hardware forms: Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), and Programmable Logic Array (PLA). The processor can integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and Modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required for display; and the modem handles wireless communication. It is understood that the modem can also be implemented as a separate chip without being integrated into the processor.
[0052] The memory may include random access memory (RAM) or read-only memory (ROM). Optionally, the memory may include non-transitory computer-readable storage medium. The memory can be used to store instructions, programs, code, code sets, or instruction sets. The memory may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound acquisition function, sound playback function, image processing function, information interaction function, etc.), instructions for implementing the various method embodiments described below, etc.; the data storage area may store data involved in the various method embodiments described below, etc.
[0053] For example, the communication component may include at least one of the following: a WiFi communication component, a mobile (including 3G / 4G / 5G) communication component. The aforementioned communication component is used to establish a wireless communication connection between the first terminal 110 and the second terminal 120, and to transmit data through the wireless communication connection.
[0054] For example, the above communication system also includes a server ( Figure 1(Not shown in the image), the first terminal 110 is connected to the server via a communication network, and the second terminal 120 is also connected to the server via a communication network. This communication network can be a wired network or a wireless network. For example, the first terminal 110 and the second terminal 120 can forward data through the server; for instance, the first terminal 110 sends spatial audio data to the server, and the server sends the spatial audio data to the second terminal 120. For example, the wired network can be a metropolitan area network (MAN), a local area network (LAN), a fiber optic network, etc.; the wireless network can be a mobile communication network or WiFi.
[0055] For example, at least two microphones are provided at different positions on the first terminal 110, and at least two speakers are provided at different positions on the second terminal 120. When a real-time call is established between the first terminal 110 and the second terminal 120, the first terminal 110 collects the user's voice during the call in real time through the at least two microphones to obtain at least two channels of call audio. The positions of the at least two microphones of the first terminal 110 are known. Thus, the first terminal can determine the user's real-time position in space relative to the first terminal based on the at least two channels of call audio. Then, it generates spatial audio data based on the real-time position and the at least two channels of call audio, and sends the spatial audio data to the second terminal 120. After receiving the spatial audio data, the second terminal 120 processes the spatial audio data, renders the at least two channels of call audio based on the real-time position, generates at least two channels of spatial audio corresponding to the at least two speakers, and then plays the at least two channels of spatial audio through the at least two speakers.
[0056] For example, terminals include, but are not limited to, at least one of the following: smart terminals, tablet computers, laptop computers, smartwatches, e-readers, smart robots, and in-vehicle devices. The aforementioned servers may include at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. Those skilled in the art will understand that the number of terminals in the aforementioned communication system can be more or less. For example, the aforementioned communication system may have only one terminal, or dozens or hundreds, or more. This application does not limit the number or type of terminals in the communication system.
[0057] The call audio processing method provided in this application uses spatial audio for real-time calls. For example, its call principle is as follows: Figure 2As shown, during a real-time call between the first terminal and the second terminal, with the first terminal 110 as the base point, the first user moves around the first terminal 110, and the position of the first user relative to the first terminal 110 changes. As the first user moves, the first user also conducts a real-time call with the second user using the second terminal 120 through the first terminal 110. The first terminal 110 collects the call audio of the first user through at least two microphones, obtaining at least two channels of call audio. Then, a dual-microphone sound source localization method is used to perform sound source localization based on at least two channels of call audio, which is actually the location of the first user's position (including the position of the first user's mouth), determining the real-time position of the first user relative to the first terminal 110, and simultaneously processing at least two channels of call audio, converting at least two channels of call audio into at least two-channel audio signals; generating spatial audio data containing the real-time position information of the first user and at least two channels of audio signals, and sending the spatial audio data to the second terminal 120.
[0058] If the second terminal 120 is in speaker mode, it processes at least two channels of audio signals based on the aforementioned real-time location information to generate at least two channels of spatial audio corresponding to at least two speakers of the second terminal 120, and then plays the at least two channels of spatial audio through the at least two speakers. If the second terminal 120 is in headphone playback mode, it processes at least two channels of audio signals based on the aforementioned real-time location information to generate at least two channels of spatial audio matching the headphone playback mode, and then plays the at least two channels of spatial audio through headphones. The aforementioned speaker mode includes a mode that uses the terminal's built-in speakers for audio playback.
[0059] The above-mentioned call audio processing method can be divided into three steps, such as... Figure 3 As shown:
[0060] Step 220: Capture call audio.
[0061] For example, the first terminal 110 is provided with at least two built-in microphones, and both built-in microphones are located on the long axis of the first terminal 110, for example, on... Figure 1 On the long axis of the first terminal 110, there are three built-in microphones arranged in a vertical order.
[0062] The first terminal 110 captures the user's voice during a call in real time using each of at least two built-in microphones, obtaining at least two audio streams. For example, the first terminal 110 captures the first user's voice during a real-time call using each built-in microphone, obtaining at least two audio streams.
[0063] Step 240: Audio spatial processing.
[0064] After acquiring at least two audio feeds, the first terminal 110 calculates the real-time location information of the first user in space based on these two audio feeds, such as the distance and direction information of the first user. The distance information can refer to a relative distance, such as the distance between the first user and the first terminal. The direction information can refer to a relative direction; for example, the relative direction is the direction the first terminal is in when observing the first user from its current location, such as southwest of the first terminal. Alternatively, the direction information can be a relative azimuth angle; for example, the relative azimuth angle is the azimuth angle of the first user when observing the first user from its current location, such as 45 degrees east of south of the first terminal. For example, the real-time location information can include at least one of the distance between the first user and the microphone, and the distance between the first user and the first terminal; the first user's direction information can include at least one of the following: the first user's direction relative to the microphone, the first user's direction relative to the first terminal, the first user's azimuth angle relative to the microphone, and the first user's azimuth angle relative to the first terminal.
[0065] For example, since the microphone is likely to pick up external environmental noise when collecting the user's voice, after obtaining at least two call audios, the first terminal first removes environmental noise other than human voice from the at least two call audios to obtain at least two call audios containing pure human voice; then, based on the above at least two call audios containing pure human voice, spatial audio is generated.
[0066] For example, when at least two audio recordings are collected, the first terminal 110 calculates the real-time location information of the first user relative to the first terminal 110 based on the at least two audio recordings. For example, it calculates the distance between the first user and the first terminal 110, and the direction of the first user relative to the first terminal 110, such as the first user being 3 meters to the left and in front of the first terminal 110.
[0067] For example, when at least three audio streams of a call are collected, the first terminal 110 determines a set of spatial location information based on every two audio streams from the at least three audio streams, thus obtaining at least two sets of spatial location information; and generates the real-time location information of the first user based on the at least two sets of spatial location information. For example, the spatial location information includes the position and direction information of the first user relative to the first terminal 110. The average of the position and direction information in the at least two sets of spatial location information is calculated to obtain a set of spatial location information as the real-time location information of the first user.
[0068] The first terminal 110 will also convert at least two call audio signals into at least two-channel audio signals. For example, the first terminal 110 can process N call audio signals to generate N-channel audio signals, where N is a positive integer greater than 1; for example, after noise reduction processing of two call audio signals, a stereo audio signal is generated; or, for example, after combining two call audio signals into a mono audio signal, noise reduction processing is performed on the mono audio signal, and then the stereo audio signal is extracted from the noise-reduced mono audio signal.
[0069] Alternatively, the first terminal 110 can process N channels of call audio to generate an M-channel audio signal, where M is a positive integer greater than N; for example, after noise reduction processing of two channels of call audio, a stereo audio signal is generated; the stereo audio signal is combined into a mono audio signal; the stereo audio signal and the mono audio signal are merged to obtain a three-channel audio signal; or, for example, the mono audio signal is divided according to the frequency domain and then a three-channel audio signal is extracted.
[0070] Alternatively, the first terminal 110 can process N channels of call audio to generate a P-channel audio signal, where P is a positive integer less than N; for example, it can process three channels of call audio with noise reduction and then synthesize them into a mono audio signal; or, for example, it can synthesize three channels of call audio into a mono audio signal and then copy the mono audio signal to generate a stereo audio signal.
[0071] The first terminal 110 can use at least one of the following methods—synthesis, copying, and cutting—to convert N channels of call audio into an M-channel audio signal, or an N-channel audio signal, or a P-channel audio signal.
[0072] After obtaining the real-time location information of the first user and at least two channels of audio signals, the first terminal 110 combines them to generate spatial audio data. For example, the audio signal at each moment corresponds to a set of real-time location information, and the real-time location information of the first user and the audio signals at least two channels are aligned according to the time correspondence to generate spatial audio data.
[0073] Subsequently, the first terminal 110 sends spatial audio data to the second terminal 120. The second terminal and the first terminal are user devices in the same real-time communication. For example, the first terminal 110 encodes the spatial audio data and sends the encoded spatial audio data to the second terminal 120.
[0074] It should be noted that the steps of generating the real-time location information of the first user and generating at least two channels of audio signals can be performed simultaneously or sequentially. In the embodiments of this application, the execution order of the two is not limited.
[0075] Step 3, 260: Audio space restoration.
[0076] After receiving the spatial audio data, the second terminal 120 renders at least two channels of audio signals based on the real-time location information of the first user to generate at least two channels of spatial audio, and then plays the at least two channels of spatial audio.
[0077] For example, the spatial audio data includes at least two channels of audio signal. If the current playback mode of the second terminal 120 supports the playback of at least two channels of spatial audio, the second terminal 120 directly renders the at least two channels of audio signal based on the real-time location information of the first user to generate at least two channels of spatial audio. If the current playback mode of the second terminal 120 does not support the playback of at least two channels of spatial audio, the at least two channels of audio signal are converted into K-channel audio signal, and the K-channel audio signal is rendered based on the real-time location information of the first user to generate K-channel spatial audio. K channels represent the number of channels supported in the current playback mode, and K is a positive integer. The terminal's playback mode includes at least one of an external speaker mode and an external device playback mode. When playing audio, the second terminal 120 can use either an external speaker mode or an external device playback mode. For example, the second terminal 120 can use at least two built-in speakers to play at least two channels of spatial audio; or, for example, the second terminal 120 can use two headphone speakers to play at least two channels of spatial audio. The external device may include at least one of wired headphones, wired speakers, Bluetooth headphones, and Bluetooth speakers.
[0078] During the generation of spatial audio, the real-time location information of the first user is determined using the first terminal 110 as a reference. During the reconstruction of spatial audio, for example, the second terminal 120, using its own location and direction as references, renders at least two channels of audio signals based on the first user's real-time location information to generate at least two channels of spatial audio. The real-time location information of the first user is then reconstructed from the at least two channels of spatial audio. For example, if the first user speaks towards the screen of the first terminal 110, and the first user is in front of the screen of the first terminal 110, the reconstructed at least two channels of spatial audio will show the first user in front of the screen of the second terminal 120, meaning the second terminal 120 is considered as the first terminal 110, thus reconstructing the audio scene of the first user speaking in front of the screen of the first terminal 110.
[0079] In summary, the call audio processing method provided in this embodiment involves a first terminal in a real-time call acquiring the user's voice through at least two microphones to obtain at least two audio streams. Since the at least two microphones are positioned at different known locations, the first terminal can determine the user's real-time location information based on these at least two audio streams. This generates spatial audio data containing the user's real-time location information, which is then sent to a second terminal in the same real-time call. The second terminal then recreates the spatial audio field based on the spatial audio data, allowing the user of the second terminal to perceive the user's spatial position relative to the first terminal.
[0080] Figure 4 A flowchart illustrating a method for processing call audio according to an exemplary embodiment of this application is shown. This method can be applied to... Figure 1 In the first terminal shown, the method includes:
[0081] Step 320: Collect the user's voice during the call in real time using at least two microphones to obtain at least two audio streams of the call.
[0082] For example, the above-mentioned at least two microphones are at least two built-in microphones, and the at least two built-in microphones are set in different positions of the first terminal; for example, the built-in microphone 1 on the mobile terminal is set above the screen, and the built-in microphone 2 is set below the screen.
[0083] With the recording function of at least two built-in microphones enabled, the first terminal collects the call audio from the sound source in real time through each built-in microphone, and marks the collected call audio with a timestamp to record the time of collection of the call audio, thus obtaining at least two channels of call audio collected by at least two built-in microphones.
[0084] Step 340: Generate spatial audio data based on at least two call audio streams. Spatial audio data refers to audio data that includes the user's real-time location information in space.
[0085] For example, since there may be a lot of environmental noise in the environment where the call audio is collected, after the first terminal collects at least two call audio streams, it first performs noise reduction processing on the at least two call audio streams to obtain at least two noise-reduced call audio streams, and then generates spatial audio data based on the at least two noise-reduced call audio streams. For example, the aforementioned environmental noise includes sounds other than human voices.
[0086] For example, when the call audio is a conversation between users, in order to make the human voice in the call audio clearer, the human voice in the call audio can be enhanced, and then the noise in at least two enhanced call audios can be removed to obtain at least two noise-reduced call audios. Spatial audio data is generated based on the at least two noise-reduced call audios. After the human voice enhancement processing, it is easier to separate the noise in the call audio, thereby achieving a better noise reduction effect.
[0087] Optionally, the first terminal determines the real-time location information of the first user relative to the first terminal based on at least two audio channels; and converts the at least two audio channels into at least two-channel audio signals; and generates spatial audio data based on the real-time location information and the at least two-channel audio signals.
[0088] Optionally, at least two microphones are at least two built-in microphones; with the recording function enabled on at least two built-in microphones, the user's voice during the call is captured in real time through the at least two built-in microphones to obtain at least two audio streams. For example, if no external microphone is connected, only the built-in microphones are used to capture the user's voice.
[0089] For example, the real-time location information of the first user includes at least one of distance information and direction information of the first user relative to the first terminal. The distance information may refer to relative distance; the direction information may refer to relative direction or azimuth angle. For example, the distance information of the first user may include at least one of the distance between the first user and the microphone, and the distance between the first user and the first terminal; the direction information of the first user may include at least one of the direction information of the first user relative to the microphone, the direction information of the first user relative to the first terminal, the azimuth angle information of the first user relative to the microphone, and the azimuth angle information of the first user relative to the first terminal.
[0090] For example, when at least three audio streams of a call are collected, the first terminal 110 can also determine a set of spatial location information based on every two audio streams from the at least three audio streams, thus obtaining at least two sets of spatial location information; and generate the real-time location information of the first user based on the at least two sets of spatial location information. For example, the spatial location information includes the position and direction information of the sound source relative to the first terminal, and the average of the position and direction information in the at least two sets of spatial location information is calculated to obtain a set of spatial location information as the real-time location information of the first user.
[0091] For example, the first terminal uses only the position of the built-in microphone to determine the real-time location information of the first user. Alternatively, the first terminal uses the positions of both the built-in and external microphones to determine the real-time location of the first user. The preset positional relationship between the external microphone and the user's vocal tract can be determined based on the connection method of the external microphone. Then, the real-time location information of the first user is determined based on the known position of the built-in microphone and the preset positional relationship. For example, when wired headphones are connected, the preset positional relationship includes the distance between the microphone on the wired headphones and the user's vocal tract.
[0092] For example, the first terminal can process N channels of call audio to generate an N-channel audio signal, where N is a positive integer greater than 1; or, process N channels of call audio to generate an M-channel audio signal, where M is a positive integer greater than N; or, process N channels of call audio to generate a P-channel audio signal, where P is a positive integer less than N.
[0093] It should be noted that the steps of generating the real-time location information of the first user and generating at least two channels of audio signals can be performed simultaneously or sequentially. In the embodiments of this application, the execution order of the two is not limited.
[0094] After generating real-time location information and at least two audio channels, the first terminal generates spatial audio data containing at least two audio channels and the real-time location information. For example, the first terminal combines the first user's real-time location information and the at least two audio channels to generate spatial audio data; for instance, it can combine the first user's real-time location information and the at least two audio channels according to a timestamp correspondence. For example, in at least two audio streams, a set of real-time location information is calculated for each 100-millisecond (ms) audio segment, and the audio segments within that 100ms are combined with that set of real-time location information.
[0095] Step 360: Send the spatial audio data to the second terminal.
[0096] The first terminal transmits spatial audio data to the second terminal via wired or wireless communication. The second terminal is a user device that is in the same real-time communication as the first terminal; for example, the second terminal is a terminal that is conducting a voice or audio / video call with the first terminal. Optionally, the second terminal may include at least one, that is, there may be one, two, or more second terminals.
[0097] For example, the first terminal encodes the spatial audio data and sends the encoded spatial audio data to the second terminal.
[0098] In summary, the call audio processing method provided in this embodiment involves a first terminal in a real-time call acquiring the voice of a first user through at least two microphones to obtain at least two call audio streams. Since the at least two microphones are positioned at different locations on the terminal, the terminal can determine the real-time location information of the first user based on these at least two call audio streams. This generates spatial audio data containing the real-time location information of the first user, which is then sent to a second terminal in the call. The second terminal then recreates the spatial audio field based on the spatial audio data, allowing the spatial audio played by the second terminal to reproduce the sound effect when the sound source propagates to the first terminal. In the scenario where the first user uses the first terminal to converse with the second user, the second user using the second terminal experiences an immersive feeling of face-to-face conversation with the first user.
[0099] In the scenario of real-time acquisition and transmission of user voice, the location of the first user sometimes changes continuously and sometimes changes discontinuously. The first terminal can update the location information of the first user to the second terminal when the location of the first user changes, and will not update the location information of the first user to the second terminal when the location of the first user does not change.
[0100] Therefore, step 340 above can be implemented as steps 342 to 346, such as... Figure 5 As shown, the steps are as follows:
[0101] Step 342: Based on at least two audio channels of the call at the first moment, determine the user's first real-time location information relative to the first terminal at the first moment; and convert the at least two audio channels of the call into audio signals with at least two channels.
[0102] Step 344: In response to the difference between the first real-time location information and the second real-time location information, generate first spatial audio data containing at least two audio channels and the first real-time location information.
[0103] The first terminal compares the first real-time location information with the second real-time location information. If the first real-time location information and the second real-time location information are different, it generates first spatial audio data containing at least two audio channels and the first real-time location information.
[0104] The second real-time location information is the user's real-time location relative to the first terminal at a second moment, where the second moment is the moment preceding the first moment.
[0105] Step 346: In response to the first real-time location information being the same as the second real-time location information, generate second spatial audio data containing at least two audio channels.
[0106] The second spatial audio data is used to instruct the second terminal to render and generate spatial audio using the second real-time location information at the second moment.
[0107] In other embodiments, the positional change between the first real-time location information and the second real-time location information is greater than a change threshold, which is a positional change that cannot be ignored. The first terminal generates first spatial audio data containing at least two channels of audio signals and the first real-time location information. For example, the positional change between the first real-time location information and the second real-time location information being greater than the change threshold can be a distance change in relative distance greater than a distance change threshold, and / or a directional change in relative direction greater than a directional change threshold.
[0108] The positional change between the first real-time location information and the second real-time location information is less than or equal to a change threshold, which is a negligible positional change. The first terminal generates second spatial audio data containing audio signals of at least two channels. For example, the positional change between the first real-time location information and the second real-time location information being less than or equal to the change threshold can be a distance change in relative distance being less than or equal to a distance change threshold, and a directional change in relative direction being less than or equal to a directional change threshold.
[0109] In summary, the call audio processing method provided in this embodiment allows the first terminal to update the real-time location information of the first user to the second terminal when the real-time location information of the first user changes at the current moment. If the real-time location information of the first user does not change at the current moment, the real-time location information of the first user in the second terminal is not updated, so that the second terminal still uses the real-time location information of the previous moment. This can ensure the restoration of spatial audio and reduce the transmission resource occupation during real-time calls.
[0110] Another scenario exists where the first terminal is connected to at least one external microphone. For example, if a wired headset is connected to the first terminal, and the headset has a microphone, this microphone is the external microphone connected to the first terminal. Normally, when an external microphone is connected, the first terminal disables its built-in microphone. However, in this embodiment, the first terminal will simultaneously enable both the built-in and external microphones. Figure 6 The diagram shown is a flowchart of a call audio processing method provided in an exemplary embodiment of this application, illustrating the method implementation in the aforementioned case. This method can be applied to... Figure 1 In the first terminal shown, the method includes:
[0111] Step 400: Collect the user's voice during the call in real time using at least two built-in microphones to obtain at least two audio streams of the call.
[0112] The first terminal has at least two built-in microphones located in different positions; for example, on a mobile terminal, built-in microphone 1 is located above the screen, and built-in microphone 2 is located below the screen.
[0113] With the recording function of at least two built-in microphones enabled, the first terminal collects the user's voice during the real-time call through each built-in microphone, and timestamps the collected call audio to record the time of collection of the call audio, thus obtaining at least two channels of call audio collected by at least two built-in microphones.
[0114] Step 420: Collect the user's voice during the call in real time using at least one external microphone to obtain at least one call audio, wherein the at least one call audio is the call audio collected at the same time as at least two call audios.
[0115] The first terminal enables the recording function of at least one external microphone. While collecting call audio through the built-in microphone, it also collects the user's voice during the real-time call through the external microphone, thus obtaining at least one call audio stream.
[0116] Step 440: Based on at least two audio feeds, determine the user's real-time location information in space relative to the first terminal.
[0117] Optionally, the real-time location information of the first user includes at least one of distance information and direction information of the first user relative to the first terminal. For example, the distance information may refer to relative distance; the direction information may refer to relative direction or azimuth. For instance, the distance information of the first user may include at least one of the distance between the first user and the microphone, and the distance between the first user and the first terminal; the direction information of the first user may include at least one of the direction information of the first user relative to the microphone, the direction information of the first user relative to the first terminal, the azimuth information of the first user relative to the microphone, and the azimuth information of the first user relative to the first terminal.
[0118] For example, when at least three audio streams of a call are collected, the first terminal 110 can also determine a set of spatial location information based on every two audio streams from the at least three audio streams, thus obtaining at least two sets of spatial location information; and generate the real-time location information of the first user based on the at least two sets of spatial location information. For example, the spatial location information includes the position and direction information of the sound source relative to the first terminal, and the average of the position and direction information in the at least two sets of spatial location information is calculated to obtain a set of spatial location information as the real-time location information of the first user.
[0119] Step 460: Generate spatial audio data based on at least one call audio and the real-time location information of the first user.
[0120] For example, the first terminal combines at least one audio stream of a call with the real-time location information of the first user to generate spatial audio data, wherein the spatial audio data includes at least one audio stream of a call and the real-time location information of the first user. For instance, the first terminal performs noise reduction processing on one audio stream of a call captured by an external microphone to obtain a noise-reduced mono audio stream, and generates spatial audio data containing the mono audio stream and the real-time location information.
[0121] Alternatively, the first terminal can convert at least one audio stream into at least two audio channels, generating spatial audio data containing at least two audio channels and real-time location information. For example, the first terminal can perform noise reduction processing on one audio stream captured by an external microphone to obtain a mono audio stream, then copy or cut the mono audio stream according to frequency bands to generate a stereo audio signal; generating spatial audio data containing the stereo audio signal and real-time location information.
[0122] It should be noted that when converting at least one audio channel into at least two audio channels, the generation of the at least two audio channels can be performed simultaneously with step 440, or sequentially with step 440, and then spatial audio data can be generated based on the at least two audio channels and real-time location information.
[0123] In some embodiments, the first terminal may also generate a mono audio signal or a multi-channel audio signal based on at least two call audio streams and at least one call audio stream. Multi-channel refers to two or more audio channels. For example, the first terminal synthesizes at least two call audio streams and at least one call audio stream to generate a mono audio signal; based on the mono audio signal and spatial information, it generates spatial audio data. Alternatively, the first terminal treats each call audio stream as a single audio channel, combines at least two call audio streams and at least one call audio stream to generate at least a three-channel audio signal; based on the at least three-channel audio signal and spatial information, it generates spatial audio data. Another example is that the first terminal synthesizes at least two call audio streams and at least one call audio stream to generate a mono audio signal; it then cuts the mono audio signal according to at least two preset frequency bands to generate at least a two-channel audio signal; based on the at least two-channel audio signal and spatial information, it generates spatial audio data.
[0124] Step 480: Send the spatial audio data to the second terminal.
[0125] The first terminal transmits spatial audio data to the second terminal via wired or wireless communication. The second terminal is a user device that is in the same real-time call as the first terminal; for example, the second terminal is a terminal that is conducting a real-time call or real-time video call with the first terminal. Optionally, the second terminal may include at least one, that is, there may be one, two, or more second terminals.
[0126] For example, the first terminal encodes the spatial audio data and sends the encoded spatial audio data to the second terminal.
[0127] For example, a mobile terminal has two built-in microphones, and a headset is connected to the mobile terminal. The headset has an external microphone. The method for processing call audio can be illustrated as follows: Figure 7 The steps in the process shown are as follows:
[0128] Step 41: Turn on the recording function of the headphones.
[0129] Step 42: Collect call audio using the external microphone on the headset.
[0130] Step 43: Enable the recording function on your mobile device.
[0131] Step 44: Collect two audio streams of the call using at least two built-in microphones on the mobile terminal; calculate the real-time location information of the first user based on the two audio streams; and generate spatial audio data containing the real-time location information and the two audio streams of the call.
[0132] Step 45: Overlay the call audio collected by the external microphone with the two call audios collected by the built-in microphone, as well as the real-time location information of the first user, to generate spatial audio data.
[0133] Step 46: Send spatial audio data to the second terminal.
[0134] In other embodiments, the first terminal determines the user's first real-time location information relative to the first terminal at a first moment based on at least two call audios at a first moment; in response to the first real-time location information being different from the second real-time location information, it generates first spatial audio data containing at least one call audio and the first real-time location information; or, in response to the first real-time location information being different from the second real-time location information, it generates second spatial audio data containing at least one call audio.
[0135] Alternatively, the first terminal determines the user's first real-time location information relative to the first terminal at the first moment based on at least two call audios at the first moment, and converts at least one call audio into an audio signal with at least two channels; in response to a difference between the first real-time location information and the second real-time location information, it generates first spatial audio data containing the audio signal with at least two channels and the first real-time location information; or, in response to a difference between the first real-time location information and the second real-time location information, it generates second spatial audio data containing the audio signal with at least two channels.
[0136] The second real-time location information is the user's real-time location relative to the first terminal at a second moment, and the second moment is the moment before the first moment; the second spatial audio data is used to instruct the second terminal to render and generate spatial audio using the second real-time location information at the second moment.
[0137] That is, when the real-time location information of the first user changes, the real-time location information of the first user is updated to the second terminal; otherwise, the real-time location information of the first user is not updated to the second terminal, so as to reduce the consumption of transmission resources during real-time calls.
[0138] In summary, the call audio processing method provided in this embodiment involves a first terminal in a real-time call acquiring at least two channels of call audio from a sound source using at least two built-in microphones. These two built-in microphones are positioned at different locations within the first terminal, allowing the first terminal to determine the real-time location information of the first user based on these at least two channels of call audio. Furthermore, at least one channel of call audio is acquired using an external microphone. Based on this at least one channel of call audio and the real-time location information, spatial audio data containing the real-time location information is generated. This spatial audio data is then sent to a second terminal in the real-time call. The second terminal then recreates the spatial audio field based on the spatial audio data, allowing the spatial audio played by the second terminal to reproduce the sound effect of a human voice propagating to the first terminal. In the scenario where the first user uses the first terminal to converse with the second user, the second user using the second terminal experiences an immersive feeling of face-to-face conversation with the first user.
[0139] Secondly, since external microphones are usually specifically connected to the terminal to be closer to the user's voice, the call audio captured by an external microphone is of better quality. Therefore, the spatial audio data generated from the call audio captured by an external microphone can improve the quality of spatial audio.
[0140] As shown above, when the first terminal is not connected to an external microphone, it can be used as follows: Figure 4 The illustrated embodiment; when an external microphone is connected to the first terminal, it can be used as follows: Figure 5 The example shown.
[0141] Figure 8 A flowchart illustrating a method for processing call audio according to an exemplary embodiment of this application is shown. This method can be applied to... Figure 1 In the second terminal shown, the method includes:
[0142] Step 620: Receive spatial audio data sent by the first terminal, which includes the user's real-time location information in space and the user's call audio. The call audio is obtained by collecting the user's voice during the call. The first terminal and the second terminal are user devices in the same real-time call.
[0143] For example, the second terminal receives spatial audio data sent by the first terminal via a wired or wireless network.
[0144] Step 640: Based on the spatial audio data, generate at least two channels of spatial audio corresponding to at least two speakers.
[0145] Optionally, the spatial audio data includes the real-time location information of the first user and at least two audio channels; the second terminal uses the real-time location information to render the at least two audio channels to generate at least two spatial audio channels.
[0146] Optionally, the spatial audio data includes the real-time location information of the first user and at least one call audio; the second terminal generates at least two audio signals corresponding to at least two built-in speakers based on at least one call audio; and renders the at least two audio signals using the real-time location information to generate at least two spatial audio.
[0147] For example, in response to the spatial audio data at the first moment containing at least two channels of audio signals and first real-time location information, the second terminal updates the real-time location information of the first user from the second real-time location information to the first real-time location information; and renders the at least two channels of audio signals using the first real-time location information to generate at least two channels of spatial audio. In response to the spatial audio data at the first moment containing at least two channels of audio signals, the second terminal obtains the second real-time location information; and renders the at least two channels of audio signals using the second real-time location information to generate at least two channels of spatial audio. Here, the first real-time location information is the user's real-time location information relative to the first terminal at the first moment; the second real-time location information is the user's real-time location information relative to the first terminal at the second moment, where the second moment is the moment preceding the first moment.
[0148] For example, in response to the spatial audio data at the first moment including at least one call audio and first real-time location information, the second terminal updates the real-time location information of the first user from the second real-time location information to the first real-time location information; and generates at least two-channel audio signals based on at least one call audio; and renders the at least two-channel audio signals using the first real-time location information to generate at least two-channel spatial audio. In response to the spatial audio data at the first moment including at least one call audio, the second terminal obtains the second real-time location information, and generates at least two-channel audio signals based on at least one call audio; and renders the at least two-channel audio signals using the second real-time location information to generate at least two-channel spatial audio. Wherein, the first real-time location information is the user's real-time location information relative to the first terminal at the first moment; the second real-time location information is the user's real-time location information relative to the first terminal at the second moment, and the second moment is the moment preceding the first moment.
[0149] For example, the second terminal obtains the current channel mode of the audio playback, determines that the channel mode of the audio signal in the spatial audio data does not conform to the current channel mode, performs conversion processing on the audio signal to generate a multi-channel audio signal that matches the current number of channels, and then renders the multi-channel audio signal based on real-time position information to generate multi-channel spatial audio; or determines that the channel mode of the audio signal in the spatial audio data conforms to the current channel mode, and directly renders the audio signal of each channel based on real-time position information to generate multi-channel spatial audio.
[0150] For example, the current channel mode is determined based on the number of built-in speakers of the second terminal; or, the current channel mode is determined based on the selection settings of the second user, and the number of channels used in the current channel mode is less than or equal to the number of built-in speakers of the second terminal.
[0151] Optionally, the real-time location information of the first user includes at least one of distance information and direction information of the first user relative to the first terminal. For example, the distance information may refer to relative distance; the direction information may refer to relative direction or azimuth. For instance, the distance information of the first user may include at least one of the distance between the first user and the microphone, and the distance between the first user and the first terminal; the direction information of the first user may include at least one of the direction information of the first user relative to the microphone, the direction information of the first user relative to the first terminal, the azimuth information of the first user relative to the microphone, and the azimuth information of the first user relative to the first terminal.
[0152] Step 660: Play at least two channels of spatial audio through at least two speakers.
[0153] In some embodiments, the second terminal may also play at least two channels of spatial audio via an external speaker, such as headphones; in this case, the at least two channels of spatial audio are generated based on the headphone's channel pattern.
[0154] In summary, the call audio processing method provided in this embodiment involves a first terminal in a real-time call acquiring user voice data through at least two microphones to obtain at least two audio streams. These microphones are positioned at different locations within the first terminal. Therefore, the first terminal can determine the real-time location information of the first user based on these at least two audio streams, generating spatial audio data containing the real-time location information. This spatial audio data is then sent to a second terminal in the real-time call. The second terminal then recreates the spatial audio field based on the spatial audio data, allowing the spatial audio played by the second terminal to reproduce the sound effect of human voice propagating to the first terminal. In the scenario where the first user uses the first terminal to converse with the second user, the second user using the second terminal experiences an immersive feeling of face-to-face conversation with the first user.
[0155] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.
[0156] Figure 9 This illustration shows a structural block diagram of a call audio processing apparatus provided in an exemplary embodiment of this application. The apparatus can be implemented as all or part of a first terminal via software, hardware, or a combination of both, wherein at least two microphones of the first terminal are disposed in different locations. The apparatus includes:
[0157] The acquisition module 720 is used to acquire the user's voice during a call in real time through the at least two microphones, and obtain at least two audio streams of the call.
[0158] The generation module 740 is used to generate spatial audio data based on the at least two call audios, wherein the spatial audio data refers to audio data containing the user's real-time location information in space.
[0159] The sending module 760 is used to send the spatial audio data to a second terminal, wherein the second terminal and the first terminal are user equipment in the same real-time call.
[0160] In some embodiments, the generation module 740 is configured to:
[0161] Based on the at least two audio feeds, determine the user's real-time location information relative to the first terminal in space; and
[0162] Convert the at least two audio channels into at least two audio signals;
[0163] The spatial audio data is generated based on the real-time location information and the audio signals of at least two channels.
[0164] In some embodiments, the generation module 740 is configured to:
[0165] Based on the at least two audio streams from the first moment, determine the user's first real-time location information relative to the first terminal at the first moment; and
[0166] Convert the at least two audio channels into at least two audio signals;
[0167] In response to the difference between the first real-time location information and the second real-time location information, first spatial audio data containing the at least two audio channels and the first real-time location information is generated; wherein, the second real-time location information is the real-time location information of the user relative to the first terminal at a second time, and the second time is the time before the first time.
[0168] In some embodiments, the generation module 740 is configured to:
[0169] Based on the at least two audio streams from the first moment, determine the user's first real-time location information relative to the first terminal at the first moment; and
[0170] Convert the at least two audio channels into at least two audio signals;
[0171] In response to the first real-time location information being the same as the second real-time location information, second spatial audio data containing the audio signals of the at least two channels is generated. The second spatial audio data is used to instruct the second terminal to render and generate spatial audio using the second real-time location information at the second moment.
[0172] In some embodiments, the real-time location information of the first user includes at least one of the distance information and direction information of the user relative to the first terminal.
[0173] In some embodiments, the at least two microphones are at least two built-in microphones;
[0174] The generation module 740 is used to collect the user's voice during a call in real time through the at least two built-in microphones when the recording function of the at least two built-in microphones is enabled, and obtain the at least two channels of call audio.
[0175] In some embodiments, the first terminal is further connected to at least one external microphone;
[0176] The acquisition module 720 is configured to, when the recording function of the at least one external microphone and the at least two internal microphones is enabled, acquire the user's voice during a call in real time through the at least one external microphone to obtain at least one audio stream; and acquire the user's voice during a call in real time through the at least two internal microphones to obtain at least two audio streams; wherein the at least one audio stream is the audio stream acquired at the same time as the at least two audio streams.
[0177] The generation module 740 is configured to determine the real-time location information of the user relative to the first terminal in space based on the at least two call audios; and generate the spatial audio data based on the at least one call audio and the real-time location information.
[0178] The transmitting module 760 is used to transmit the spatial audio data to the second terminal.
[0179] Figure 10 This illustration shows a structural block diagram of a call audio processing apparatus provided in an exemplary embodiment of this application. The apparatus can be implemented as all or part of a second terminal via software, hardware, or a combination of both, wherein at least two speakers of the second terminal are disposed in different locations. The apparatus includes:
[0180] The receiving module 920 is used to receive spatial audio data sent by the first terminal. The spatial audio data includes the user's real-time location information in space and the user's call audio. The call audio is obtained by collecting the user's voice during the call. The first terminal and the second terminal are user devices in the same real-time call.
[0181] The generation module 940 is used to generate at least two channels of spatial audio corresponding to the at least two speakers based on the spatial audio data;
[0182] The playback module 960 is used to play the at least two channels of spatial audio through the at least two speakers.
[0183] In some embodiments, the spatial audio data includes the real-time location information of the first user and at least two channels of audio signal;
[0184] The generation module 940 is used to render the at least two-channel audio signal using the real-time location information to generate the at least two-channel spatial audio.
[0185] In some embodiments, the spatial audio data includes the real-time location information of the first user and at least one audio call.
[0186] The generation module 940 is used to generate at least two audio signals corresponding to at least two built-in speakers based on the at least one call audio; and to render the at least two audio signals using the real-time location information to generate the at least two spatial audio channels.
[0187] In some embodiments, the real-time location information of the first user includes at least one of the distance information and direction information of the user relative to the first terminal.
[0188] Figure 11 This illustration shows a schematic diagram of a computer device provided in an exemplary embodiment of this application. The computer device may be a device that performs a call audio processing method as provided in this application. Exemplarily, the computer device may be a first terminal or a second terminal. Specifically:
[0189] Computer device 1000 includes a central processing unit (CPU) 1001, a system memory 1004 including random access memory (RAM) 1002 and read-only memory (ROM) 1003, and a system bus 1005 connecting the system memory 1004 and the central processing unit 1001. Computer device 1000 also includes a basic input / output system (I / O system) 1006 that facilitates information transfer between various devices within the computer, and a mass storage device 1007 for storing the operating system 1013, application programs 1014, and other program modules 1015.
[0190] The basic input / output system 1006 includes a display 1008 for displaying information and an input device 1009 for user input, such as a mouse or keyboard. Both the display 1008 and the input device 1009 are connected to the central processing unit 1001 via an input / output controller 1010 connected to the system bus 1005. The basic input / output system 1006 may also include the input / output controller 1010 for receiving and processing input from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 1010 also provides output to a display screen, printer, or other types of output devices.
[0191] Mass storage device 1007 is connected to central processing unit 1001 via a mass storage controller (not shown) connected to system bus 1005. Mass storage device 1007 and its associated computer-readable media provide non-volatile storage for computer device 1000. That is, mass storage device 1007 may include computer-readable media (not shown) such as hard disk or compact disc read-only memory (CD-ROM) drive.
[0192] Computer-readable media can include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include RAM, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other solid-state storage technologies, CD-ROM, digital versatile optical disc (DVD), or solid-state drives (SSD), other optical storage, magnetic tape cassettes, magnetic tape, disk storage, or other magnetic storage devices. Random access memory can include resistive random access memory (ReRAM) and dynamic random access memory (DRAM). Of course, those skilled in the art will recognize that computer storage media are not limited to the above-mentioned types. The system memory 1004 and the mass storage device 1007 mentioned above can be collectively referred to as memory.
[0193] According to various embodiments of this application, the computer device 1000 can also be connected to a remote computer on a network, such as the Internet. That is, the computer device 1000 can be connected to the network 1012 via the network interface unit 1011 connected to the system bus 1005, or the network interface unit 1011 can be used to connect to other types of networks or remote computer systems (not shown).
[0194] The aforementioned memory also includes one or more programs, which are stored in the memory and configured to be executed by the CPU to implement the call audio processing method described above.
[0195] This application also provides a computer-readable storage medium storing at least one instruction, which is loaded and executed by a processor to implement the call audio processing method described in the above embodiments.
[0196] Optionally, the computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), solid-state drives (SSDs), or optical discs, etc. The random access memory may include resistive random access memory (ReRAM) and dynamic random access memory (DRAM).
[0197] It should be noted that the call audio processing device provided in the above embodiments is only illustrated by the division of the above functional modules when executing the call audio processing method. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the call audio processing device and the call audio processing method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.
[0198] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0199] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0200] The above description is merely an exemplary embodiment that can be implemented in this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for processing call audio, characterized in that, The method is executed by a first terminal, wherein at least two microphones of the first terminal are positioned at different locations, and the method includes: The user's voice is captured in real time through the at least two microphones to obtain at least two audio streams of the call. Spatial audio data is generated based on the at least two call audio streams. The spatial audio data refers to audio data containing the user's real-time location information in space. Generating spatial audio data based on the at least two call audio streams includes: determining the user's real-time location information in space relative to the first terminal based on the at least two call audio streams; converting the at least two call audio streams into at least two-channel audio signals; and generating the spatial audio data based on the real-time location information and the at least two-channel audio signals. The spatial audio data is sent to a second terminal, which is a user device in the same real-time call as the first terminal. The second terminal is used to render the at least two-channel audio signal based on the user's real-time location information to generate at least two-channel spatial audio, and to restore the user's real-time location information relative to the first terminal in the at least two-channel spatial audio.
2. The method according to claim 1, characterized in that, Determining the user's real-time location information relative to the first terminal in space based on the at least two audio feeds includes: Based on the at least two audio streams of the call at the first moment, determine the first real-time location information of the user relative to the first terminal at the first moment; The process of generating the spatial audio data based on the real-time location information and the at least two-channel audio signals includes: In response to the difference between the first real-time location information and the second real-time location information, first spatial audio data containing the at least two audio channels and the first real-time location information is generated; wherein, the second real-time location information is the real-time location information of the user relative to the first terminal at a second time, and the second time is the time before the first time.
3. The method according to claim 2, characterized in that, The method further includes: In response to the first real-time location information being the same as the second real-time location information, second spatial audio data containing the audio signals of the at least two channels is generated. The second spatial audio data is used to instruct the second terminal to render and generate spatial audio using the second real-time location information at the second moment.
4. The method according to any one of claims 1 to 3, characterized in that, The real-time location information includes at least one of the distance information and direction information of the user relative to the first terminal.
5. The method according to any one of claims 1 to 3, characterized in that, The at least two microphones are at least two built-in microphones; The process of acquiring at least two audio streams of a user's voice during a call in real time using at least two microphones includes: With the recording function enabled on the at least two built-in microphones, the user's voice during the call is captured in real time through the at least two built-in microphones to obtain the at least two audio channels of the call.
6. The method according to claim 5, characterized in that, The first terminal is also connected to at least one external microphone; The method further includes: With the recording function of at least one external microphone and at least two built-in microphones all enabled, the user's voice during the call is captured in real time through the at least one external microphone to obtain at least one audio stream of the call; and The user's voice is captured in real time through the at least two built-in microphones to obtain the at least two audio streams; wherein, the at least one audio stream is the audio stream captured at the same time as the at least two audio streams. Based on the at least two audio recordings, the user's real-time location information relative to the first terminal in space is determined; The spatial audio data is generated based on the at least one audio call and the real-time location information; The spatial audio data is sent to the second terminal.
7. A method for processing call audio, characterized in that, The method is performed by a second terminal, wherein at least two speakers of the second terminal are positioned at different locations, and the method includes: The system receives spatial audio data sent by a first terminal. The spatial audio data includes the user's real-time location information in space and the user's call audio. The call audio is obtained by collecting the user's voice during a call. The first terminal and the second terminal are user devices in the same real-time call. Based on the spatial audio data, at least two channels of spatial audio corresponding to the at least two speakers are generated; when the spatial audio data includes the real-time location information and at least two channels of audio signals, the at least two channels of audio signals are obtained based on at least two channels of call audio, which are used to determine the user's real-time location information relative to the first terminal in space; generating at least two channels of spatial audio corresponding to the at least two speakers based on the spatial audio data includes: rendering the at least two channels of audio signals using the real-time location information to generate the at least two channels of spatial audio; The at least two channels of spatial audio are played through the at least two speakers.
8. The method according to claim 7, characterized in that, The spatial audio data includes the real-time location information and at least one audio call. The step of generating at least two channels of spatial audio corresponding to the at least two built-in speakers based on the spatial audio data includes: Based on the at least one call audio, generate at least two audio signals corresponding to the at least two built-in speakers; The real-time location information is used to render the audio signals of the at least two channels to generate the spatial audio of the at least two channels.
9. The method according to claim 7 or 8, characterized in that, The real-time location information includes at least one of the distance information and direction information of the user relative to the first terminal.
10. A device for processing call audio, characterized in that, The device is disposed in a first terminal, wherein at least two microphones of the first terminal are disposed in different positions, and the device includes: The acquisition module is used to acquire the user's voice during a call in real time through the at least two microphones to obtain at least two audio streams of the call; A generation module is configured to generate spatial audio data based on the at least two call audio streams. The spatial audio data refers to audio data containing real-time location information of the user in space. Generating spatial audio data based on the at least two call audio streams includes: determining the real-time location information of the user relative to the first terminal in space based on the at least two call audio streams; converting the at least two call audio streams into at least two-channel audio signals; and generating the spatial audio data based on the real-time location information and the at least two-channel audio signals. The sending module is used to send the spatial audio data to a second terminal, which is a user device in the same real-time call as the first terminal. The second terminal is used to render the at least two-channel audio signal based on the user's real-time location information to generate at least two-channel spatial audio, and to restore the user's real-time location information relative to the first terminal in the at least two-channel spatial audio.
11. A device for processing call audio, characterized in that, The device is disposed in a second terminal, wherein at least two speakers of the second terminal are disposed in different positions, and the device includes: The receiving module is used to receive spatial audio data sent by the first terminal. The spatial audio data includes the user's real-time location information in space and the user's call audio. The call audio is obtained by collecting the user's voice during the call. The first terminal and the second terminal are user devices in the same real-time call. A generation module is configured to generate at least two channels of spatial audio corresponding to the at least two speakers based on the spatial audio data; when the spatial audio data includes the real-time location information and at least two channels of audio signals, the at least two channels of audio signals are obtained by conversion based on at least two channels of call audio, and the at least two channels of call audio are used to determine the real-time location information of the user relative to the first terminal in space; the step of generating at least two channels of spatial audio corresponding to the at least two speakers based on the spatial audio data includes: rendering the at least two channels of audio signals using the real-time location information to generate the at least two channels of spatial audio; A playback module for playing the at least two channels of spatial audio through the at least two speakers.
12. A terminal, characterized in that, The terminal includes a processor and a memory connected to the processor. The memory stores program instructions. When the processor executes the program instructions, it implements the call audio processing method as described in any one of claims 1 to 6, or the call audio processing method as described in any one of claims 7 to 9.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program instructions that, when executed by a processor, implement the call audio processing method as described in any one of claims 1 to 6, or the call audio processing method as described in any one of claims 7 to 9.
14. A computer program product, characterized in that, The computer program product includes computer instructions stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the call audio processing method as described in any one of claims 1 to 6, or the call audio processing method as described in any one of claims 7 to 9.
Citation Information
Patent Citations
Two stage audio focus for spatial audio processing
CN110537221A
Device and method for Rotating camera and microphone configurations
CN113014797A
Spatial audio data processing method and device, storage medium and electronic equipment
CN113660063A
Gain Control in Spatial Audio Systems
US20190289420A1