Audio processing method and device based on Android system, electronic equipment and storage medium
By acquiring signals from the acquisition link and playback link of the audio sharing terminal on the Android system, echo cancellation processing is performed, and the processed signal is sent to the second application terminal, the problem of not being able to play shared audio through its own speaker during audio sharing is solved, and high-quality audio sharing and playback is achieved.
Patent Information
- Application Number
- CN202311612314.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-28
- Publication Date
- 2025-05-30
AI Technical Summary
When sharing audio on Android systems, echo cancellation can usually be performed through the audio signal in front of the speaker's physical output interface, and the shared audio cannot be played through its own speaker while sharing audio.
By collecting the near-end signal of the call stream from the acquisition link of the first application terminal, and obtaining the near-end signal of the shared audio from the playback link, performing echo cancellation processing, obtaining the near-end target call stream audio signal, and sending the shared audio near-end signal and the near-end target call stream audio signal to the second application terminal for playback processing.
It realizes the shared audio playback through its own speaker while sharing audio, avoids echo problems and improves call quality.
Smart Images

Figure CN120075362A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio data, and in particular, to an audio processing method, an audio processing device, an electronic device, and a computer-readable storage medium based on the Android system. Background Art
[0002] Audio conferences have gradually been integrated into people's lives and work, providing a fast and intuitive communication method for users in different locations. With the rapid development of audio conferences, audio sharing technology is also becoming increasingly mature. Audio sharing technology refers to the first application terminal sharing the audio and video signals of a third-party audio and video application (a conference application that is not currently in use) to each second application terminal.
[0003] Currently, when performing audio sharing on the Android system, usually only the audio signal before entering the physical output interface of the speaker is subjected to echo cancellation, and only the audio and video signals of a third-party audio and video application (a conference application that is not currently in use) are played to prevent each second application terminal from receiving the audio signal sent by itself, thereby generating an echo. Further, if shared audio is played through the speaker while performing audio sharing, the second application terminal will hear the double accent of the two shared audios, that is, the second application terminal first hears the audio of the shared audio and then hears the echo of the shared audio generated by the call flow end. Summary of the Invention
[0004] The main technical problem to be solved by this application is to provide an audio processing method, an audio processing device, an electronic device, and a computer-readable storage medium based on the Android system, which can realize playing the shared audio through its own speaker while performing audio sharing.
[0005] To solve the above technical problem, a technical solution adopted by this application is: providing an audio processing method based on the Android system, the method includes: collecting a call flow proximal signal from a collection link included in a first application terminal carrying the application program and obtaining a shared audio proximal signal from a playback link of the first application terminal; performing echo cancellation processing on the call flow proximal signal according to the shared audio proximal signal to obtain a proximal target call flow audio signal; sending the shared audio proximal signal and the proximal target call flow audio signal to a second application terminal carrying the application program for playback processing.
[0006] To solve the above technical problems, another technical solution adopted by this application is: to provide an audio processing device, which includes an acquisition module, an echo processing module, and a playback module. The acquisition module is used to acquire the proximal signal of the call stream from the acquisition link included in the first application terminal carrying the application program and obtain the proximal signal of the shared audio from the playback link of the first application terminal; the echo processing module is used to perform echo cancellation processing on the proximal signal of the call stream according to the proximal signal of the shared audio to obtain the proximal target call stream audio signal; the playback module is used to send the proximal signal of the shared audio and the proximal target call stream audio signal to the second application terminal carrying the application program for playback processing.
[0007] To solve the above technical problems, another technical solution adopted by this application is: to provide an electronic device, which includes a memory and a processor. The memory stores program instructions, and the processor retrieves the program instructions from the memory to execute the above-mentioned audio processing method based on the Android system.
[0008] To solve the above technical problems, another technical solution adopted by this application is: to provide a computer-readable storage medium including stored program data, and when the program data is executed by a processor, it is used to implement the above-mentioned audio processing method based on the Android system.
[0009] The beneficial effects of this application are as follows: acquire the proximal signal of the call stream from the acquisition link included in the first application terminal carrying the application program and obtain the proximal signal of the shared audio from the playback link of the first application terminal; perform echo cancellation processing on the proximal signal of the call stream according to the proximal signal of the shared audio to obtain the proximal target call stream audio signal; send the proximal signal of the shared audio and the proximal target call stream audio signal to the second application terminal carrying the application program for playback processing. By performing echo cancellation processing on the proximal signal of the call stream, it is possible to play the shared audio through its own speaker while sharing the audio. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings, where:
[0011] Figure 1 is a schematic flowchart of an exemplary embodiment of the audio processing method based on the Android system shown in this application;
[0012] Figure 2 isFigure 1 Flow schematic diagram of an exemplary embodiment of step S110 in the illustrated Android system-based audio processing method;
[0013] Figure 3 Is Figure 2 Flow schematic diagram of an exemplary embodiment of step S220 in the illustrated Android system-based audio processing method;
[0014] Figure 4 Is Figure 3 Specific schematic diagram of an exemplary embodiment of step S220 in the illustrated Android system-based audio processing method;
[0015] Figure 5 Is Figure 1 Flow schematic diagram of an exemplary embodiment before step S120 in the illustrated Android system-based audio processing method;
[0016] Figure 6 Is Figure 5 Flow schematic diagram of an exemplary embodiment of step S420 in the illustrated Android system-based audio processing method;
[0017] Figure 7 Is Figure 1 Flow schematic diagram of an exemplary embodiment after step S110 in the illustrated Android system-based audio processing method;
[0018] Figure 8 Is Figure 1 Flow schematic diagram of an exemplary embodiment of step S120 in the illustrated Android system-based audio processing method;
[0019] Figure 9 Is a specific schematic diagram of an application scenario of the Android system-based audio processing method illustrated in the present application;
[0020] Figure 10 Is a flow schematic diagram of an exemplary embodiment of the audio processing device provided in the present application;
[0021] Figure 11 Is a structural schematic diagram of an embodiment of the electronic device provided in the present application;
[0022] Figure 12 Is a structural schematic diagram of an embodiment of the computer-readable storage medium provided in the present application. Detailed implementation manners
[0023] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. It can be understood that the specific embodiments described herein are only used to explain the present application, rather than limiting the present application. Additionally, it should be noted that for the convenience of description, only the parts related to the present application rather than all the structures are shown in the accompanying drawings. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0024] First of all, it should be noted that the present application proposes an audio processing method based on the Android system, which can be applied to echo cancellation processing during audio sharing. Specifically, it performs echo cancellation processing on the collected near-end call flow audio signal according to the obtained shared audio near-end signal to obtain the near-end target call flow audio signal, and plays the shared audio near-end signal and the near-end target call flow audio signal through the second application terminal for playback processing, so as to realize playing the shared audio through its own speaker while sharing the audio. Please refer to Figure 1 , Figure 1 is a schematic flowchart of an exemplary embodiment of the audio processing method based on the Android system shown in the present application. The audio processing method based on the Android system in this embodiment can be applied to an audio processing device, such as a conference application, or can also be applied to a server with data processing capabilities.
[0025] Specifically, the audio processing method based on the Android system in this embodiment includes the following steps:
[0026] S110: Collect the near-end call flow audio signal from the acquisition link included in the first application terminal carrying the application program and obtain the shared audio near-end signal from the playback link of the first application terminal.
[0027] The application program refers to a software program with specific functions running on the Android system. Exemplarily, the application program can be a conference software, a social software, etc.
[0028] The first application terminal refers to the terminal that shares the shared audio in the audio conference. Exemplarily, the first application terminal can be intelligent devices such as a mobile phone, a personal computer, and a conference tablet that can carry the application program.
[0029] The second application terminal refers to the terminal that receives the shared audio shared by the first application terminal in the audio conference. Exemplarily, the second application terminal can also be a mobile phone, a personal computer, a tablet, etc. that can carry the application program. It should be noted that there can be multiple second application terminals, that is, the shared audio shared by the first application terminal in the audio conference can be played and processed on multiple second application terminals simultaneously.
[0030] The acquisition link refers to the link through which the acquisition device configured on the first application terminal acquires audio signals. Among them, the acquisition device configured on the first application terminal can be a microphone.
[0031] The playback link includes the link for playing the audio signals transmitted from the second application terminal to the first application terminal and the audio signals of the third-party audio and video applications on the first application terminal through the playback device configured on the first application terminal. Among them, the playback device configured on the first application terminal can be a speaker, a sound system, etc.
[0032] The near-end signal of the call flow refers to the audio signal directly acquired through the acquisition device configured on the first application terminal. Exemplarily, the near-end signal of the call flow includes the shared audio near-end signal acquired through the acquisition device and other external sounds except the audio signals played by the playback device. Exemplarily, taking the application program in the embodiments of the present application as a conferencing software, the audio signals acquired by the acquisition device include the voices of the users of the first application terminal in the current conference.
[0033] The shared audio near-end signal refers to the audio signal that is about to enter the physical output interface of the playback device configured on the first application terminal. Exemplarily, the shared audio near-end signal includes the audio signals of the third-party audio and video applications on the first application terminal and the audio signals of the remote target call flow. Exemplarily, taking the application program in the embodiments of the present application as a conferencing software, the shared audio near-end signal includes the voices of the users of the second application terminal in the current conference.
[0034] The audio processing device acquires the shared audio near-end signal and other external sounds except the audio signals played by the playback device from the acquisition link of the first application terminal where audio sharing is in progress, and acquires the audio signals of the third-party audio and video applications on the first application terminal from the playback link of the first application terminal.
[0035] S120: Perform echo cancellation processing on the near-end signal of the call flow according to the shared audio near-end signal to obtain the near-end target call flow audio signal.
[0036] Echo refers to the situation where the second application terminal receives the audio signals it has sent due to the reflection or propagation delay of the audio signals during the transmission process.
[0037] The near-end target call flow audio signal refers to other external sound signals acquired through the acquisition device except the audio signals played by the playback device. Exemplarily, taking the application program in the embodiments of the present application as a conferencing software, the near-end target call flow audio signal includes the voice signal of the user of the first application terminal in the current conference.
[0038] While the shared audio is being played through a physical speaker, the microphone continuously captures the sound of the shared audio being played in the speaker. Therefore, the proximal signal of the call stream captured by the microphone includes the proximal signal of the shared audio and other external sound signals other than the audio signal played by the playback device. At this time, if the echo cancellation is not performed on the proximal signal of the call stream, the second application terminal will receive the accent of the proximal signal of the shared audio, affecting the call quality.
[0039] The audio processing device performs echo cancellation processing on the proximal signal of the call stream according to the proximal signal of the shared audio and the distal target call stream audio signal to obtain the proximal target call stream audio signal.
[0040] S130: Send the proximal signal of the shared audio and the proximal target call stream audio signal to the second application terminal hosting the application for playback processing.
[0041] After obtaining the proximal target call stream audio signal, the proximal target call stream audio signal is encoded to obtain a call stream encoding, and the proximal signal of the shared audio is encoded to obtain a proximal encoding of the shared audio. The call stream encoding and the proximal encoding of the shared audio are transmitted to the second application terminal through the server, and the second application terminal decodes and plays them.
[0042] The audio processing device obtains the proximal signal of the call stream and the proximal signal of the shared audio from the first application terminal, processes the obtained audio signals to obtain the proximal target call stream audio signal, and the second application terminal performs playback processing on the proximal signal of the shared audio and the proximal target call stream audio signal.
[0043] It can be seen that the audio processing method based on the Android system in the embodiment of the present application collects the proximal signal of the call stream from the acquisition link included in the first application terminal hosting the application and obtains the proximal signal of the shared audio from the playback link of the first application terminal; performs echo cancellation processing on the proximal signal of the call stream according to the proximal signal of the shared audio to obtain the proximal target call stream audio signal; sends the proximal signal of the shared audio and the proximal target call stream audio signal to the second application terminal hosting the application for playback processing.
[0044] Based on the above embodiment, the embodiment of the present application adopts Figure 2 The flowchart details how to obtain the proximal signal of the shared audio according to the version of the Android system. Please refer to Figure 2 , Figure 2 is Figure 1 A schematic flowchart of an exemplary embodiment of step S110 in the audio processing method based on the Android system is shown. Specifically, the process of obtaining the proximal signal of the shared audio from the playback link of the first application terminal in step S110 specifically includes the following steps:
[0045] First of all, it should be noted that the playback link includes an audio framework layer, a hardware abstraction layer, and a digital signal processing layer. The audio framework layer, the hardware abstraction layer, and the digital signal processing layer refer to the framework levels for audio playback in the Android system.
[0046] S210: Based on the preset permissions of the Android system, obtain the shared audio near-end signal from the audio framework layer, the hardware abstraction layer, or the digital signal processing layer of the playback link.
[0047] The preset permission refers to the permission of the Android system to support the function of recording sound within the system. Exemplarily, the version of the Android system can to a certain extent determine the preset permissions of the Android system. Among them, the version of the Android system refers to the different versions released by the Android system of the developed mobile device at different times. Exemplarily, generally speaking, Android system versions lower than or equal to Android 9 do not support the function of recording sound within the system.
[0048] After detecting the version of the Android system, the audio processing device determines the preset permissions of the Android system. If the preset permissions of the Android system support recording sound, it obtains the shared audio near-end signal from the audio framework layer, the hardware abstraction layer, or the digital signal processing layer of the playback link.
[0049] It can be seen that the audio processing method based on the Android system in the embodiments of the present application obtains the shared audio near-end signal from the audio framework layer, the hardware abstraction layer, or the digital signal processing layer of the playback link based on the preset permissions of the Android system. This improves the adaptability between obtaining the shared audio near-end signal in the Android system and the preset permissions of the Android system.
[0050] Based on the above embodiments, the embodiments of the present application use Figure 3 The flowchart details how to obtain the shared audio near-end signal according to the permission information of the application. Please refer to Figure 3 , Figure 3 is Figure 2 The schematic flowchart of an exemplary embodiment of step S210 in the audio processing method based on the Android system shown. Specifically, the process of step S210 obtaining the shared audio near-end signal from the audio framework layer, the hardware abstraction layer, or the digital signal processing layer of the playback link based on the preset permissions of the Android system specifically includes the following steps:
[0051] S310: Detect the permission information of the application, where the permission information includes the permission for the Android system to allow the application to obtain data from the corresponding layer of the playback link.
[0052] Permission information refers to the information that the Android system authorizes an application to access the Android system and user data. Exemplarily, an application can access the Android system through an API interface. Among them, the API interface is a call interface provided by the Android system to the application, and the application makes the Android system execute the commands of the application by calling the API interface of the Android system.
[0053] In the embodiment of the present application, the permission information includes the permission for the Android system to allow the application to obtain data from the corresponding layer of the playback link, that is, the permission for the Android system to open for the application to obtain the shared audio near-end signal. Exemplarily, the permission of the digital signal processing layer in the playback link is greater than that of the hardware abstraction layer, and the permission of the hardware abstraction layer is greater than that of the audio framework layer.
[0054] The audio processing device detects the permission for the Android system to allow the application to obtain the shared audio near-end signal from the corresponding layer of the playback link.
[0055] S320: Match the target layer from the audio framework layer, the hardware abstraction layer, and the digital signal processing layer based on the permission information, and the target layer includes one of the audio framework layer, the hardware abstraction layer, and the digital signal processing layer.
[0056] The permission information determines the acquisition level for the application to obtain the shared audio near-end signal. The hierarchical permissions of the audio framework layer, the hardware abstraction layer, and the digital signal processing layer increase in turn.
[0057] The target layer refers to the level at which the application obtains the shared audio near-end signal. The levels for obtaining the shared audio near-end signal corresponding to applications with different permissions are different.
[0058] What is obtained through the above three levels is the mix of the shared audio signal and the remote target call flow audio signal, that is, the shared audio near-end signal. In order to obtain the remote target call flow audio signal separately, the remote target call flow audio signal can be obtained from the playback thread in the audio framework layer. Exemplarily, after obtaining the mix of the shared audio signal and the remote target call flow audio signal in the audio framework layer, the addMatchingUsage() function and the execludeUsage() function can be used to determine whether to obtain the audio signals included in the mix. For example, the addMatchingUsage() function can be used to determine to obtain any audio signal in the mix, and the execludeUsage() function can be used to determine not to obtain any audio signal in the mix. Exemplarily, the addMatchingUsage() function can be used to determine to obtain only the remote target call flow audio signal in the mix.
[0059] The audio processing device matches a target layer from the audio framework layer, the hardware abstraction layer, and the digital signal processing layer according to the permission information of the application. The target layer includes one of the audio framework layer, the hardware abstraction layer, and the digital signal processing layer.
[0060] S330: Obtain the shared audio near-end signal from the target layer.
[0061] If the Android system only opens to the audio framework layer for the application, the application can only obtain the shared audio near-end signal from the audio framework layer; if the Android system opens to the hardware abstraction layer for the application, the application can obtain the shared audio near-end signal from the audio framework layer or the hardware abstraction layer; if the Android system opens to the digital signal processing layer for the application, the application can obtain the shared audio near-end signal from the audio framework layer, the hardware abstraction layer, or the digital signal processing layer.
[0062] The audio processing device determines the acquisition level according to the degree of permission opening of the Android system to the application, and selects one of the opened levels to obtain the shared audio near-end signal.
[0063] To elaborate in detail the process of obtaining the shared audio near-end signal applied to the audio processing method based on the Android system in this application, the Figure 4 following flow chart is used for further illustration, details are as follows:
[0064] In the Android system, the playback path of the audio signal is from Application (application layer) to framework (audio framework layer) to HAL (hardware abstraction layer) to DSP (digital signal processing layer) and then to SPK (speaker). The acquisition link of the audio signal is from MIC (microphone) to DSP (digital signal processing layer) to HAL (hardware abstraction layer) to framework (audio framework layer) and then to Application (application layer). Application, that is, the application program, includes conferencing (conference software) and Music APK (other non-conference applications). The AudioTrack (audio signal) generated by conferencing and Music APK is mixed to obtain Mixer (mixing). If the shared audio near-end signal is obtained in the framework (audio framework layer), the mixing is processed in the PlaybackThreads (playback thread) to obtain the required shared audio near-end audio signal, and the audio signal is looped back to the conferencing (conference software); if the shared audio near-end signal is obtained in the HAL (hardware abstraction layer), the Input (shared audio near-end signal) is obtained from the Output (output data) in the HAL, and the Input is looped back to the conferencing (conference software); if the shared audio near-end signal is obtained in the DSP (digital signal processing layer), the EcRef (shared audio near-end signal) is obtained from the DSP, and the EcRef is looped back to the conferencing (conference software).
[0065] It can be seen that the audio processing method based on the Android system in the embodiment of the present application detects the permission information of the application program. The permission information includes the permission for the Android system to allow the application program to obtain data from the corresponding layer of the playback link; based on the permission information, the target layer is matched from the audio framework layer, the hardware abstraction layer, and the digital signal processing layer. The target layer includes one of the audio framework layer, the hardware abstraction layer, and the digital signal processing layer; the shared audio near-end signal is obtained from the target layer. Thus, three levels are provided to obtain the shared audio near-end signal, which can ensure that application programs with different permissions can obtain the shared audio near-end signal.
[0066] Based on the above embodiment, the embodiment of the present application uses Figure 5 The flowchart details how to adjust the delay value between the shared audio near-end signal and the call flow near-end signal. Please refer to Figure 5 , Figure 5 is Figure 1Schematic diagram of an exemplary embodiment before step S120 in the shown audio processing method based on the Android system. Specifically, before performing echo cancellation processing on the shared audio near-end signal according to the remote target call flow audio signal in step S120 to obtain the shared audio signal, the audio processing method based on the Android system in the embodiments of the present application specifically includes the following steps:
[0067] To distinguish the delay values in each embodiment, the delay value between the shared audio near-end signal and the call flow near-end signal in this embodiment is denoted as the first delay value, and the delay value between the shared audio near-end signal and the remote target call flow audio signal is denoted as the second delay value.
[0068] S410: Obtain the first delay value between the shared audio near-end signal and the call flow near-end signal. The shared audio near-end signal is obtained at the first timestamp in the current shared audio near-end queue, and the current shared audio near-end queue includes at least one timestamp for collecting the shared audio near-end signal.
[0069] The first delay value refers to the numerical value of the delay between the shared audio near-end signal and the call flow near-end signal. The positive or negative of the first delay value in this embodiment is not limited. The first delay value includes a positive delay value and a negative delay value. When the shared audio near-end signal is earlier than the call flow near-end signal, the first delay value between the two can be a positive delay value. When the call flow near-end signal is earlier than the shared audio near-end signal, the first delay value between the two can be a negative delay value. This embodiment does not limit how to determine the first delay value, as long as the first delay value can be obtained.
[0070] The current shared audio near-end queue refers to the queue storing the shared audio near-end signal. The current shared audio near-end queue includes at least one timestamp for collecting the shared audio near-end signal. Exemplarily, at the current moment, the collected shared audio near-end signal is cached into the shared audio near-end queue to form the current shared audio near-end queue.
[0071] The timestamp refers to the absolute time point when the audio is converted by the DAC (Digital-to-Analog Converter) or ADC (Analog-to-Digital Converter). Exemplarily, the timestamp of the shared audio near-end signal is the absolute time point when it is converted by the DAC; the timestamp of the call flow near-end signal is the absolute time point when it is converted by the ADC.
[0072] The audio processing device obtains the first delay value between the shared audio near-end signal and the call flow near-end signal. The shared audio near-end signal is obtained at the first timestamp in the current shared audio near-end queue, and the current shared audio near-end queue stores the shared audio near-end signal and includes at least one timestamp for collecting the shared audio near-end signal.
[0073] S420: In response to the absolute value of the first delay value being greater than the first set delay value, adjust the timestamps of other timestamps in the current shared audio near-end queue according to the timestamp offset between the first delay value and the first set delay value, where the sorting time of the first timestamp in the current shared audio near-end queue is earlier than the sorting time of other timestamps in the current shared audio near-end queue.
[0074] First, compare the first delay value between the obtained shared audio near-end signal and the call flow near-end signal with the first set delay value. If the absolute value of the first delay value is greater than the first set delay value, perform timestamp adjustment through the timestamp offset between the first delay value and the first set delay value; otherwise, do not process the first delay value.
[0075] The first set delay value refers to the delay value that is preset to allow between the shared audio near-end signal and the call flow near-end signal.
[0076] Determine the calculation method of the timestamp offset between the first delay value and the first set delay value according to the positive or negative of the first delay value. Exemplarily, when the first delay value is positive, the difference between the first delay value and the first set delay value is used as the timestamp offset between the first delay value and the first set delay value; when the first delay value is negative, the sum of the absolute value of the first delay value and the first set delay value is used as the timestamp offset between the first delay value and the first set delay value.
[0077] Exemplarily, the timestamp of the shared audio near-end signal can be represented as T 0 , the timestamp of the call flow near-end signal can be represented as T 1 , the first delay value can be represented as ΔT 1 , the first set delay value can be represented as ΔT 0 , and the timestamp offset between the first delay value and the first set delay value can be represented as ΔT.
[0078] When the timestamp of the shared audio near-end signal is earlier than the timestamp of the call flow near-end signal, the first delay value is positive, and the calculation of the first delay value satisfies the following formula: ΔT 1 = T 1 - T 0 ; the calculation of the timestamp offset between the first delay value and the first set delay value satisfies the following formula: ΔT = ΔT 1 - ΔT 0 .
[0079] When the timestamp of the call flow near-end signal is earlier than the timestamp of the shared audio near-end signal, the first delay value is negative, and the calculation of the first delay value satisfies the following formula: ΔT 1 = T 1 - T 0; The calculation of the timestamp offset between the first delay value and the first set delay value satisfies the following formula: ΔT = |ΔT 1 | + ΔT 0 .
[0080] Other timestamps refer to the timestamps corresponding to other shared audio proximal signals in the current shared audio proximal queue except the first timestamp.
[0081] The sorting time refers to the reading time for echo cancellation by using the shared audio proximal signals in the current shared audio proximal queue as the reference signals for the proximal signals of the call flow. When reading the shared audio proximal signals from the current shared audio proximal queue, the signals can be popped out of the stack in sequence based on the sorting times of the respective shared audio proximal signals in the current shared audio proximal queue.
[0082] Adjust the other timestamps in the current shared audio proximal queue according to the timestamp offset between the first delay value and the first set delay value, so as to ensure that the shared audio proximal signals corresponding to the other timestamps are read from the current shared audio proximal queue earlier than the corresponding proximal signals of the call flow from the call flow proximal queue, and the difference between the timestamps of reading the shared audio proximal signals and the timestamps of reading the proximal signals of the call flow is less than or equal to the first set delay value. Wherein, the call flow proximal queue stores the proximal signals of the call flow.
[0083] The audio processing device determines whether the absolute value of the first delay value between the acquired shared audio proximal signal and the proximal signal of the call flow is greater than the first set delay value. If it is greater, adjust the other timestamps in the current shared audio proximal queue according to the first delay value, so as to ensure that the timestamp of reading the shared audio proximal signal is earlier than the timestamp of reading the corresponding proximal signal of the call flow, and at the same time, the difference between the timestamps of reading the shared audio proximal signal and the timestamps of reading the proximal signal of the call flow is less than or equal to the first set delay value.
[0084] S430: Use the first timestamp in the current shared audio proximal queue and the adjusted other timestamps as the timestamps in the latest shared audio proximal queue.
[0085] The audio processing device uses the first timestamp in the current shared audio proximal queue and the adjusted other timestamps as the timestamps in the latest shared audio proximal queue, and reads the shared audio proximal signals according to the timestamps in the latest shared audio proximal queue.
[0086] It can be seen that the audio processing method based on the Android system in the embodiment of the present application obtains the first delay value between the shared audio near-end signal and the call flow near-end signal. The shared audio near-end signal is obtained at the first timestamp in the current shared audio near-end queue, and the current shared audio near-end queue includes at least one timestamp for collecting the shared audio near-end signal. In response to the absolute value of the first delay value being greater than the first set delay value, adjust the other timestamps in the current shared audio near-end queue according to the timestamp offset between the first delay value and the first set delay value. The sorting time of the first timestamp in the current shared audio near-end queue is earlier than the sorting time of the other timestamps in the current shared audio near-end queue. Take the first timestamp in the current shared audio near-end queue and the adjusted other timestamps as the timestamps in the latest shared audio near-end queue. Thereby, the delay value between the shared audio near-end signal and the call flow near-end signal can be adjusted to ensure that the delay of the shared audio near-end signal and the call flow near-end signal of the remote target call is within the allowable range, so as to ensure the normal operation of echo cancellation.
[0087] Based on the above embodiments, the embodiments of the present application adopt Figure 6 The flowchart details how to adjust the other timestamps in the current shared audio near-end queue. Please refer to Figure 6 , Figure 6 is Figure 5 The flowchart shows an exemplary embodiment of step S420 in the audio processing method based on the Android system. Specifically, in step S420, the process of adjusting the other timestamps in the current shared audio near-end queue according to the timestamp offset between the first delay value and the first set delay value specifically includes the following steps:
[0088] First, it should be noted that the other timestamps include the second timestamp and the third timestamp arranged in ascending order of sorting time.
[0089] The sorting time of the second timestamp is earlier than that of the third timestamp.
[0090] S510: Take the sum of the timestamp offset and the second timestamp as the adjusted second timestamp.
[0091] Exemplarily, the timestamp offset calculated at time point A is ΔT, and the second timestamp is represented as T 2 , and the adjusted second timestamp can be represented as T NEW2 = T 2 -ΔT; at time point B, due to other system reasons, the timestamp offset required for delay adjustment is ΔT′, then the adjusted second timestamp at time point B can be represented as T N ′ EW2 = T NEW2 -ΔT′.
[0092] The audio processing device uses the sum of the timestamp offset at the current moment and the second timestamp as the adjusted second timestamp.
[0093] S520: Use the sum of the timestamp offset and the third timestamp as the adjusted third timestamp.
[0094] The calculation method of the third timestamp is the same as that of the second timestamp in step S510. Therefore, the calculation method of step S510 can be referred to for calculating the third timestamp, and it will not be elaborated here.
[0095] The audio processing device uses the sum of the timestamp offset at the current moment and the third timestamp as the adjusted third timestamp.
[0096] It can be seen that in the audio processing method based on the Android system in the embodiments of the present application, the sum of the timestamp offset and the second timestamp is used as the adjusted second timestamp; the sum of the timestamp offset and the third timestamp is used as the adjusted third timestamp. Other timestamps are adjusted accordingly to ensure that the delay between the shared audio near-end signal and the call flow near-end signal is within the allowable range, thereby ensuring the normal operation of echo cancellation.
[0097] Based on the above embodiments, the embodiments of the present application adopt Figure 7 The flowchart details how to obtain the shared audio signal. It should be noted that this embodiment can be combined with the above embodiments. Please refer to Figure 7 , Figure 7 is Figure 1 a schematic flowchart of an exemplary embodiment after step S110 in the audio processing method based on the Android system shown. Specifically, it includes the following steps:
[0098] S610: Obtain the remote target call flow audio signal in the shared audio near-end signal, and the remote target call flow audio signal comes from the second application terminal.
[0099] During a conference call, if an external sound signal is collected from the second application terminal, the collected external sound signal will be played and processed through the first application terminal. Therefore, when the second application terminal does not collect an external sound signal, the shared audio near-end signal collected from the playback link of the first application terminal only contains the shared audio signal; when the second application terminal collects an external sound signal, the shared audio near-end signal collected from the playback link of the first application terminal contains the shared audio signal and the remote target call flow audio signal.
[0100] The remote target call flow audio signal refers to the audio signal from the second application terminal. Exemplarily, taking the application program in the embodiments of this application as a conferencing software, the remote target call flow audio signal includes the voice of the user of the second application terminal speaking in the current conference.
[0101] The audio processing device obtains the audio signal from the second application terminal from the playback link of the first application terminal.
[0102] S620: Perform echo cancellation processing on the shared audio proximal signal according to the remote target call flow audio signal to obtain the shared audio signal.
[0103] The shared audio signal refers to the third-party audio signal played by the first application terminal through a non-conference application. Exemplarily, the shared audio signal can be the audio played by other audio applications such as a music application, a video application, etc. Currently, when the first application terminal conducts a conference call, it can share the audio on this terminal to other application terminals participating in the conference. At this time, the audio signal shared to other application terminals in the conference is the shared audio signal.
[0104] The remote target call flow audio signal comes from the second application terminal. In order to prevent the second application terminal from receiving the remote target call flow audio signal again and thus generating an echo, it is necessary to perform echo cancellation processing on the remote target call flow audio signal in the shared audio proximal signal so that the second application terminal only receives the shared audio signal.
[0105] The audio processing device performs echo cancellation processing on the shared audio proximal signal according to the remote target call flow audio signal from the second application terminal to obtain the shared audio signal.
[0106] In one embodiment, before performing echo cancellation processing on the shared audio proximal signal according to the remote target call flow audio signal to obtain the shared audio signal, it further includes estimating the delay of the remote target call flow audio signal with respect to the shared audio proximal signal.
[0107] Specifically, obtain the second delay value between the remote target call flow audio signal and the shared audio proximal signal. The remote target call flow audio signal is obtained at the fourth timestamp in the current remote call flow audio queue, and the current remote call flow audio queue includes at least one timestamp for collecting the remote call flow audio signal.
[0108] Among them, the second delay value refers to the value of the delay between the audio signal of the remote target call flow and the shared audio proximal signal. In this embodiment, the positive or negative of the second delay value is not limited. The second delay value includes a positive delay value and a negative delay value. When the audio signal of the remote target call flow is earlier than the shared audio proximal signal, the second delay value between the two can be a positive delay value. When the shared audio proximal signal is earlier than the audio signal of the remote target call flow, the second delay value between the two can be a negative delay value. This embodiment does not limit how to determine the first delay value, as long as the first delay value can be obtained. The current remote call flow audio queue refers to the queue storing the audio signals of the remote call flow. The current remote call flow audio queue includes at least one timestamp for collecting the audio signals of the remote call flow. Exemplarily, at the current moment, the collected audio signals of the remote call flow are cached into the remote call flow audio queue to form the current remote call flow audio queue.
[0109] In response to the absolute value of the second delay value being greater than the second set delay value, adjust other timestamps in the current remote call flow audio queue according to the timestamp offset between the second delay value and the second set delay value. The sorting time of the fourth timestamp in the current remote call flow audio queue is earlier than the sorting time of other timestamps in the current remote call flow audio queue.
[0110] The second set delay value refers to the delay value that is preset and allowed to exist between the audio signal of the remote target call flow and the shared audio proximal signal.
[0111] Other timestamps refer to the timestamps corresponding to other current remote target call flow audio signals in the current remote call flow audio queue except the fourth timestamp.
[0112] The sorting time refers to the reading time for echo cancellation by using the audio signal of the remote target call flow in the current remote call flow audio queue as the reference signal for the shared audio proximal signal. When reading the audio signal of the remote target call flow from the current remote call flow audio queue, it can be popped out of the stack in sequence based on the sorting time of each audio signal of the remote target call flow in the current remote call flow audio queue.
[0113] In this embodiment, the steps of calculating the timestamp offset between the second delay value and the second set delay value and adjusting other timestamps in the current remote call flow audio queue by the timestamp offset between the second delay value and the second set delay value can be specifically referred to step S420, and will not be elaborated in detail in this embodiment.
[0114] Take the fourth timestamp in the current remote call flow audio queue and the adjusted other timestamps as the timestamps in the latest remote call flow audio queue.
[0115] The audio processing device uses the first timestamp in the current remote call stream audio queue and the adjusted other timestamps as the timestamps in the latest remote call stream audio queue, and reads the remote target call stream audio signal according to the timestamps in the latest remote call stream audio queue.
[0116] S630: Send the shared audio signal and the proximal target call stream audio signal to the second application terminal hosting the application for playback processing.
[0117] After obtaining the shared audio signal and the proximal target call stream audio signal, the application encodes the shared audio signal to obtain a shared audio encoding, encodes the proximal target call stream audio signal to obtain a call stream audio encoding, and transmits the shared audio encoding and the call stream audio encoding to the second application terminal through the server, and the second application terminal decodes and plays them.
[0118] It can be seen that the audio processing method based on the Android system in the embodiments of the present application obtains the remote target call stream audio signal in the shared audio proximal signal, and the remote target call stream audio signal comes from the second application terminal; performs echo cancellation processing on the shared audio proximal signal according to the remote target call stream audio signal to obtain a shared audio signal; performs echo cancellation processing on the call stream proximal signal according to the shared audio proximal signal to obtain a proximal target call stream audio signal; and plays the shared audio signal and the proximal target call stream audio signal through the second application terminal. By performing echo cancellation processing on the shared audio proximal signal and the call stream proximal signal, relatively pure shared audio signals and proximal target call stream audio signals can be obtained, reducing noise and improving call quality.
[0119] Based on the above embodiments, the embodiments of the present application use Figure 8 The flowchart details how to perform echo cancellation processing on the call stream proximal signal. Please refer to Figure 8 , Figure 8 is Figure 1 A schematic flowchart of an exemplary embodiment of step S120 in the audio processing method based on the Android system shown. Specifically, in step S120, the process of performing echo cancellation processing on the call stream proximal signal according to the shared audio proximal signal to obtain a proximal target call stream audio signal specifically includes the following steps:
[0120] S710: Perform echo path simulation on the shared audio proximal signal to obtain a shared audio proximal estimated echo signal.
[0121] The echo path simulation can use the methods commonly used in the art. Exemplarily, the echo path simulation can be performed through an adaptive estimation method (such as a block frequency domain adaptive filter algorithm) to obtain a shared audio proximal estimated echo signal.
[0122] The audio playback device uses an echo cancellation algorithm commonly used by those skilled in the art to obtain an analog echo path, simulates the echo path for the shared audio near-end signal, and obtains an estimated echo signal for the shared audio near-end.
[0123] S720: Based on the estimated echo signal for the shared audio near-end, perform cancellation processing on the near-end signal of the call flow to obtain the near-end target call flow audio signal.
[0124] The audio playback device obtains an estimated echo signal for the shared audio near-end through echo path simulation, and performs cancellation processing on the far estimated echo signal for the shared audio near-end in the near-end signal of the call flow according to the estimated echo signal for the shared audio near-end, so as to obtain only the near-end target call flow audio signal.
[0125] It can be seen that the audio processing method based on the Android system in the embodiments of the present application performs echo path simulation on the shared audio near-end signal to obtain an estimated echo signal for the shared audio near-end; based on the estimated echo signal for the shared audio near-end, perform cancellation processing on the near-end signal of the call flow to obtain the near-end target call flow audio signal. Thus, it is possible to prevent the second application terminal from receiving the shared audio near-end signal and generating echoes, thereby improving the call quality.
[0126] To elaborate in detail on the application of the present application to the audio processing method based on the Android system, Figure 9 the following further describes it with the flowchart shown below:
[0127] The first application terminal obtains the far-end target call flow audio signal and the shared audio near-end signal from the second application terminal in the playback link. Among them, the shared audio near-end signal contains the shared audio signal played by other non-conference applications and the far-end target call flow audio signal; the first application terminal collects the near-end signal of the call flow in the collection link. Among them, the near-end signal of the call flow contains the shared audio near-end signal and the near-end target call flow audio signal. Perform echo cancellation processing on the shared audio near-end signal according to the far-end target call flow audio signal to obtain the shared audio signal, and perform echo cancellation processing on the near-end signal of the call flow according to the shared audio near-end signal to obtain the near-end target call flow audio signal. Encode the shared audio signal to obtain the shared audio encoding, encode the near-end target call flow audio signal to obtain the call flow encoding, and send the shared audio encoding and the call flow encoding to the second application terminal for decoding and then playback processing.
[0128] Please refer to Figure 10 , Figure 10It is a schematic flowchart of an exemplary embodiment of the audio processing device provided by this application. The audio processing device 10 includes an acquisition module 101, an echo processing module 102, and a playback module 103. The acquisition module 101 is configured to acquire a near-end signal of a call stream from an acquisition link included in a first application terminal carrying an application program and acquire a shared audio near-end signal from a playback link of the first application terminal; the echo processing module 102 is configured to perform echo cancellation processing on the near-end signal of the call stream according to the shared audio near-end signal to obtain a near-end target call stream audio signal; the playback module 103 is configured to send the shared audio near-end signal and the near-end target call stream audio signal to a second application terminal carrying the application program for playback processing.
[0129] In the above solution, the audio processing device according to the embodiment of this application acquires a near-end signal of a call stream from an acquisition link included in a first application terminal carrying an application program and acquires a shared audio near-end signal from a playback link of the first application terminal; performs echo cancellation processing on the near-end signal of the call stream according to the shared audio near-end signal to obtain a near-end target call stream audio signal; and sends the shared audio signal and the near-end target call stream audio signal to a second application terminal carrying the application program for playback processing. By performing echo cancellation processing on the near-end signal of the call stream, it is possible to play the shared audio through its own speaker while sharing the audio.
[0130] Among them, the functions of each module can be referred to in the embodiment of the audio processing method based on the Android system, which will not be elaborated here.
[0131] To implement the audio processing method based on the Android system in the above embodiment, this application proposes another electronic device. For details, please refer to Figure 11 , Figure 11 It is a schematic structural diagram of an embodiment of the electronic device provided by this application.
[0132] The electronic device 11 includes a memory 1101 and a processor 1102. Among them, the memory 1101 and the processor 1102 are coupled.
[0133] The memory 1101 is used to store program data, and the processor 1102 is used to execute the program data to implement the audio processing method based on the Android system in the above embodiment.
[0134] In this embodiment, the processor 1102 can also be referred to as a CPU (Central Processing Unit). The processor 1102 may be an integrated circuit chip with signal processing capabilities. The processor 1102 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor 1102 can also be any conventional processor, etc.
[0135] This application also provides a computer-readable storage medium 12, as Figure 12 shown. The computer-readable storage medium 12 is used to store program data 1201. When the program data 1201 is executed by the processor, it is used to implement the audio processing method based on the Android system in the method embodiment of this application.
[0136] Regarding the method involved in the method embodiment of the audio processing method based on the Android system in this application, when it exists in the form of a software functional unit and is sold or used as an independent product during implementation, it can be stored in a device, such as a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, etc., which can store program codes.
[0137] The above are only the embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or equivalent process transformation made using the content of the specification and drawings of this application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of this application.
Claims
1. An audio processing method based on the Android system, characterized in that, the method is applied to an application, and the method includes: collecting a call flow proximal signal from a collection link included in a first application terminal carrying the application and obtaining a shared audio proximal signal from a playback link of the first application terminal; performing echo cancellation processing on the call flow proximal signal according to the shared audio proximal signal to obtain a proximal target call flow audio signal; sending the shared audio proximal signal and the proximal target call flow audio signal to a second application terminal carrying the application for playback processing.
2. The audio processing method based on the Android system according to claim 1, characterized in that, the playback link includes an audio framework layer, a hardware abstraction layer, and a digital signal processing layer, and the step of obtaining the shared audio proximal signal from the playback link of the first application terminal includes: obtaining the shared audio proximal signal from the audio framework layer, the hardware abstraction layer, or the digital signal processing layer of the playback link based on a preset permission of the Android system.
3. The audio processing method based on the Android system according to claim 2, characterized in that, the step of obtaining the shared audio proximal signal from the audio framework layer, the hardware abstraction layer, or the digital signal processing layer of the playback link based on the preset permission of the Android system further includes: detecting permission information of the application, where the permission information includes a permission for the Android system to allow the application to obtain data from a corresponding layer of the playback link; matching a target layer from the audio framework layer, the hardware abstraction layer, and the digital signal processing layer based on the permission information, where the target layer includes one of the audio framework layer, the hardware abstraction layer, and the digital signal processing layer; obtaining the shared audio proximal signal from the target layer.
4. The audio processing method based on the Android system according to claim 1, characterized in that, before the step of performing echo cancellation processing on the call flow proximal signal according to the shared audio proximal signal to obtain a proximal target call flow audio signal, the method further includes: obtaining a delay value between the shared audio proximal signal and the call flow proximal signal, where the shared audio proximal signal is obtained at a first timestamp in a current shared audio proximal queue, and the current shared audio proximal queue includes at least one timestamp for collecting the shared audio proximal signal; in response to the absolute value of the delay value being greater than a set delay value, adjusting other timestamps in the current shared audio proximal queue according to a timestamp offset between the delay value and the set delay value, where the first timestamp is earlier than the other timestamps in the current shared audio proximal queue in terms of sorting time; using the first timestamp in the current shared audio proximal queue and the adjusted other timestamps as timestamps in a latest shared audio proximal queue.
5. The audio processing method based on the Android system according to claim 4, wherein, the other timestamps include a second timestamp and a third timestamp arranged in ascending order of sorting time, and the step of adjusting the other timestamps in the current shared audio proximal queue according to the timestamp offset between the delay value and the set delay value includes: taking the sum of the timestamp offset and the second timestamp as the adjusted second timestamp; taking the sum of the timestamp offset and the third timestamp as the adjusted third timestamp.
6. The audio processing method based on the Android system according to claim 1, wherein, the method further includes: acquiring a remote target call flow audio signal in the shared audio proximal signal, and the remote target call flow audio signal comes from the second application terminal; performing echo cancellation processing on the shared audio proximal signal according to the remote target call flow audio signal to obtain a shared audio signal; the step of sending the shared audio proximal signal and the proximal target call flow audio signal to the second application terminal carrying the application program for playback processing includes: sending the shared audio signal and the proximal target call flow audio signal to the second application terminal carrying the application program for playback processing.
7. The audio processing method based on the Android system according to claim 1, wherein, the step of performing echo cancellation processing on the call flow proximal signal according to the shared audio proximal signal to obtain a proximal target call flow audio signal includes: performing echo path simulation on the shared audio proximal signal to obtain a shared audio proximal estimated echo signal; performing cancellation processing on the call flow proximal signal based on the shared audio proximal estimated echo signal to obtain the proximal target call flow audio signal.
8. An audio processing device, wherein, the audio processing device includes: an acquisition module, configured to acquire a call flow proximal signal from an acquisition link included in a first application terminal carrying the application program and obtain a shared audio proximal signal from a playback link of the first application terminal; a first echo processing module, configured to perform echo cancellation processing on the shared audio proximal signal according to the remote target call flow audio signal to obtain a shared audio signal; a second echo processing module, configured to perform echo cancellation processing on the call flow proximal signal according to the shared audio proximal signal to obtain a proximal target call flow audio signal; a playback module, configured to perform playback processing on the shared audio signal and the proximal target call flow audio signal through the second application terminal.
9. An electronic device, wherein, includes: a memory and a processor, wherein the memory stores program instructions, and the processor retrieves the program instructions from the memory to execute the audio processing method based on the Android system according to any one of claims 1-7.
10. A computer-readable storage medium, wherein, includes: Stored with program data, when the program data is executed by a processor, it is used to implement the audio processing method based on the Android system described in any one of claims 1-7.