Method and device for recording application audio for mobile terminal

By directly recording the playback and audio data of VoIP applications on mobile terminals, generating and displaying audio files, the problems of cumbersome VoIP call recording and privacy leaks are solved, recording speed and user experience are improved, and power consumption is reduced.

CN121173902APending Publication Date: 2025-12-19SAMSUNG GUANGZHOU MOBILE R&D CENT +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410796228.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-06-19
Publication Date
2025-12-19

AI Technical Summary

Technical Problem

The existing VoIP call voice recording process is cumbersome, has a poor user experience, and suffers from privacy leaks and high power consumption.

Method used

A method and apparatus for recording application audio on a mobile terminal are provided, which directly records the playback data and recording data of VoIP applications, generates an audio file, and displays it in a multimedia application, avoiding the need to use other screen recording tools.

Benefits of technology

It improves recording speed and user experience, prevents privacy leaks, and reduces power consumption during the recording process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121173902A_ABST
    Figure CN121173902A_ABST
Patent Text Reader

Abstract

The invention discloses an application audio recording method and device for a mobile terminal. The method comprises the steps of determining at least one preset application supporting an application audio recording function in response to the fact that the application audio recording function is started; recording playing data and recording data of the preset application in response to the condition that the preset application is in the call state; generating a recording file based on the recorded playing data and recording data; and displaying the sound recording file in a file list of a multimedia application.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates generally to the field of electronic technology, and more particularly, to a method and apparatus for recording application audio for a mobile terminal. BACKGROUND

[0002] With the development and popularity of various intelligent terminals, voice and video calls (i.e., VoIP calls, such as WeChat voice / video calls, QQ voice / video calls, Tencent meetings, etc.) as a means of remote communication and collaboration, the application scenarios in modern society are expanding. VoIP calls enable real-time audio and video transmission, allowing people to communicate and collaborate face-to-face at different locations.

[0003] VoIP calls can achieve efficient remote meetings / office work and communication by means of the Internet and mobile devices. Compared with traditional communication methods, VoIP calls are not limited by time and space, and can effectively improve user communication efficiency and decision-making efficiency. For important VoIP calls, users may want to record the voice of the VoIP call to record and further organize the call content (e.g., transcription, translation, abstract extraction, etc.). However, the process of recording the voice of the VoIP call in the related art is very cumbersome, and the user experience is not good. SUMMARY

[0004] Embodiments of the present disclosure provide a method and apparatus for recording application audio for a mobile terminal, which can quickly and accurately record audio data of a preset application.

[0005] According to a first aspect of embodiments of the present disclosure, a method for recording application audio for a mobile terminal is provided, the method comprising: in response to a recording application audio function being turned on, determining at least one preset application that supports the recording application audio function; in response to the preset application being in a call state, recording playing data and recording data of the preset application; generating a recording file based on the recorded playing data and recording data; and displaying the recording file in a file list of a multimedia application.

[0006] Optionally, the method further comprises: in response to the recording application audio function being turned on, registering a listening audio state event to an operating system of the mobile terminal to listen to audio states of each application of the mobile terminal; and in response to listening to a change in the audio state of the preset application, determining that the preset application is in the call state.

[0007] Optionally, in response to the preset application being in the call state, the step of recording the playing data and the recording data of the preset application comprises: in response to the preset application being in the call state, obtaining the playing data and the recording data of the preset application; and recording the playing data and the recording data of the preset application, respectively.

[0008] Optionally, the step of obtaining the playing data and the recording data of the preset application comprises: receiving the preset application from a remote device and backing up audio data played through a speaker to obtain the playing data of the preset application; and / or bypassing audio data collected through a microphone to obtain the recording data of the preset application.

[0009] Optionally, the step of bypassing the audio data input from the microphone comprises: in response to the preset application being in a state of allowing use of the microphone, bypassing the audio data collected through the microphone; and in response to the preset application being in a state of not allowing use of the microphone, prohibiting bypassing the audio data input through the microphone.

[0010] Optionally, the step of generating the recording file based on the recorded playing data and the recording data comprises: mixing and encoding the recorded playing data and the recording data in real time to generate the recording file.

[0011] Optionally, the step of generating the recording file based on the recorded playing data and the recording data further comprises: mixing and encoding the recorded playing data and the recording data to obtain a recorded audio file; and preprocessing the recorded audio file to generate an audio file recognizable by the multimedia playing application as the recording file.

[0012] According to a second aspect of embodiments of the present disclosure, there is provided an apparatus for recording application audio for a mobile terminal, the apparatus comprising: an application determining module configured to determine at least one preset application supporting a recording application audio function in response to the recording application audio function being turned on; a recording module configured to record playing data and recording data of the preset application in response to the preset application being in a call state; a recording file generating module configured to generate a recording file based on the recorded playing data and the recording data; and a display control module configured to display the recording file in a file list of a multimedia application.

[0013] Optionally, the apparatus further comprises: a registration and listening module configured to register a listening audio state event to an operating system of the mobile terminal to listen to audio states of each application of the mobile terminal in response to the recording application audio function being turned on; and a call state determining module configured to determine that the preset application is in the call state in response to listening to a change in the audio state of the preset application.

[0014] Optionally, the recording module is configured to obtain the playing data and the recording data of the preset application in response to the preset application being in the call state; and record the playing data and the recording data of the preset application, respectively.

[0015] Optionally, the recording module is further configured to: back up the audio data received by the preset application from a remote device and played through a speaker to obtain the playback data of the preset application; and / or perform bypass extraction on the audio data collected through a microphone to obtain the recording data of the preset application.

[0016] Optionally, the recording module is further configured to: bypass the extraction of audio data acquired through the microphone in response to the preset application being in a state where the microphone is allowed; and prohibit the bypass extraction of audio data input through the microphone in response to the preset application being in a state where the microphone is not allowed.

[0017] Optionally, the audio recording file generation module is configured to mix and encode the recorded playback data and audio recording data in real time to generate an audio recording file.

[0018] Optionally, the audio recording file generation module is further configured to: mix and encode the recorded playback data and audio recording data in real time to obtain a recorded audio file; and preprocess the recorded audio file to generate an audio file that the multimedia playback application can recognize, as the audio recording file.

[0019] According to a third aspect of the embodiments of the present disclosure, a computer-readable storage medium storing a computer program is provided, which, when executed by a processor, implements the method for recording application audio for a mobile terminal as described above.

[0020] According to a fourth aspect of the embodiments of the present disclosure, an electronic device is provided, the electronic device comprising: at least one processor; at least one memory storing a computer program that, when executed by the at least one processor, implements the method for recording application audio for a mobile terminal as described above.

[0021] The method and apparatus for recording application audio on a mobile terminal according to embodiments of this disclosure can directly record application audio data without using other screen recording tools, thereby avoiding cumbersome recording operations and post-editing work, improving recording speed and user experience. Furthermore, since the application audio data is recorded directly without using other screen recording tools, interference from other applications can be avoided, preventing the leakage of privacy information, and reducing screen recording power consumption.

[0022] Further aspects and / or advantages of the general concept of this disclosure will be set forth in part in the description which follows, and in part will be clear from the description or may be learned by practice of the general concept of this disclosure. Attached Figure Description

[0023] The above and other objects and features of exemplary embodiments of this disclosure will become clearer from the following description taken in conjunction with the accompanying drawings, which exemplarily illustrate the embodiments.

[0024] Figure 1 This is a flowchart illustrating a method for recording application audio for a mobile terminal according to an embodiment of the present disclosure.

[0025] Figure 2 This is an illustration showing an example of enabling the audio recording function of an application.

[0026] Figure 3 This is a diagram illustrating examples of one or more applications that support recording application audio functionality.

[0027] Figure 4 This is a diagram illustrating the existing audio processing architecture for VoIP calls.

[0028] Figure 5A This is a diagram showing an example of preparing to record audio in the normal mode (default mode) of the recording application's audio function.

[0029] Figure 5B This is an illustration showing an example of recording in normal mode.

[0030] Figure 5C This is a diagram showing an example of pausing recording in normal mode.

[0031] Figure 5D This is a diagram showing an example of preparing to record in mini mode.

[0032] Figure 5E This is an illustration showing an example of recording in mini mode.

[0033] Figure 5F This is a diagram showing an example of preparing to record in Plus mode.

[0034] Figure 5G This is an illustration showing an example of recording in Plus mode.

[0035] Figure 5H This is an illustration showing an example of pausing recording in Plus mode.

[0036] Figure 6 This is a diagram illustrating the processing structure for acquiring playback and recording data.

[0037] Figure 7 This is a block diagram illustrating an apparatus for recording application audio for a mobile terminal according to an embodiment of the present disclosure.

[0038] Figure 8 This is a block diagram illustrating an electronic device according to an embodiment of the present disclosure. Detailed Implementation

[0039] In the following description, various embodiments of the present disclosure are illustrated with reference to the accompanying drawings, wherein the same reference numerals are used to denote the same or similar elements, features, and structures. However, the present disclosure is not intended to be limited to the specific embodiments described herein, and it is intended that the present disclosure cover all modifications, equivalents, and / or substitutions of the present disclosure, provided they fall within the scope of the appended claims and their equivalents. The terms and words used in the following description and claims are not limited to their dictionary meanings, but are used only to enable a clear and consistent understanding of the present disclosure. Therefore, it will be apparent to those skilled in the art that the following description of various embodiments of the present disclosure is for illustrative purposes only and is not intended to limit the purpose of the present disclosure as defined by the appended claims and their equivalents.

[0040] It should be understood that, unless the context clearly indicates otherwise, the singular form includes the plural form. The terms “comprising,” “including,” and “having” as used herein indicate the presence of a disclosed function, operation, or element, but do not exclude other functions, operations, or elements.

[0041] For example, the expression “A or B” or “at least one of A and / or B” can indicate A and B, or A or B. For example, the expression “A or B” or “at least one of A and / or B” can indicate (1) A, (2) B or (3) both A and B.

[0042] In various embodiments of this disclosure, it is intended that when a component (e.g., a first component) is referred to as being "coupled" or "connected" to, or being "coupled" or "connected" to, another component (e.g., a second component), the component may be directly connected to, or may be connected via, another component (e.g., a third component). Conversely, when a component (e.g., a first component) is referred to as being "directly coupled" or "directly connected" to, or being directly coupled to or directly connected to, another component (e.g., a second component), there is no other component (e.g., a third component) between the component and the other component.

[0043] The expression “configured as” used in describing the various embodiments of this disclosure may be used interchangeably, for example, with expressions such as “suitable for,” “capable of,” “designed to,” “suitable for,” “manufactured as,” and “capable,” depending on the context. The term “configured as” may not necessarily indicate that the hardware is “specifically designed for.” Rather, in some cases, the expression “a device configured as…” may indicate that the device and another device or part are “capable of…”. For example, the expression “a processor configured to perform A, B, and C” may indicate a dedicated processor (e.g., an embedded processor) for performing the respective operations or a general-purpose processor (e.g., a central processing unit, CPU, or application processor (AP)) for performing the respective operations by executing at least one software program stored in memory.

[0044] The terminology used herein is intended to describe certain embodiments of this disclosure but is not intended to limit the scope of other embodiments. Unless otherwise indicated herein, all terms used herein (including technical or scientific terms) are to have the same meaning as commonly understood by one of ordinary skill in the art. Generally, terms as defined in dictionaries should be considered to have the same meaning as in the context of the relevant art and should not be interpreted differently or as having an overly formal meaning unless expressly defined herein. In no event should the terminology defined in this disclosure be construed as excluding embodiments of this disclosure.

[0045] Below, we will first describe the existing methods for recording application audio and their shortcomings.

[0046] Suppose user A uses a VoIP app (e.g., WeChat) to conduct VoIP calls with users B and C, who are located in different offices, to discuss project progress and wants to record the VoIP call to avoid forgetting important information during later reports. Because VoIP calls have high priority, users cannot directly record VoIP call audio through multimedia applications (e.g., but not limited to voice recorders). Specifically, multimedia applications cannot directly access the audio output of the VoIP app; and for the audio input of the VoIP app, when the highest-priority application (i.e., the VoIP app) uses its microphone (e.g., a local microphone or an external microphone), the multimedia application cannot use its microphone.

[0047] To record VoIP call audio, User A can only perform the following operations: (1) Start a screen recording tool (note that the screen recording tool needs to be set to "record media and microphone sound") to record video; (2) Extract audio data from the recorded video file; (3) Remove interference sounds from other applications besides the VoIP APP (e.g., notification sounds); (4) Convert the audio data after removing interference sounds into an audio file by performing special processing (e.g., inserting special fields); (5) Optionally, use various existing AI technologies to further process the audio file such as transcription, translation, and extraction of summaries.

[0048] Therefore, it is evident that there is no direct method for recording VoIP call audio in existing technologies, making it easy to miss or forget important information. On the other hand, recording VoIP call audio is cumbersome and requires a lot of post-editing work to obtain the recording file; at the same time, recording tools cannot record the audio of the app making the VoIP call separately, and the sounds of other apps will also be recorded (such as notification sounds of various apps, voices of other VoIP apps), which may cause interference and leak privacy information; in addition, long-term screen recording will increase additional power consumption, and the generated video file will also occupy a large amount of memory.

[0049] In view of this, the present disclosure provides a method and apparatus for recording application audio on a mobile terminal, so as to quickly and accurately record the application's audio data. The method and apparatus for recording application audio on a mobile terminal according to embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.

[0050] Figure 1 This is a flowchart illustrating a method for recording application audio for a mobile terminal according to an embodiment of the present disclosure. The method for recording application audio for a mobile terminal can be executed by an application audio recording APP (hereinafter referred to as a Controller APP) installed on the mobile terminal. The Controller APP can be a system-level APP of the mobile terminal, such as, but not limited to, a voice assistant, a recorder, etc. Alternatively, the Controller APP can also be a newly developed system-level APP of the mobile terminal.

[0051] Reference Figure 1 In step S101, in response to the activation of the application audio recording function, at least one preset application that supports the application audio recording function is determined. Specifically, when the application audio recording function is activated, the Controller APP can register with the mobile terminal's operating system to listen for audio status events to monitor the application's audio status; when a change in the audio status of a preset application is detected, it can be determined that the preset application is in a call state. Here, the application refers to a VoIP APP installed on the mobile terminal, but this disclosure is not limited to this. Figure 2This is an illustration showing an example of enabling the audio recording function in an application. For example... Figure 2 As shown, in response to user actions, the recording application's audio function (i.e., Record voice) can change from the off state shown on the left to the on state shown on the right. Here, user actions may include, but are not limited to, swiping, tapping, and gestures.

[0052] Furthermore, when the application audio recording function is enabled, one or more applications that support the application audio recording function can be identified. Figure 3 This diagram illustrates an example of identifying one or more applications that support the audio recording function. When the audio recording function is enabled, at least one VoIP app that supports it can be selected. Here, in response to the user's selection, WeChat and QQ can be identified as apps that support the audio recording function. On the other hand, in response to the user's selection, DingTalk and Tencent Meeting can be identified as apps that do not support the audio recording function. Optionally, after identifying one or more apps that support the audio recording function, a list of apps that support the audio recording function (i.e., mPackageList) can be automatically generated. Optionally, when the audio recording function is enabled, all VoIP apps can be automatically selected as apps that support the audio recording function.

[0053] Return to reference Figure 1 In step S102, in response to the preset application being in a call state, playback data and recording data of the preset application are recorded. Here, the preset application is an application determined to support the recording function of application audio. According to embodiments of this disclosure, playback data may include audio data received from a remote device and played through a speaker, and recording data may include audio data collected through a microphone. Therefore, step S102 may specifically include: in response to the preset application being in a call state, acquiring the playback data and recording data of the preset application; and then recording the playback data and recording data of the preset application respectively. Here, the speaker and microphone may be the mobile terminal's own speaker and microphone, or they may be external speakers and microphones connected to the mobile terminal (e.g., Bluetooth headsets, Bluetooth speakers, etc.).

[0054] Specifically, when a VoIP app (e.g., APP X) on a mobile terminal is detected receiving / initiating a VoIP call, mPackageList can be used to determine whether APP X supports application audio recording. If APP X supports application audio recording, its playback and recording data can be obtained in real time and then recorded separately. This will be described in detail later. However, if APP X does not support application audio recording, existing audio processing for VoIP calls can be performed without applying audio recording to the VoIP call.

[0055] Figure 4 This is a diagram illustrating the existing audio processing architecture for VoIP calls. For example... Figure 4 As shown, the audio processing architecture during a VoIP call can include operations performed at the application layer, application framework layer, and hardware abstraction layer (HAL). The left-hand flow represents the playback process, and the right-hand flow represents the recording process. Specifically, during a VoIP call, audio data provided by the remote device can be played back via App Playback, the PlaybackThread, and the Audio HAL and output through the speaker. Audio data captured by the microphone can be recorded and provided to the remote device via the Audio HAL, the RecordThread, and App Recording. Here, App Playback, PlaybackThread, Audio HAL, RecordThread, and App Recording, as part of the audio processing architecture, are well known to those skilled in the art and can be implemented using various existing methods; therefore, their detailed descriptions are omitted here.

[0056] According to embodiments of this disclosure, when APP X is detected to be in a call state and APP X supports the application audio recording function, a recording function icon can be displayed on the screen of the mobile terminal. In response to the recording function icon being selected (e.g., but not limited to clicking, swiping, etc.), the playback data and recording data of APP X are acquired and recorded to generate an audio file. According to embodiments of this disclosure, the application audio recording function may include multiple modes, such as, but not limited to, normal mode, mini mode, Plus mode, etc. When the application audio recording function is enabled, the mode can be switched by swiping on the screen of the mobile terminal. However, this disclosure is not limited to this, and various modes can be switched in other ways. Figure 5AThis is a diagram illustrating an example of preparing to record audio in the normal mode (default mode) of a recording application's audio function. Figure 5B This is an illustration showing an example of recording in normal mode. Figure 5C This is an illustration showing an example of pausing recording in normal mode. (Example:) Figure 5A As shown, a recording icon is displayed on the mobile device's screen. At this time, the recording icon displays recording options (a dot with a microphone icon in the image) and a text description ("Record" in the image). When the recording option on the recording icon is clicked (for example, but not limited to this), recording application audio begins, such as... Figure 5B As shown. At this time, the recording function icon displays the pause option (as shown in the image). The recording options include an end option (as shown by the ■ in the image) and the recording time. Clicking the pause option on the recording icon pauses the application's audio recording, such as... Figure 5C As shown in the image. At this point, the recording icon displays recording options, an end option, and the recording time. Optionally, clicking the recording option continues recording application audio, while clicking the end option stops recording application audio. Figure 5D This is an illustration showing an example of preparing to record in mini mode. Figure 5E This is an illustration showing an example of recording in mini mode. (Example:) Figure 5D As shown, the recording function icon is displayed on the mobile device's screen. At this time, the recording function icon only displays the recording option. When the recording option of the recording function icon is clicked (for example, but not limited to this), recording of the application's audio begins, such as... Figure 5E As shown. At this time, the recording icon only displays the recording time. Optionally, when pausing app audio recording (e.g., by tapping the recording icon), the recording icon in mini mode also only displays the recording time. Figure 5F This is an illustration showing an example of preparing to record in Plus mode. Figure 5G This is an illustration showing an example of recording in Plus mode. Figure 5H This is an illustration showing an example of pausing recording in Plus mode. (Example:) Figure 5F As shown, the recording icon is displayed on the mobile device's screen. At this time, the recording icon displays recording options, a text description, and a close option (as shown by the × in the image). When the recording option on the recording icon is clicked, recording of the application's audio begins, as shown... Figure 5G As shown. At this time, the recording icon displays pause, end, recording time, and close options. When the pause option on the recording icon is clicked, recording of the application's audio is paused, as shown. Figure 5HAs shown in the image. At this point, the recording icon displays recording options, an end option, recording time, and a close option. Optionally, clicking the recording option continues recording the application's audio; clicking the end option stops recording the application's audio and returns to the initial screen where you prepared to record, as shown. Figure 5F The interface shown. On the other hand, when the close option is selected, recording application audio can be stopped, an audio file can be automatically generated, and recording of application audio can be exited (for example, application audio cannot be recorded again in the current VoIP call of APP X). Optionally, if the application audio recording function supports automatic recording (i.e., Figure 3 As shown in the image, "Auto record audio" is enabled. Once Auto record audio is activated, the recording icon will be displayed on the mobile terminal's screen, and the app will automatically start acquiring playback and recording data from APP X to record application audio without requiring manual clicking of the recording function icon to start recording.

[0057] The following describes in detail the process of acquiring playback and recording data from a preset application (e.g., APP X) in real time.

[0058] For the audio output of APP X, the audio data received by APP X from a remote device and played through a speaker needs to be backed up to obtain the playback data of APP X. At this point, the backed-up audio data can be used as the playback data of APP X. On the other hand, for the audio input of APP X, the audio data collected through the microphone needs to be bypassed and extracted to obtain the recording data of APP X. At this point, the bypassed audio data can be used as the recording data of APP X.

[0059] Figure 6 This is a diagram illustrating the processing structure for acquiring playback and recording data.

[0060] Reference Figure 6As described above, the audio data provided by the remote device can be output through the speaker via APP Playback, PlaybackThread, and Audio HAL. At this time, in order to obtain playback data (i.e., mRecorderContinuity), the audio data played through the speaker can be backed up, and the data stream of the backed-up audio data can be switched to the Remote Submix HAL layer. The MonoPipe (i.e., audio pipe) in the Remote Submix HAL layer enables audio data connection between the VoIP APP and the Controller APP. That is, the Remote Submix HAL layer can use the MonoPipe to provide the data stream of the backed-up audio data to a RecordThread (first RecordThread). According to embodiments of this disclosure, one VoIP APP corresponds to one MonoPipe. In other words, when recording application audio for multiple applications, the number of MonoPipes can be the same as the number of applications being recorded. Further, after switching the data stream of the backed-up playback data to the Remote Submix HAL layer, the Controller APP can obtain the backed-up playback data from the Remote Submix HAL layer through the first RecordThread and the first App Recording. Here, the Remote Submix HAL layer and MonoPipe, as part of the audio processing architecture, are well known to those skilled in the art and can be implemented in various existing ways, so their detailed description is omitted here. Furthermore, while switching the backup playback data stream to the Remote Submix HAL layer, the original data stream played from the speaker can be retained, thus not interfering with the normal playback of APP X.

[0061] On the other hand, refer to Figure 6As described above, audio data captured by the microphone can be transmitted to the remote device as recording data for the VoIP app via Audio HAL, RecordThread, and the VoIP app's APP Recording. Here, the RecordThread on the transmission path of the recording data is the second RecordThread, which is not the same RecordThread as the first RecordThread on the transmission path of the backup playback data mentioned above. In order to obtain the recording data (i.e., mRecorderMic), the recording data can be extracted from the second RecordThread through the second APP Recording of the Controller app. At the same time, the path for APP X's APP Recording to extract recording data from the second RecordThread is retained, so as not to interfere with the normal recording and transmission of APP X. Here, the first and second APP Recording of the Controller app are essentially the same as the APP Recording of APP X, and can be implemented by those skilled in the art using various existing methods, such as MediaRecorder, AudioRecord, etc. To more conveniently control audio data streams, the Controller App can implement App Recording using AudioRecord. However, the above description is merely illustrative, and this disclosure does not impose any restrictions on the implementation of App Recording.

[0062] Optionally, the Controller App can have the AudioManager, an audio control interface provided by the operating system. By setting corresponding parameters through AudioManager, the Controller App can control whether to acquire audio data played through the speaker by the VoIP App and whether to acquire audio data captured by the microphone by the VoIP App. On the other hand, the PlaybackThread and each RecordThread in the Framework layer described above are configured within AudioFlinger, which is responsible for managing and scheduling the transmission of audio data. AudioFlinger is also configured with SecAudioParamFlinger, used to parse and execute commands issued by AudioManager to control whether to provide audio data played through the speaker and audio data captured by the microphone to the Controller App. Furthermore, AudioPolicy and AudioContinuity are also set up in the Framework layer. AudioPolicy provides audio input / output management, that is, it controls which device and corresponding HAL layer the audio data in the mobile terminal selects. In other words, AudioPolicy can control the Controller App to acquire the corresponding recording and playback data so that the Controller App can record the recording and playback data. AudioContinuity can be used to record parameters set in the Controller App's AudioManager (including the Controller App's package name, the UID of applications that support recording application audio, etc.), so that AudioPolicy and AudioFlinger can correctly control the audio data streams of the Controller App and VoIP App based on the information recorded by AudioContinuity.

[0063] The Remote Submix HAL layer, MonoPipe, AudioManager, SecAudioParamFlinger, AudioPolicy, and AudioContinuity, as described above, are well known to those skilled in the art as part of an audio processing architecture. Embodiments of this disclosure utilize these components to implement methods for recording application audio for mobile terminals. Since they can be implemented in various existing methods, their detailed descriptions are omitted here.

[0064] In this way, playback and audio data can be obtained without the need for other screen recording tools to record the application's audio data, thereby avoiding tedious recording operations and post-editing work, improving recording speed and user experience.

[0065] According to embodiments of this disclosure, to protect user privacy, the extraction of audio data is controlled by setting parameters to determine whether the Controller App is allowed to extract audio data captured through the microphone. Specifically, when the recording application's audio function is not enabled, the Controller App's audio data extraction parameter is set to "SPECIFY_RECORDING_STATE_UNKNOWN". In this case, the Controller App's ability to extract audio data captured through the microphone is determined according to existing microphone usage priority principles.

[0066] Taking Android as an example, for regular applications (such as WeChat, QQ, and a voice recorder), the microphone usage priority principle is: Audio mode owner > Privacy sensitive > TOP > Latest. When an application is set as the Audio mode owner, it has the highest microphone usage priority. When an application is set as privacy sensitive, it has the second highest microphone usage priority. When an application runs in the system foreground (TOP), it has a higher microphone usage priority than an application running in the system background. When two applications have the same priority level, the application that started recording more recently (Latest) has a higher microphone usage priority than the application that started recording earlier. The microphone usage priority principle described above is only an example; other microphone usage priority principles can also be used, and this disclosure does not impose any restrictions on them.

[0067] On the other hand, when the application's audio recording function is enabled, the Controller App monitors the microphone usage of applications in the mPackageList in real time: only when at least one application in the mPackageList is in the "microphone allowed" state, the Controller App's audio data extraction parameter is set to "SPECIFY_RECORDING_STATE_ALLOWED", thus allowing the Controller App to extract audio data captured through the microphone; otherwise, the Controller App's audio data extraction parameter is set to "SPECIFY_RECORDING_STATE_FORBIDDEN", ​​prohibiting the Controller App from extracting audio data captured through the microphone.

[0068] Furthermore, when the application's audio recording function is enabled and the preset application in mPackageList (e.g., APPX) is in the "microphone allowed" state, the Controller APP's audio data extraction parameter is set to "SPECIFY_RECORDING_STATE_ALLOWED" (i.e., the first value) to allow the Controller APP to bypass the audio data captured through the microphone. If, during the recording of the preset application's audio, the preset application's state changes from "microphone allowed" to "microphone not allowed," then the Controller APP's audio data extraction parameter is set to "SPECIFY_RECORDING_STATE_FORBIDDEN" (i.e., the second value) to prevent the Controller APP from bypassing the audio data captured through the microphone. In this case, the recording data obtained by the Controller APP is silent data. For example, if APPX is making a VoIP call and APPX is allowed to use the microphone, the Controller APP's audio data extraction parameter can be set to "SPECIFY_RECORDING_STATE_ALLOWED" to bypass the audio data captured through the microphone. However, if during a VoIP call, another application that does not support recording application audio (e.g., APP Y) performs a VoIP call and is allowed to use the microphone (e.g., when Tencent Meeting is accessed during a WeChat call and Tencent Meeting has the highest microphone usage priority, WeChat's status changes from "Microphone allowed" to "Microphone not allowed"), then the Controller APP's audio data extraction parameter can be set to "SPECIFY_RECORDING_STATE_FORBIDDEN", ​​thereby pausing the bypass extraction of audio data collected through the microphone.

[0069] This way, on the one hand, interference from other applications on the audio being recorded can be avoided, and on the other hand, privacy information can be effectively prevented from being leaked.

[0070] Return to reference Figure 1 After recording the playback and audio data of the preset application, in step S103, an audio file can be generated based on the recorded playback and audio data. Specifically, the acquired playback and audio data of the preset application can be mixed and encoded in real time to generate an audio file. Here, various existing mixing and encoding schemes can be used to mix and encode the playback and audio data, and this disclosure does not impose any restrictions on the mixing and encoding schemes. For example, the mixed playback and audio data can be encoded into an m4a format audio file.

[0071] According to embodiments of this disclosure, the preset application may include two or more applications. In this case, if two or more applications are in a call state simultaneously, in step S102, playback data and audio recording data of each application included in the preset application can be recorded respectively, and in step S103, an audio recording file of each application included in the preset application can be generated based on the playback data and audio recording data of that application.

[0072] Next, in step S104, the generated audio file can be displayed in the file list of the multimedia application. Here, the multimedia application can be the system recorder built into the mobile terminal. However, considering that some mobile terminal manufacturers' built-in recorders will perform special field verification on the audio file, the audio file needs to be preprocessed when generating the audio file so that it can be recognized by the recorder. In this case, in step S103, the recorded playback data and recording data can first be mixed and encoded in real time to obtain a recorded audio file. Then, the recorded audio file can be preprocessed to generate an audio file that the multimedia playback application can recognize, which can then be used as the audio file. For example, special fields can be embedded in the recorded audio file using an M4a editor (i.e., M4aEditor) to generate an audio file that the recorder can recognize. Then, the generated audio file can be added to the file list of the multimedia application by updating the mobile terminal's system database. For example, the generated audio file can be added to the recorder's file list by triggering the operating system to update the mobile terminal's system database through different methods such as notifying the operating system, direct insertion, or triggering a scan using the ForceSync function. In this way, when users open the recorder, they can see the recorded VoIP call audio files.

[0073] According to embodiments of this disclosure, audio recordings can be shared, and various AI technologies can be used to perform AI operations such as transcription, translation, and abstraction on the recordings according to user needs. The AI ​​operations described herein are all existing AI operations, and therefore their detailed descriptions are omitted.

[0074] Figure 7 This is a block diagram illustrating an apparatus for recording application audio for a mobile terminal according to an embodiment of the present disclosure.

[0075] Reference Figure 7An apparatus 700 for recording application audio on a mobile terminal may include an application determination module 710, a recording module 720, an audio recording file generation module 730, and a display control module 740. The application determination module 710, in response to the activation of the application audio recording function, determines at least one preset application that supports the application audio recording function. The recording module 720, in response to the preset application being in a call state, records the playback data and recording data of the preset application. The audio recording file generation module 730 generates an audio recording file based on the recorded playback data and recording data. The display control module 740 displays the audio recording file in the file list of the multimedia application.

[0076] According to embodiments of this disclosure, the device 700 may further include a registration and monitoring module (not shown) and a call status determination module (not shown). The registration and monitoring module may register an audio status monitoring event with the operating system of the mobile terminal in response to the activation of the recording application's audio function, thereby monitoring the audio status of various applications on the mobile terminal. The call status determination module may determine that a preset application is in a call state in response to detecting a change in the audio status of a preset application.

[0077] According to embodiments of this disclosure, the recording module 720 can, in response to a preset application being in a call state, acquire playback data and recording data of the preset application, and then record the playback data and recording data of the preset application respectively. Further, the recording module 720 can back up audio data received by the preset application from a remote device and played through a speaker to obtain the playback data of the preset application; and / or perform bypass extraction on audio data collected through a microphone to obtain the recording data of the preset application. Optionally, the recording module 720 can, in response to a preset application being in a microphone-enabled state, perform bypass extraction on audio data collected through a microphone; and can, in response to a preset application being in a microphone-disabled state, prohibit bypass extraction on audio data input through a microphone.

[0078] According to embodiments of this disclosure, the audio recording file generation module 730 can mix and encode recorded playback data and audio recording data in real time to generate an audio recording file. Further, the audio recording file generation module 730 can mix and encode recorded playback data and audio recording data in real time to obtain a recorded audio file; then, the recorded audio file can be preprocessed to generate an audio file that a multimedia playback application can recognize, which serves as the audio recording file.

[0079] The method for recording application audio for a mobile terminal according to embodiments of the present disclosure can be programmed into a computer program and stored on a computer-readable storage medium. When the computer program is executed by a processor, the method for recording application audio for a mobile terminal as described above can be implemented. Examples of computer-readable storage media include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc storage, hard disk drive (HDD), solid-state drive (SSD), card storage (such as multimedia cards, secure digital (SD) cards, or ultra-fast digital (XD) cards), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid-state drive, and any other device configured to store a computer program and any associated data, data files, and data structures in a non-transitory manner and to provide the computer program and any associated data, data files, and data structures to a processor or computer so that the processor or computer can execute the computer program. In one example, the computer program and any associated data, data files, and data structures are distributed across a networked computer system, such that the computer program and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner through one or more processors or computers.

[0080] Figure 8 This is a block diagram illustrating an electronic device according to an embodiment of the present disclosure.

[0081] Reference Figure 8 The electronic device 800 may include a memory 801 and a processor 802. The memory 801 stores a computer program that, when executed by the processor 802, implements a method for recording application audio for a mobile terminal according to embodiments of the present disclosure.

[0082] As an example, electronic device 800 may be a smartphone, tablet, personal digital assistant, or other mobile electronic device capable of executing the aforementioned computer programs. In electronic device 800, processor 802 may include a central processing unit (CPU), graphics processing unit (GPU), programmable logic device, dedicated processor system, microcontroller, or microprocessor. As an example, and not a limitation, processor may also include analog processors, digital processors, microprocessors, multi-core processors, processor arrays, network processors, etc. Processor 802 may execute instructions or code stored in memory 801, which may also store data. Instructions and data may also be sent and received via a network through a network interface device, which may employ any known transmission protocol. Memory 801 may be integrated with processor 802, for example, by arranging RAM or flash memory within an integrated circuit microprocessor. Furthermore, memory 801 may include separate devices, such as external disk drives, storage arrays, or other storage devices usable by any database system. The memory 801 and the processor 802 may be operatively coupled or may communicate with each other, for example, through I / O ports, network connections, etc., so that the processor 802 can read files stored in the memory.

[0083] In addition, the electronic device 800 may also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, touch input device, etc.). All components of the electronic device 800 can be interconnected via a bus and / or network.

[0084] The method and apparatus for recording application audio on a mobile terminal according to embodiments of this disclosure can directly record application audio data without using other screen recording tools, thereby avoiding cumbersome recording operations and post-editing work, improving recording speed and user experience. Furthermore, since the application audio data is recorded directly without using other screen recording tools, interference from other applications can be avoided, preventing the leakage of privacy information, and reducing screen recording power consumption.

[0085] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0086] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A method for recording application audio on a mobile terminal, characterized in that, The method includes: In response to the activation of the application audio recording function, at least one preset application that supports the application audio recording function is identified. In response to the preset application being in a call state, the playback data and audio recording data of the preset application are recorded; Based on the recorded playback data and audio recording data, generate an audio file; The audio file is displayed in the file list of the multimedia application.

2. The method as described in claim 1, characterized in that, The method further includes: In response to the audio recording function of the application being enabled, register with the mobile terminal's operating system to listen for audio status events in order to monitor the audio status of various applications on the mobile terminal. In response to a change in the audio state of the preset application, it is determined that the preset application is in a call state.

3. The method as described in claim 1, characterized in that, In response to the preset application being in a call state, the steps of recording the playback data and audio data of the preset application include: In response to the preset application being in a call state, the playback data and recording data of the preset application are obtained; Record the playback data and audio data of the preset application respectively.

4. The method as described in claim 3, characterized in that, The steps for obtaining the playback data and recording data of the preset application include: Backup the audio data received by the preset application from the remote device and played through the speaker to obtain the playback data of the preset application; and / or The audio data collected through the microphone is bypassed and extracted to obtain the recording data of the preset application.

5. The method as described in claim 4, characterized in that, The steps for bypassing audio data acquired through a microphone include: In response to the preset application being in a state where microphone use is allowed, the audio data collected through the microphone is extracted by bypass. In response to the preset application being in a state where microphone use is not allowed, bypass extraction of audio data input through the microphone is prohibited.

6. The method as described in claim 1, characterized in that, The steps for generating an audio file based on the recorded playback data and audio recording data include: The recorded playback data and audio recording data are mixed and encoded in real time to generate an audio file.

7. The method as described in claim 6, characterized in that, The step of generating an audio file based on the recorded playback data and audio recording data further includes: The recorded playback data and audio recording data are mixed and encoded in real time to obtain the recorded audio file; The recorded audio file is preprocessed to generate an audio file that the multimedia playback application can recognize, which is then used as the recording file.

8. An apparatus for recording application audio on a mobile terminal, characterized in that, The device includes: The application determination module is configured to: in response to the activation of the application audio recording function, determine at least one preset application that supports the application audio recording function; The recording module is configured to: record the playback data and audio data of the preset application in response to the preset application being in a call state; The audio recording file generation module is configured to generate audio recording files based on the recorded playback data and audio recording data. The display control module is configured to display the audio file in the file list of a multimedia application.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method for recording application audio for a mobile terminal as described in any one of claims 1 to 7.

10. An electronic device, characterized in that, The electronic device includes: processor; A memory storing a computer program that, when executed by the processor, implements the method for recording application audio for a mobile terminal as described in any one of claims 1 to 7.