Audio processing device and terminal equipment
Through the audio service module and the sound effect unit in the sound effect hardware abstraction layer, the first audio data is converted into the second audio data, which solves the problem that multi-channel audio data cannot be obtained through a small number of recording channels in the prior art, and realizes simplified multi-channel audio data acquisition and recording and playback of spatial audio data.
Patent Information
- Application Number
- CN202311465170.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-03
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, more channels of audio data cannot be obtained through fewer recording channels in the device, resulting in complex ways of obtaining multi-channel audio data.
Through the audio service module and the sound effect hardware abstraction layer, including the sound effect unit, the first audio data (the number of channels is smaller than the second audio data) is converted into the second audio data (the number of channels is larger than the first audio data), thereby realizing the acquisition of multi-channel audio data through a small number of recording channels.
The process of acquiring multi-channel audio data is simplified, the application range is improved, and the recording and playback of spatial audio data can be realized.
Smart Images

Figure CN119946499A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to but are not limited to electronic technology, and in particular, to an audio processing device and a terminal equipment. Background Art
[0002] In multimedia and communication systems, multi-channel audio is played more and more frequently. For example, taking multi-channel audio as spatial audio, spatial audio technology is an extension technology based on stereo surround sound. Compared with stereo surround sound, it has more support for sound decoding and channel technology, and can more easily present the layering and stereoscopic sense of sound. There is no solution in the related technology on how to obtain audio data of more channels through fewer recording channels in the device. Summary of the invention
[0003] The embodiments of the present application provide an audio processing device and a terminal device.
[0004] In a first aspect, an embodiment of the present application provides an audio processing device, the audio processing device comprising: an audio service module and a sound effect hardware abstraction layer, the sound effect hardware abstraction layer comprising a sound effect unit;
[0005] The audio service module is used to: determine first audio data, and send the first audio data to the sound effect unit;
[0006] The audio effect unit is used to: convert the first audio data into second audio data, and send the second audio data to the audio service module; wherein the number of channels corresponding to the first audio data is smaller than the number of channels corresponding to the second audio data.
[0007] In a second aspect, an embodiment of the present application provides a terminal device, comprising the above-mentioned audio processing device.
[0008] In an embodiment of the present application, the audio service module is used to determine the first audio data and send the first audio data to the sound effect unit; the sound effect unit is used to convert the first audio data into the second audio data and send the second audio data to the audio service module; wherein the number of channels corresponding to the first audio data is less than the number of channels corresponding to the second audio data. In this way, the audio service module in the audio processing device can determine the first audio data, and the sound effect unit in the audio processing device can determine the corresponding second audio data, and the number of channels corresponding to the second audio data is greater than the number of channels corresponding to the first audio data, thereby being able to obtain audio data of more channels through fewer recording channels in the device, solving the problem in the related art that the audio data of the fixed channel can only be obtained through the recording channel of the fixed channel, thereby resulting in a complicated method for obtaining multi-channel audio data. Therefore, the present application can improve the application scope of obtaining multi-channel audio data. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The drawings described herein are used to provide further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute improper limitations on the present application.
[0010] Figure 1 A schematic diagram of the structure of an audio processing device provided in an embodiment of the present application;
[0011] Figure 2 A schematic diagram of the structure of another audio processing device provided in an embodiment of the present application;
[0012] Figure 3 A structural diagram of another audio processing device provided in an embodiment of the present application;
[0013] Figure 4 A schematic diagram of the structure of another audio processing device provided in an embodiment of the present application;
[0014] Figure 5 A schematic diagram of the structure of an audio processing device provided in another embodiment of the present application;
[0015] Figure 6 A structural schematic diagram of an audio processing device provided in yet another embodiment of the present application;
[0016] Figure 7 A structural schematic diagram of an audio processing device provided in yet another embodiment of the present application;
[0017] Figure 8 A schematic diagram of the structure of another audio processing device provided in another embodiment of the present application;
[0018] Figure 9a A flowchart of an audio processing method provided in an embodiment of the present application;
[0019] Figure 9b A schematic diagram of the structure of another audio processing device provided in another embodiment of the present application;
[0020] Fig.10 A schematic diagram of an audio processing process provided in an embodiment of the present application;
[0021] Fig.11 A schematic diagram of a static structure corresponding to an audio processing flow provided in an embodiment of the present application;
[0022] Fig.12 A flowchart of a recording channel creation method provided in an embodiment of the present application;
[0023] Fig.13A flowchart of a method for configuring an upmixer buffer provider provided in an embodiment of the present application;
[0024] Fig.14 A flowchart of a recording channel release method provided in an embodiment of the present application;
[0025] Fig.15 A flowchart of a recording path data processing method provided in an embodiment of the present application;
[0026] Fig.16 A schematic flow chart of a data processing method of an upmixer buffer provider provided in an embodiment of the present application;
[0027] Fig.17 A schematic diagram of file locations corresponding to an audio processing flow provided in an embodiment of the present application;
[0028] Fig.18 A schematic diagram of the structure of a terminal device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0029] The technical solution of the present application will be described in detail below through embodiments and in conjunction with the accompanying drawings. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0030] It should be noted that in the examples of the present application, "first", "second", etc. are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
[0031] In addition, the technical solutions described in the embodiments of the present application can be combined arbitrarily without conflict. In the description of the present application, "multiple" means two or more, unless otherwise clearly and specifically defined.
[0032] In the related art, the audio processing device or terminal device used to obtain two-channel audio data can only record mono audio or two-channel audio, but cannot record spatial audio. In addition, in the related art, although there are terminal devices that support spatial audio playback with head tracking effects, the source is limited and can only be provided by a third party, and the terminal device cannot independently record the source of spatial audio.
[0033] The technical solution provided by the embodiment of the present application can enable users to easily record spatial audio sources for playback and publishing. For example, in the related art, a mobile phone generally has two microphones, so when recording audio data through a mobile phone, only dual-channel audio data can be recorded, and the final audio obtained is also dual-channel audio. The technicians of this case discovered this problem and proposed the present solution, which is to obtain spatial audio data through the two microphones of the mobile phone, so that users can obtain spatial audio data independently through the mobile phone. In this way, at a time when audio and video are exploding, the spatial audio data recorded by a user through his or her mobile phone can be played by himself or other users, so that the user at the playback end has an immersive feeling. Such a solution is very meaningful and does not exist at present.
[0034] The audio processing device in the embodiment of the present application may be included in a terminal device, or the audio processing device in the embodiment of the present application may be a terminal device. The terminal device in any embodiment of the present application may include one of the following or a combination of at least two of them: Internet of Things (IoT) devices, satellite terminals, wireless local loop (WLL) stations, personal digital assistants (PDA), handheld devices with wireless communication functions, computing devices or other processing devices connected to wireless modems, servers, mobile phones, tablet computers (Pad), computers with wireless transceiver functions, handheld computers, desktop computers, personal digital assistants, portable media players, smart speakers, navigation devices, smart watches, smart glasses, smart necklaces and other wearable devices, learning machines, translation pens, translation machines, point reading machines, pedometers, digital TVs, virtual reality (VR) terminal devices, augmented reality (AR) terminal devices, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, and wireless terminals in smart cities. The wireless terminals in the Internet of Vehicles (IoV) system include wireless terminals in smart cities, wireless terminals in smart homes, and vehicles, on-board equipment, on-board modules, wireless modems, handheld devices, customer premises equipment (CPE), smart home appliances, etc. in the IoV system.
[0035] Figure 1 A schematic diagram of the structure of an audio processing device provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, the audio processing device 10 includes: an audio service module 11 and a sound effect hardware abstraction layer 12 , and the sound effect hardware abstraction layer 12 includes a sound effect unit 121 .
[0036] The audio service module 11 is used to: determine first audio data; send the first audio data to the sound effect unit 121;
[0037] The audio effect unit 121 is used to: convert the first audio data into second audio data, and send the second audio data to the audio service module 11; wherein the number of channels corresponding to the first audio data is smaller than the number of channels corresponding to the second audio data.
[0038] Exemplarily, the first audio data may be two-channel audio data, and the second audio data may be spatial audio data. Also exemplarily, the first audio data may be monophonic audio data, and the second audio data may be spatial audio data. Also exemplarily, the first audio data may be three-channel audio data, and the second audio data may be spatial audio data having more than three channels. Also exemplarily, the first audio data may be monophonic audio data, and the second audio data may be two-channel audio data. It should be noted that the present application does not limit the number of channels corresponding to the first audio data and the second audio data, as long as the number of channels corresponding to the first audio data is less than the number of channels corresponding to the second audio data, it should be within the protection scope of the present application.
[0039] In any embodiment of the present application, the channel has a one-to-one correspondence with the channel. For example, the number of channels corresponding to N-channel audio data is also N, where N is an integer greater than or equal to 1.
[0040] Exemplarily, the number of channels corresponding to the first audio data is 2, and the number of channels corresponding to the second audio data is 6 or 8. Exemplarily, the number of channels 2 corresponds to 2 channels, and the number of channels 6 or 8 corresponds to 5.1 channels or 7.1 channels, respectively. Among them, the 5.1 channel consists of 6 independent audio channels, including left front, center front, right front, left rear, right rear and subwoofer channels. This configuration is suitable for small to medium-sized rooms and provides a range of surround sound effects, allowing the audience to feel the sound effects in the front, back and left and right directions. The subwoofer channel is specifically responsible for processing bass frequencies, enhancing the bass effect, and making the audio richer and more dynamic. Among them, the 7.1 channel adds two additional audio channels on the basis of 5.1, namely the left rear center and right rear center channels. This can provide the audience with more precise and accurate rear sound effects. The 7.1 channel system is suitable for larger rooms or users who need a higher level of audio experience. Through the additional audio channels, the 7.1 channel system can better reproduce the surround sound effects, allowing the audience to experience a more real and realistic audio experience.
[0041] In some embodiments, converting the first audio data into the second audio data may include: converting the first audio data into the seventh audio data, spatializing the seventh audio data to obtain the second audio data; wherein the number of channels corresponding to the seventh audio data is the same as the number of channels corresponding to the second audio data. In some embodiments, spatializing the seventh audio data may include performing at least one of the following processing on the seventh audio data: equalizer adjustment, gain adjustment, reverberation adjustment, and sound effect adjustment. Wherein, the sound effect adjustment is performed to match the second audio data with the default playback mode of the audio processing device or the playback mode set by the user through the terminal device. The playback mode may include at least one of the following: retro mode, vinyl mode, record mode, youth mode, alcohol mode, etc. In other embodiments, in some embodiments, spatializing the seventh audio data may include: spatializing the seventh audio data based on a target algorithm; wherein the target algorithm may include one of the following: vector base Amplitude Panning (VBAP), distance-based amplitude positioning (DBAP), wave field synthesis (WFS), etc.
[0042] In some embodiments, converting the first audio data into the second audio data / seventh audio data may include: identifying the audio and video information of each sound source in the first audio data; and determining the second audio data / seventh audio data according to the audio and video information of each sound source. In some embodiments, the audio and video information of the sound source may include at least one of the following: sound feature information of the sound source, position change information of the sound source relative to the terminal device, position information of the sound source relative to the terminal device, name information of the sound source, etc.
[0043] It should be noted that the present application does not limit the specific implementation method of converting the first audio data into the second audio data / seventh audio data, and any implementation method should be within the protection scope of the present application.
[0044] In an embodiment of the present application, the audio service module is used to determine the first audio data and send the first audio data to the sound effect unit; the sound effect unit is used to convert the first audio data into the second audio data and send the second audio data to the audio service module; wherein the number of channels corresponding to the first audio data is less than the number of channels corresponding to the second audio data. In this way, the audio service module in the audio processing device can determine the first audio data, and the sound effect unit in the audio processing device can determine the corresponding second audio data, and the number of channels corresponding to the second audio data is greater than the number of channels corresponding to the first audio data, thereby being able to obtain audio data of more channels through fewer recording channels in the device, solving the problem in the related art that the audio data of the fixed channel can only be obtained through the recording channel of the fixed channel, thereby resulting in a complicated method for obtaining multi-channel audio data. Therefore, the present application can improve the application scope of obtaining multi-channel audio data.
[0045] Figure 2 A structural diagram of another audio processing device provided in an embodiment of the present application, such as Figure 2 As shown, the audio processing device 10 includes: an audio service module 11 and a sound effect hardware abstraction layer 12, and the sound effect hardware abstraction layer 12 includes a sound effect unit 121; in addition, the audio processing device also includes: a recording hardware abstraction layer 13.
[0046] The audio service module 11 is further used to: obtain third audio data from the recording hardware abstraction layer 13, determine the first audio data according to the third audio data; determine fourth audio data according to the second audio data;
[0047] The number of channels corresponding to the first audio data is the same as the number of channels corresponding to the third audio data, and the number of channels corresponding to the second audio data is the same as the number of channels corresponding to the fourth audio data.
[0048] In some embodiments, the audio processing device 10 also includes an audio acquisition module 14, and the audio acquisition module 14 can acquire original audio data. Exemplarily, the original audio data may include original dual-channel audio data. For example, the audio acquisition module 14 can collect original dual-channel audio data. Exemplarily, the audio acquisition module 14 may include a microphone, and the microphone is used to collect original dual-channel audio data. For another example, the audio acquisition module 14 can be connected to other devices for communication and receive original dual-channel audio data sent by other devices. Exemplarily, other devices may include microphones or Bluetooth devices. Bluetooth devices may include Bluetooth headsets, Bluetooth speakers, smart watches and other devices. During the implementation of this application, the audio acquisition module 14 of this solution can acquire dual-channel audio data, but cannot acquire spatial audio data.
[0049] In some embodiments, the recording hardware abstraction layer 13 may receive original audio data from the audio acquisition module 14 , and provide an interface corresponding to the third audio data to the audio service module 11 , so that the audio service module 11 obtains the third audio data from the recording hardware abstraction layer 13 .
[0050] In some embodiments, the recording hardware abstraction layer 13 may receive original audio data from the audio acquisition module 14, and determine the original audio data as the third audio data. In other embodiments, the recording hardware abstraction layer 13 may receive original audio data from the audio acquisition module 14, and perform audio processing on the original audio data to obtain the third audio data. The audio processing may include one or more of the following: noise suppression processing, pitch rise and fall processing, audio enhancement processing, timbre processing, and audio special effects processing.
[0051] In some embodiments, determining the first audio data according to the third audio data may include: determining the third audio data as the first audio data. In other embodiments, determining the first audio data according to the third audio data may include: resampling and / or quantization bit conversion of the third audio data to determine the first audio data.
[0052] In any embodiment of the present application, the number of quantization bits may be referred to as bit width.
[0053] In some embodiments, the sampling rate of the third audio data is the first sampling rate, and the number of quantization bits of the third audio data is the first quantization bit number. In some embodiments, the sampling rate of the first audio data is the second sampling rate, and the number of quantization bits of the first audio data is the second quantization bit number. In some embodiments, the first sampling rate is the same as the second sampling rate, and / or the first quantization bit number is the same as the second quantization bit number. In other embodiments, the first sampling rate is different from the second sampling rate, and / or the first quantization bit number is different from the second quantization bit number.
[0054] In some embodiments, the sampling rate supported by the sound effect unit 121 includes a second sampling rate, and the number of quantization bits supported by the sound effect unit 121 includes a second number of quantization bits.
[0055] In some embodiments, the sampling rate supported by the sound effect unit 121 is a sampling rate, which is a second sampling rate, and / or the number of quantization bits supported by the sound effect unit 121 is a quantization bit number, which is a second quantization bit number.
[0056] In some embodiments, the sampling rates supported by the sound effect unit 121 include multiple sampling rates, the multiple sampling rates include a second sampling rate, and / or the number of quantization bits supported by the sound effect unit 121 includes a second number of quantization bits, and the multiple number of quantization bits include a second number of quantization bits. In some implementations, the second sampling rate may be a sampling rate with the smallest absolute value of the difference from the first sampling rate among the multiple sampling rates, and / or the second number of quantization bits may be a quantization bit with the smallest absolute value of the difference from the first number of quantization bits among the multiple number of quantization bits.
[0057] In some embodiments, the sampling rate required for recording is a third sampling rate, and the number of quantization bits required for recording is a third number of quantization bits.
[0058] In some embodiments, the second sampling rate is the same as the third sampling rate, and / or the second quantization bit number is the same as the third quantization bit number. In other embodiments, the second sampling rate is different from the third sampling rate, and / or the second quantization bit number is different from the third quantization bit number.
[0059] In some embodiments, the sampling rate supported by the sound effect unit 121 includes multiple sampling rates, the multiple sampling rates include the second sampling rate, and / or the number of quantization bits supported by the sound effect unit 121 includes the second number of quantization bits, and the multiple number of quantization bits includes the second number of quantization bits. In some implementations, the second sampling rate can be the sampling rate with the smallest absolute value of the difference with the third sampling rate among the multiple sampling rates, and / or the second number of quantization bits can be the number of quantization bits with the smallest absolute value of the difference with the third number of quantization bits among the multiple number of quantization bits. In other embodiments, the second sampling rate can be between the first sampling rate and the third sampling rate, and / or the second number of quantization bits can be between the first number of quantization bits and the third number of quantization bits. For example, the difference between the second sampling rate and the first sampling rate, and the difference between the third sampling rate and the second sampling rate are less than or equal to the first threshold, and / or the difference between the second number of quantization bits and the first number of quantization bits, and the difference between the third number of quantization bits and the second number of quantization bits are less than or equal to the second threshold.
[0060] In some embodiments, the third sampling rate and / or third number of quantization bits required for recording may be configured by the user through the terminal device before triggering video recording or triggering audio recording, or may be called by the user from the default configuration of the terminal device when triggering video recording or triggering audio recording.
[0061] In some embodiments, determining the fourth audio data according to the second audio data may include: determining the second audio data as the fourth audio data. In other embodiments, determining the fourth audio data according to the second audio data may include: resampling and / or quantization bit conversion of the second audio data to determine the fourth audio data.
[0062] In some embodiments, the sampling rate of the second audio data is the second sampling rate, and the number of quantization bits of the second audio data is the second quantization bit number. In some embodiments, the sampling rate of the fourth audio data is the third sampling rate, and the number of quantization bits of the fourth audio data is the third quantization bit number. In some embodiments, the second sampling rate is the same as the third sampling rate, and / or the second quantization bit number is the same as the third quantization bit number. In other embodiments, the second sampling rate is different from the third sampling rate, and / or the second quantization bit number is different from the third quantization bit number.
[0063] Figure 3 A structural diagram of another audio processing device provided in an embodiment of the present application is shown in FIG. Figure 3 As shown, the audio processing device 10 includes: a recording hardware abstraction layer 13, an audio service module 11, an audio effect hardware abstraction layer 12 and an audio acquisition module 14, wherein the audio effect hardware abstraction layer 12 includes an audio effect unit 121. The audio processing device 10 also includes a media service module 15 and an application module 16.
[0064] The audio service module 11 is also used to send the fourth audio data to the media service module 15; the media service module 15 is used to perform audio encoding on the fourth audio data and determine the target space audio data; the application module 16 is used to obtain the target space audio data from the media service module 15 and generate a target audio file corresponding to the target space audio data.
[0065] In some embodiments, the second audio data determined by the sound effect unit 121 may be spatial audio data that has not been spatially rendered, and the media service module 15 is further configured to perform audio encoding on the second audio data after spatial rendering to determine the target spatial audio data. In some embodiments, the second audio data determined by the sound effect unit 121 may be spatial audio data that has been spatially rendered, and the media service module 15 does not need to perform spatial rendering, but instead performs audio encoding on the fourth audio data to determine the target spatial audio data.
[0066] In some embodiments, the format of the fourth audio data may be a Pulse Code Modulation (PCM) format. In some embodiments, the format of the target space audio data may be one of the following: Moving Picture Experts Group Audio Layer III (MP3), Advanced Audio Coding (AAC), Waveform Audio File Format (WAV), etc.
[0067] In some embodiments, the application module 16 may include a recording application module or a camera application module.
[0068] In some embodiments, when the user triggers video recording, a target video file corresponding to the target spatial audio data is generated. In some embodiments, when the user triggers audio recording, a target audio file corresponding to the target spatial audio data is generated. In some embodiments, the target video file or the target audio file can be stored in the terminal device, and the target video file or the target audio file can be played by the user's trigger.
[0069] In some embodiments, the target audio file or the target video file can be used for playback. In some embodiments, the number of speakers of the device playing the target audio file or the target video file can be the same as or different from the number of channels corresponding to the second audio data. In some embodiments, the device playing the target audio file or the target video file can be the device recorded to the target audio file or the target video file, or it is not the device recorded to the target audio file or the target video file.
[0070] When the number of speakers of the device playing the target audio file or the target video file is different from the number of channels corresponding to the second audio data, the device can decode the target audio file or the target video file, and then convert the decoded audio data so that the number of channels corresponding to the converted audio data is the same as the number of speakers of the device. It should be noted that when the converted audio data is played, the user still feels the effect of spatial audio. For example, the audio processing device has two speakers, and the media service module of the audio processing device can be used to decode the target audio file or the target video file to obtain the decoded audio data; the audio service module can be used to convert the sampling rate and / or quantization bit number of the decoded audio data to obtain the audio data to be processed, or the audio service module can determine the decoded audio data as the audio data to be processed, and the sound effect unit can obtain the processed audio data (the corresponding number of channels is the same as the number of speakers of the audio processing device); the audio service unit can also convert the sampling rate and / or quantization bit number of the processed audio data to obtain the final data, or the audio service unit can also determine the processed audio data as the final data, and the final data is used for the speaker to play.
[0071] Figure 4 A structural diagram of another audio processing device provided in an embodiment of the present application is shown in FIG. Figure 4 As shown, the audio processing device 10 includes: a recording hardware abstraction layer 13, an audio service module 11, a sound effect hardware abstraction layer 12 and an audio acquisition module 14, wherein the sound effect hardware abstraction layer 12 includes a sound effect unit 121. The audio service module 11 includes a sound effect provider 111 and a converter 112;
[0072] The converter 112 is used to: obtain the first audio data, and send the first audio data to the sound effect provider;
[0073] The sound effect provider 111 is used to: receive the first audio data sent by the converter, and send the first audio data to the sound effect unit 121; receive the second audio data sent by the sound effect unit 121, and send the second audio data to the converter 112;
[0074] The converter 112 is used to receive the second audio data sent by the sound effect provider.
[0075] In this way, the sound effect provider 111 calls the sound effect unit 121 to obtain the second audio data, so that the method of determining the second audio data is easy to expand.
[0076] Figure 5A structural diagram of an audio processing device provided in another embodiment of the present application is shown in FIG. Figure 5 As shown, the audio processing device 10 includes: a recording hardware abstraction layer 13, an audio service module 11, a sound effect hardware abstraction layer 12 and an audio acquisition module 14, wherein the sound effect hardware abstraction layer 12 includes a sound effect unit 121. The audio service module 11 includes a sound effect provider 111, a converter 112 and a conversion provider 113;
[0077] The converter 112 is used for:
[0078] Acquire the third audio data; when the first sampling rate of the third audio data is different from the second sampling rate supported by the sound effect unit and / or the first quantization bit number is different from the second quantization bit number supported by the sound effect unit, call the conversion provider 113 to resample the third audio data and / or convert the quantization bit number to obtain the first audio data of the second sampling rate and / or the second quantization bit number; or, when the first sampling rate of the third audio data is the same as the second sampling rate supported by the sound effect unit, and the first quantization bit number is the same as the second quantization bit number supported by the sound effect unit, determine the third audio data as the first audio data; and / or,
[0079] When the second sampling rate of the second audio data is different from the third sampling rate required for recording and / or the second quantization bit number is different from the third quantization bit number required for recording, the conversion provider 113 is called to resample the second audio data and / or convert the quantization bit number to obtain the fourth audio data of the third sampling rate and / or the third quantization bit number; or, when the second sampling rate of the second audio data is the same as the third sampling rate required for recording, and the second quantization bit number is the same as the third quantization bit number required for recording, the second audio data is determined as the fourth audio data.
[0080] In some embodiments, when the first sampling rate of the third audio data is the same as the second sampling rate supported by the sound effect unit, and the first quantization bit number is the same as the second quantization bit number supported by the sound effect unit, the converter 112 may directly determine the third audio data as the first audio data. In other embodiments, when the first sampling rate of the third audio data is the same as the second sampling rate supported by the sound effect unit, and the first quantization bit number is the same as the second quantization bit number supported by the sound effect unit, the converter 112 may send the third audio data to the conversion provider 113, so that the conversion provider 113 determines the third audio data as the first audio data and sends the first audio data to the converter 112.
[0081] In some embodiments, when the second sampling rate of the second audio data is the same as the third sampling rate required for recording, and the second quantization bit number is the same as the third quantization bit number required for recording, the converter 112 may directly determine the second audio data as the fourth audio data. In other embodiments, when the first sampling rate of the third audio data is the same as the second sampling rate supported by the sound effect unit, and the first quantization bit number is the same as the second quantization bit number supported by the sound effect unit, the converter 112 may send the second audio data to the conversion provider 113, so that the conversion provider 113 determines the second audio data as the fourth audio data and sends the fourth audio data to the converter 112.
[0082] In some embodiments, the third audio data may be the same as the first audio data. In other embodiments, the third audio data may be different from the first audio data, and the third audio data needs to be resampled and / or quantized to obtain the first audio data.
[0083] In some embodiments, the conversion provider 113 may obtain the first content and / or the second content, and determine the first audio data according to the first content and / or the second content and the third audio data, wherein the first content includes the first sampling rate and the second sampling rate, and the second content includes the first quantization bit number and the second quantization bit number.
[0084] In some embodiments, the sound effect unit 121 may obtain third content, and determine the second audio data according to the third content and the first audio data, wherein the third content includes the number of channels corresponding to the first audio data and the number of channels corresponding to the second audio data.
[0085] In some embodiments, the second audio data may be the same as the fourth audio data. In other embodiments, the second audio data may be different from the fourth audio data, and the second audio data needs to be resampled and / or quantized to obtain the fourth audio data.
[0086] In some embodiments, the conversion provider 113 may obtain the fourth content and / or the fifth content, and determine the fourth audio data according to the fourth content and / or the fifth content and the second audio data, wherein the fourth content includes the second sampling rate and the third sampling rate, and the fifth content includes the second quantization bit number and the third quantization bit number.
[0087] In some embodiments, the converter 112 calls the conversion provider 113 to resample and / or convert the number of quantization bits of the third audio data to obtain the first audio data of the second sampling rate and / or the second number of quantization bits, which may include: when the first sampling rate of the third audio data is different from the second sampling rate supported by the sound effect unit, the converter 112 configures the first content to the conversion provider 113, the first content includes the first sampling rate and the second sampling rate, and / or, when the first quantization bit of the third audio data is different from the second quantization bit supported by the sound effect unit, the converter 112 configures the second content to the conversion provider 113, the second content includes the first quantization bit and the second quantization bit. In addition, the converter 112 also sends the third audio data to the conversion provider 113, so that the conversion provider 113 resamples and / or converts the number of quantization bits of the third audio data according to the first content and / or the second content to obtain the first audio data of the second sampling rate and / or the second number of quantization bits. Further, the conversion provider 113 can send the first audio data to the converter 112.
[0088] In some embodiments, the converter 112 calls the conversion provider 113 to resample and / or convert the number of quantization bits of the second audio data to obtain the fourth audio data of the third sampling rate and / or the third number of quantization bits, which may include: when the second sampling rate of the second audio data is different from the third sampling rate required for recording, the converter 112 configures the fourth content to the conversion provider 113, and the fourth content includes the second sampling rate and the third sampling rate, and / or, when the second quantization bit number of the second audio data is different from the third quantization bit number required for recording, the converter 112 configures the fifth content to the sound effect provider 111, and the fifth content includes the second quantization bit number and the third quantization bit number. In addition, the converter 112 also sends the second audio data to the conversion provider 113, so that the conversion provider 113 resamples and / or converts the number of quantization bits of the second audio data according to the fourth content and / or the fifth content to obtain the fourth audio data of the third sampling rate and / or the third number of quantization bits. Further, the conversion provider 113 can send the fourth audio data to the converter 112.
[0089] It should be noted that in the above embodiment, the converter 112 calls the conversion provider 113 to resample and / or perform quantization bit conversion on the third audio data to obtain the first audio number, and / or resample and / or perform quantization bit conversion on the second audio data to obtain the fourth audio data. In other embodiments of the present application, the converter 112 can resample and / or perform quantization bit conversion on the third audio data by itself to obtain the first audio number, and / or resample and / or perform quantization bit conversion on the second audio data by itself to obtain the fourth audio data.
[0090] Figure 6 A structural diagram of an audio processing device provided in another embodiment of the present application is shown in FIG. Figure 6 As shown, the audio processing device 10 includes: a recording hardware abstraction layer 13, an audio service module 11, a sound effect hardware abstraction layer 12 and an audio acquisition module 14, wherein the sound effect hardware abstraction layer 12 includes a sound effect unit 121. The audio service module 11 also includes a sound effect provider 111, a converter 112, a conversion provider 113, and an audio manager 114;
[0091] The audio manager 114 is used to create a recording thread 115, and the recording thread 115 is used to create a recording track 116 configured corresponding to the recording thread 115, and the recording track 116 is used to configure at least one of the following to the converter 112: a first content, a second content, a third content, a fourth content, and a fifth content; the first content includes the first sampling rate and the second sampling rate, the second content includes the first quantization bit number and the second quantization bit number, the third content includes the first channel number and the second channel number, the first channel number is the channel number corresponding to the first audio data, the second channel number is the channel number corresponding to the second audio data, the fourth content includes the second sampling rate and the third sampling rate, and the fifth content includes the second quantization bit number and the third quantization bit number;
[0092] The converter 112 is used for at least one of the following: when the first sampling rate and the second sampling rate are different, configuring the first content to the conversion provider 113; when the first quantization bit number is different from the second quantization bit number, configuring the second content to the conversion provider 113; in the case of upmix recording, configuring the third content to the sound effect provider 111 so that the sound effect provider 111 configures the third content to the sound effect unit 121; when the second sampling rate and the third sampling rate are different, configuring the fourth content to the conversion provider 113; when the second quantization bit number is different from the third quantization bit number, configuring the fifth content to the conversion provider 113.
[0093] Figure 7 A structural diagram of an audio processing device provided in yet another embodiment of the present application is shown in FIG. Figure 7As shown, the audio processing device 10 includes: a recording hardware abstraction layer 13, an audio service module 11, a sound effect hardware abstraction layer 12 and an audio acquisition module 14, wherein the sound effect hardware abstraction layer 12 includes a sound effect unit 121. The audio service module 11 includes a sound effect provider 111, a converter 112, and a conversion provider 113; the conversion provider 113 includes: a resampling buffer provider 1131 and a reformatting buffer provider 1132.
[0094] The resampling buffer provider 1131 is used to: receive the third audio data sent by the converter 112; when the first sampling rate of the third audio data is different from the second sampling rate, resample the third audio data to obtain fifth audio data, or when the first sampling rate of the third audio data is the same as the second sampling rate, determine the third audio data as fifth audio data; and send the fifth audio data to the reformatting buffer provider 1132;
[0095] The reformatting buffer provider 1132 is used to: receive the fifth audio data sent by the resampling buffer provider; when the first quantization bit number of the fifth audio data is different from the second quantization bit number, convert the quantization bit number of the fifth audio data to obtain the first audio data of the second quantization bit number, or when the first quantization bit number of the fifth audio data is the same as the second quantization bit number, determine the fifth audio data as the first audio data; and send the first audio data to the converter 112.
[0096] The number of channels corresponding to the fifth audio data is the same as the number of channels corresponding to the third audio data.
[0097] In some embodiments, the reformatting buffer provider 1132 sends the first audio data to the converter 112, which may include: the reformatting buffer provider 1132 sends the first audio data directly to the converter 112, or the reformatting buffer provider 1132 sends the first audio data to the converter 112 through the resampling buffer provider 1131.
[0098] In any embodiment of the present application, the sound effect provider 111 may be an upmixer buffer provider, or the sound effect provider 111 may include an upmixer buffer provider. Exemplarily, the sound effect provider 111 belongs to a class, and the upmixer buffer provider belongs to an object.
[0099] In some embodiments, the resampling buffer provider 1131 is also used to: receive the second audio data sent by the converter 112; when the second sampling rate of the second audio data is different from the third sampling rate required for recording, resample the second audio data to obtain sixth audio data, or, when the second sampling rate of the second audio data is the same as the third sampling rate required for recording, determine the second audio data as sixth audio data; and send the sixth audio data to the reformatting buffer provider 1132.
[0100] The number of channels corresponding to the sixth audio data is the same as the number of channels corresponding to the second audio data.
[0101] In some embodiments, the reformatting buffer provider 1132 is also used to: receive the sixth audio data sent by the resampling buffer provider; when the second quantization bit number of the sixth audio data is different from the third quantization bit number required for recording, convert the quantization bit number of the sixth audio data to obtain fourth audio data, or when the second quantization bit number of the sixth audio data is the same as the third quantization bit number required for recording, determine the sixth audio data as fourth audio data; and send the fourth audio data to the converter 112.
[0102] In some embodiments, the reformatting buffer provider 1132 sending the fourth audio data to the converter 112 may include: the reformatting buffer provider 1132 sending the fourth audio data directly to the converter 112, or the reformatting buffer provider 1132 sending the fourth audio data to the converter 112 through the resampling buffer provider 1131.
[0103] Figure 8 A structural diagram of another audio processing device provided in another embodiment of the present application is shown in FIG. Figure 8 As shown, the audio processing device 10 includes: a recording hardware abstraction layer 13, an audio service module 11, a sound effect hardware abstraction layer 12 and an audio acquisition module 14, wherein the sound effect hardware abstraction layer 12 includes a sound effect unit 121. The audio service module 11 includes a sound effect provider 111, a converter 112, a conversion provider 113, and an audio manager 114; the conversion provider 113 includes: a resampling buffer provider 1131 and a reformatting buffer provider 1132.
[0104] The audio manager 114 is used to create a recording thread 115, and the recording thread 115 is used to create a recording track 116 configured corresponding to the recording thread 115, and the recording track 116 is used to configure at least one of the following to the converter 112: a first content, a second content, a third content, a fourth content, and a fifth content; the first content includes the first sampling rate and the second sampling rate, the second content includes the first quantization bit number and the second quantization bit number, the third content includes the first channel number and the second channel number, the first channel number is the channel number corresponding to the first audio data, the second channel number is the channel number corresponding to the second audio data, the fourth content includes the second sampling rate and the third sampling rate, and the fifth content includes the second quantization bit number and the third quantization bit number;
[0105] The converter 112 is used for at least one of the following: when the first sampling rate and the second sampling rate are different, configuring the first content to the resampling buffer provider 1131; when the first quantization bit number is different from the second quantization bit number, configuring the second content to the reformatting buffer provider 1132; in the case of upmix recording, configuring the third content to the sound effect provider 111 so that the sound effect provider 111 configures the third content to the sound effect unit; when the second sampling rate and the third sampling rate are different, configuring the fourth content to the resampling buffer provider 1131; when the second quantization bit number is different from the third quantization bit number, configuring the fifth content to the reformatting buffer provider 1132.
[0106] In the embodiment of the present application, the sound effect provider 111 is further used to send the following content to the sound effect unit 121: the input buffer corresponding to the first audio data, the output buffer corresponding to the second audio data, and the configuration information of the upmixed sound effect;
[0107] The audio effect unit 121 is further used to: read the first audio data from the input buffer, store the second audio data in the output buffer, and convert the first audio data into the second audio data according to the configuration information of the upmixed audio effect.
[0108] In some embodiments, the configuration information of the upmixed sound effect may include one or more of the following: sound effect algorithm information, sound effect mode information, sound effect parameter information, equalizer setting information, gain setting information, etc.
[0109] In an embodiment of the present application, the recording thread 115 is also used to delete the recording track 116, and the recording track 116 is also used to delete at least one of the following configured to the converter 112: the first content, the second content, the third content, the fourth content, and the fifth content.
[0110] In an embodiment of the present application, the converter 112 is also used for at least one of the following: deleting the first content and / or the fourth content configured to the resampling buffer provider 1131; deleting the second content and / or the fifth content configured to the reformatting buffer provider 1132; deleting the third content configured to the sound effect provider 111, so that the sound effect provider 111 deletes the third content configured to the sound effect unit.
[0111] In the embodiment of the present application, the sound effect provider 111 is further used to: close the interface corresponding to the sound effect unit 121.
[0112] In an embodiment of the present application, the recording thread 115 is also used to send instruction information to the converter 112 for instructing data conversion.
[0113] In the embodiment of the present application, the converter 112 is further used to: send a first buffer to the sound effect provider 111; send a second buffer to the reformatted buffer provider 1132, so that the reformatted buffer provider 1132 sends a third buffer to the resampling buffer provider 1131.
[0114] In an embodiment of the present application, the resampling buffer provider 1131 is also used to: store the fifth audio data and / or the sixth audio data in the third buffer; the reformatting buffer provider 1132 is used to store the first audio data and / or the fourth audio data in the second buffer; the sound effect provider 111 is used to store the second audio data in the first buffer.
[0115] In an embodiment of the present application, the sound effect provider 111 is also used to: send target content to the sound effect unit 121; the target content includes a buffer address corresponding to the output data, an audio frame rate corresponding to the output data, and a data length corresponding to the output data.
[0116] In the embodiment of the present application, the sound effect unit 121 is further used to: send the second audio data to the sound effect provider 111 according to the target content.
[0117] Figure 9a A flowchart of an audio processing method provided in an embodiment of the present application is shown as follows: Figure 9aAs shown, binaural recording data (2 channels, also known as 6 channels) is first obtained, and then spatial audio processing is performed on the binaural recording data to obtain spatial audio data (5.1 channels, also known as 5.1 channels), and then Advanced Audio Coding_High Efficiency (AAC_HE) audio encoding is performed on the spatial audio data to obtain 5.1-channel encoded audio data, and finally, spatial audio or video with spatial audio (i.e., spatial audio video) is obtained based on the 5.1-channel encoded audio data.
[0118] Figure 9b A structural diagram of another audio processing device provided in another embodiment of the present application is shown in FIG. Figure 9b As shown, the audio processing device 10 includes: an audio acquisition module 14, a recording hardware abstraction layer 13, an audio service module 11, an audio effect hardware abstraction layer 12, a media service (MediaServer) module 15 and an application module 16. The audio effect hardware abstraction layer 12 includes an audio effect unit 121. The audio effect unit 121 can be an upmixing effect in other embodiments.
[0119] The audio acquisition module 14 includes a Bluetooth module. The sound effect hardware abstraction layer 12 includes a sound effect unit 121. The audio service module 11 may include: a recording thread 115 (RecordThread), a recording track 116 (RecordTrack), a sound effect provider 111, and a converter 112. The media service module 15 includes an audio recording unit 151. The sound effect provider 111 may include a resampler buffer provider (ResamplerBufferProvider), a reformatter buffer provider (ReformatBufferProvider), and an upmixer buffer provider (UpmixerBufferProvider). In any embodiment of the present application, the provider and the producer may be replaced.
[0120] In any embodiment of the present application, the recording hardware abstraction layer 13 may include a binaural recording hardware abstraction layer 13 (Binaural_Record HAL). In any embodiment of the present application, the sound effect provider 111 may include an upmix sound effect provider 111 (mUpmixEffectProvider). In any embodiment of the present application, the converter 112 may include a recording buffer converter (RecordBufferConvertor). In any embodiment of the present application, the application module 16 may include an application used for audio or video recording, such as a camera application, a video playback application, etc.
[0121] In any embodiment of the present application, the recording buffer converter (RecordBufferConvertor) can be understood in the same way as the reformatting buffer converter (RerormatBufferConverter), or have the same meaning.
[0122] In the implementation process, the Bluetooth module is used to obtain the third audio data, the recording hardware abstraction layer 13 is used to obtain the third audio data from the Bluetooth module, and then provide an interface corresponding to the third audio data. The recording thread 115 obtains the third audio data from the recording hardware abstraction layer 13 according to the interface, and sends the third audio data to the recording track 116. The recording track 116 sends the third audio data to the converter 112. The converter 112 can determine the third audio data as the first audio data, or the converter 112 can call other units (such as buffer providers) to determine the first audio data. The converter 112 sends the first audio data to the sound effect provider 111, so that the sound effect provider 111 sends the first audio data to the sound effect unit 121. The sound effect unit 121 is used to: convert the first audio data into the second audio data, and send the second audio data to the sound effect provider 111. The sound effect provider 111 sends the second audio data to the converter 112. The converter 112 can determine the second audio data as the fourth audio data, or the converter 112 can call other units (such as buffer providers) to convert the second audio data into the fourth audio data. Then the converter 112 sends the fourth audio data to the recording track 116. The recording track 116 sends the fourth audio data to the audio recording unit 151. Exemplarily, the sampling rate of the fourth audio data is 48khz, the quantization bit number is 16bit, and the channel mask = the audio channel in 5.1 (channel mask = AUDIO_CHANNEL_IN_5POINT1). In any embodiment of the present applicant, the channel mask can correspond to the number of channels. The audio recording unit 151 is also used to: perform audio encoding on the fourth audio data, and determine the target space audio data (for example, mp3 format). The application module 16 is used to: obtain the target space audio data from the media service module 15, and generate a target audio file or a target video file corresponding to the target space audio data.
[0123] In some embodiments, the data before being input to the sound effect unit is 2-channel (2ch) data, and the data after being output from the sound effect unit is 6-channel (6ch, corresponding to the above-mentioned 5.1 channels) data.
[0124] Fig.10 A schematic diagram of an audio processing process provided in an embodiment of the present application is shown as follows: Fig.10As shown, the processing process at least involves binaural audio Hal (Binaural Audio Hal, corresponding to the above-mentioned recording hardware abstraction layer), audio service module, upmixing effect Hal and media service module. After the algorithm program (Alg_process) in the binaural stream input (BinauralStreamIn) in the binaural audio Hal obtains the third audio data, it stores the data in the buffer (Buffer) in the binaural stream input. The binaural audio Hal can store the third audio data in anonymous shared memory (Anonymous Shared Memory, ashmem). The stream (StreamlnHalHidl) of the HAL interface definition language (HAL InterfaceDefinition Language, HalHidl) reads the third audio data from ashmem and stores it in the Buffer in StreamlnHalHidl. Among them, Hidl is an interface description language used to specify the interface between HAL and its users. The recording thread reads the third audio data from the Buffer in StreamlnHalHidl and stores it in the resampling input buffer (mRsmpInBuffer) corresponding to the recording thread. The resampling buffer provider is used to read the third audio data from the resampling input buffer corresponding to the recording thread, resample, and obtain the fifth audio data (referred to as the local buffer data (mLocalBufferData) determined by the resampling buffer provider). The reformatting buffer provider determines the first audio data (referred to as the local buffer data (mLocalBufferData) determined by the reformatting buffer provider).
[0125] The effect buffer HalHidi (EffectBufferHalHidi) stores the first audio data to ashmem through the input buffer (minBuffer) in the effect buffer HalHidi. The program (process()) in the upmixing effect (for example, Oplus upmixing effect) in the sound hardware abstraction layer (for example, upmixing effect Hal) determines the second audio data based on the first audio data, and stores the second audio data in ashmem for the effect buffer HalHidi to read and store in the output buffer (mOutBuffer) in the effect buffer HalHidi. The upmixer buffer provider reads the second audio data in mOutBuffer (referred to as mLocalBufferData determined by the upmixer buffer provider) and sends it to the reformatting buffer converter (corresponding to the above-mentioned converter). The reformatting buffer converter (RerormatBufferConverter) determines the fourth audio data based on the second audio data, and stores it in the corresponding buffer (mBuf). The recording tracking reads the fourth audio data (for example, mSink.raw) from the buffer corresponding to the reformatting buffer and stores it in the corresponding buffer (mBuffer). The recording tracking stores the fourth audio data in mBuffer in the control block memory (m Control block Memory, mCblkMemory). The audio recording unit in the media service module reads the fourth audio data from mCblkMemory, determines the target space audio data based on the fourth audio data, and stores the target space audio data in the audio buffer (audioBuffer).
[0126] Fig.11 A schematic diagram of a static structure corresponding to an audio processing flow provided in an embodiment of the present application, such as Fig.11 As shown, it should be noted that, for the solution in the embodiment of the present application, three member variables are added in the recording buffer converter (RecordBufferConverter), namely, request upmix effect (mRequiresUpmixEffect), session identifier (mSessionId), and upmix effect provider (mUpmixEffectProvider).
[0127] mRequiresUpmixEffect determines whether an upmix effect is required. The default is false. When a RecordTrack is created, if the channel mask (channel_mask) is AUDIO_CHANNEL_IN_5POINT1, this variable is set to true. Fig.11In mRequiresUpmixEffect:bool, the bool data type has two values, True and False.
[0128] mSessionId indicates that if upmixing is required, it is bound to this Session Id. Fig.11 In the example, Session Id is audio session t (audio_session_t).
[0129] mUpmixEffectProvider manages at least one of the creation, parameter setting, data processing, and release of upmix effects, and interacts with the upmix effect Hal (for example, Oplus upmix effect Hal) through a HIDL (same meaning as Hidi) interface. Fig.11 In the example, the variable in mUpmixEffectProvider is the passthru buffer provider (PassthruBufferProvider*).
[0130] It should be noted that Fig.11 The other parts of the content can be understood by those skilled in the art based on the common sense in the art, and this application will not elaborate on them.
[0131] Fig.12 A flowchart of a recording channel creation method provided in an embodiment of the present application is shown as follows: Fig.12 As shown, the method includes:
[0132] S1201. The audio manager (AudioFlinger) obtains and creates a recording (createRecord()).
[0133] S1202, check recording thread I (checkRecordThread_I()) and return the result.
[0134] S1203. Create a recording thread I (createRecordTrack_I()).
[0135] S1204, create a new recording track (new RecordTrack()).
[0136] S1205. The recording tracker sends configuration information to the recording buffer.
[0137] The configuration information includes: the first sampling rate (corresponding to thread->mSampleRate), the first quantization bit number (corresponding to thread->mFormat), the first channel mask (corresponding to thread->mChannelMask), the second sampling rate (corresponding to sampleRate), the second quantization bit number (corresponding to format), the second channel mask (corresponding to channelMask), and the session ID (sessionld). Exemplarily, the configuration information may include the following: new RecordBufferConverter (thread->mChannelMask, thread->mFormat, thread->mSampleRate, channelMask, format, sampleRate, sessionld).
[0138] In the embodiment of the present application, the second sampling rate is the sampling rate supported by the sound effect unit and is the sampling rate required for recording. In the embodiment of the present application, the second quantization bit number is the quantization bit number supported by the sound effect unit and is the quantization bit number required for recording.
[0139] S1206. Update parameters (updateParameters()).
[0140] After S1206 , at least one of S1207 , S1208 , and S1209 may be executed.
[0141] S1207: If the first sampling rate is not equal to the second sampling rate, the recording buffer converter configures the first configuration information to the audio resampling, and receives a return result.
[0142] The first sampling rate may be an original sampling rate, and the second sampling rate may be a target sampling rate. In some embodiments, the first configuration information includes the target sampling rate, or includes the original sampling rate and the target sampling rate. In other embodiments, the first configuration information includes at least one of the following: floating point PCM format audio (AUDIO_FORMAT_PCM_FLOAT) (corresponding to the second quantization bit number mentioned above), a target number of channels (mDstChannelCount), and a target sampling rate (mDstSampleRate). For example, if the original sampling rate is not equal to the target sampling rate (if (mSrcSampleRate! = mDstSampleRate) the recording buffer converter configures AudioResampler::create(AUDIO_FORMAT_PCM_FLOAT, mDstChannelCount, mDstSampleRate) to the audio resampling.
[0143] S1208. If a floating point type is requested, the recording buffer converter configures the second configuration information to the reformatting buffer provider and receives a return result.
[0144] The second configuration information includes the second quantization bit number, or includes the first quantization bit number and the second quantization bit number. In some other embodiments, the second configuration information includes at least one of the following: the original channel mask (audio_channel_count_from_in_mask(mSrcChannelMask)), the original format (mSrcFormat, corresponding to the first quantization bit number mentioned above), floating point PCM format audio (AUDIO_FORMAT_PCM_FLOAT, corresponding to the second quantization bit number mentioned above), and receives the return result. For example, if (mRequiresFloat), configure NewReformatBufferProvider
[0145] (audio_channel_count_from_in_mask(mSrcChannelMask),mSrcFormat,AUDIO_FORMAT_PCM_FLOAT).
[0146] S1209: If an upmixing effect is requested, the recording buffer converter configures third configuration information to the upmixer buffer provider, and receives a return result.
[0147] In any embodiment of the present application, the upmixing effect represents spatial audio.
[0148] Wherein, the second configuration information includes a second channel mask, or includes a first channel mask and a second channel mask. In some other embodiments, the third configuration information includes at least one of the following: an original channel mask (mSrcChannelMask, corresponding to the above-mentioned first channel mask), a target channel mask (mDstChannelMask, corresponding to the above-mentioned second channel mask), a floating-point PCM format audio (AUDIO_FORMAT_PCM_FLOAT, corresponding to the above-mentioned second quantization bit number), an original sampling rate (mSrcSampleRate,), and a channel ID (mSessionid). For example, if (mRequiresUpmixEffect), configure newUpmixerBufferProvider (mSrcChannelMask, mDstChannelMask, AUDIO_FORMAT_PCM_FLOAT, mSrcSampleRate, mSessionid).
[0149] After S1209, the recording buffer converter may also be executed to return the result to the recording tracking.
[0150] S1210, the recording track creates a new resampler buffer provider (new ResamplerBufferProvider(this)) and receives the return result.
[0151] After S1210, recording tracking may be performed to return a result to the recording thread, and the recording thread may return a result to the audio manager.
[0152] It should be noted that, in any embodiment of the present application, when audio resampling can be called to implement audio resampling, there is no need for a resampling buffer provider to perform audio resampling processing.
[0153] Fig.13 A flow chart of a method for configuring an upmixer buffer provider provided in an embodiment of the present application is shown in FIG. Fig.13 As shown, the method includes:
[0154] S1301: The upmixer buffer provider receives third configuration information (eg, newUpmixerBufferProvider).
[0155] S1302: The upmixer buffer provider sends a create (create()) to the effects factory Hal interface, and receives a result of returning the effects factory (return mEffectsFactory).
[0156] S1303: The upmixer buffer provider sends a create effect (createEffect()) to the effect factory Hal interface, and receives a result of returning an upmix interface (return mUpmixInterface).
[0157] S1304 . The upmixer buffer provider sends a mirror buffer (mirrorBuffer()) to the effect factory Hal interface, and receives a result of returning an input buffer (return minBuffer).
[0158] S1305 . The upmixer buffer provider sends a mirror buffer (mirrorBuffer()) to the effect factory Hal interface, and receives a result of returning an output buffer (return mOutBuffer).
[0159] S1306. The upmixer buffer provider sends a generated buffer (setinBuffer(mInBuffer)) to the effect Hal interface, and receives a return result.
[0160] S1307. The upmixer buffer provider sends a generated buffer (setinBuffer(mOutBuffer)) to the effect Hal interface and receives a return result.
[0161] S1308. The upmixer buffer provider sends an effect command to establish configuration (command (EFFECT CMD_SET_CONFIG)) to the effect Hal interface, and receives a return result.
[0162] S1309: The upmixer buffer provider sends an effect command enable command (command (EFFECT_CMD_ENABLE)) to the effect Hal interface, and receives a return result.
[0163] Fig.14 A flowchart of a recording channel release method provided in an embodiment of the present application is shown as follows: Fig.14 As shown, the method includes:
[0164] S1401, the recording track receives the information of deleting the recording track (~RecordTrack()).
[0165] S1402, recording tracking deletes the recording buffer converter (delete mRecordBufferConverter).
[0166] S1403, recording tracking delete resampler (delete mResampler), and receiving the return result.
[0167] S1404, the recording track deletes the reformatting buffer provider (delete ReformatBufferProvider, also called delete minputConverterProvider), and receives the return result.
[0168] S1405 , the recording track deletes the upmixer buffer provider (delete UpmixerBufferProvider, also called delete mUpmixEffectProvider), and receives the returned result.
[0169] S1406: The upmixer buffer provider closes the effect Hal interface (also called m upmix interface) and receives a return result.
[0170] After S1406, the recording buffer converter returns the result to the recording tracking.
[0171] S1407, recording tracking deletes the resampler buffer provider (ResamplerBufferProvider, also called deletemRecordBufferConverter), and receives the return result.
[0172] Fig.15 A flowchart of a recording path data processing method provided in an embodiment of the present application is shown as follows: Fig.15 As shown, the method includes:
[0173] S1501. The recording thread obtains the thread queue (threadLoop()).
[0174] S1502. The recording thread reads the source data and receives the returned result.
[0175] The source data may correspond to the third audio data. In some embodiments, the source data may be read by reading ((uint8_t*)mRsmpinBuffer+rear*mFrameSize, mBufferSize, & bytesRead).
[0176] S1503. The recording thread sends a notification message to the resampling buffer provider and receives a return result.
[0177] The notification message is used to indicate the arrival of data. In some embodiments, the notification message may include sync(&framesin,&hasOverrun).
[0178] S1504. The recording thread sends a conversion notification to the recording buffer converter.
[0179] In some embodiments, the conversion notification may include convert(activeTrack->mSink.raw, activeTrack->mResamplerBufferProvider, framesOut). In some embodiments, the conversion notification is used to indicate data arrival.
[0180] After S1504, S1505 or S1513 may be executed.
[0181] S1505 . If the resampler is empty (if (mResampler == NULL)), the recording buffer converter sends a request to the upmixer buffer provider to obtain the next buffer.
[0182] For example, the recording buffer converter provides provider->getNextBuffer&buffer) to the upmixer buffer provider.
[0183] S1506: The upmixer buffer provider sends a request to obtain the next buffer to the reformatting buffer provider.
[0184] For example, the upmixer buffer provider sends mTrackBufferProvider->getNextBuffer(pBuffer) to the reformatter buffer provider.
[0185] S1507: The reformatting buffer provider sends a request to the resampling buffer provider to obtain the next buffer.
[0186] For example, the reformatting buffer provider sends mTrackBufferFrovider->getNextBuffer(pBuffer) to the resampling buffer provider.
[0187] S1508. The resampling buffer provider copies the audio frames (copyFrames()), and returns the result (corresponding to the fifth audio data mentioned above) to the reformatting buffer provider.
[0188] S1509: The reformatting buffer provider copies the audio frames (copyFrames()), and returns the result (corresponding to the first audio data mentioned above) to the upmixer buffer provider.
[0189] S1510 : The upmixer buffer provider performs conversion without resampling (convertNoResampler()) and performs storage copy by audio format (memcpy_by_audio_format()).
[0190] In some embodiments, before S1510, the upmixer buffer provider also performs an action of copying the audio frame (the copied audio frame corresponds to the second audio data mentioned above). In some embodiments, the upmixer buffer provider performing the conversion without resampling may include: the upmixer buffer provider determines the second audio data as the fourth audio data. In some embodiments, the upmixer buffer provider performing the storage copying by audio format may include: the upmixer buffer provider stores the fourth audio data in the corresponding buffer.
[0191] S1511. The upmixer buffer provider sends a release buffer to the reformatting buffer provider (eg, mTrackBufferProvider->releaseBuffer(&mBuffer)).
[0192] S1512. The reformatting buffer provider sends a release buffer (eg, mTrackBufferProvider->releaseBuffer(&mBuffer)) to the resampling buffer provider, and receives a return result.
[0193] After S1512, the reformatting buffer provider may return a result to the upmixer buffer provider, and the upmixer buffer provider may return a result to the recording buffer converter.
[0194] S1513. If the resampler is not empty (if (mResampler!=NULL)), the recording buffer converter sends the resampling information to the audio resampler.
[0195] In some embodiments, the resampling information may be resample((int32_t)mBuf,frames,provider).
[0196] In some embodiments, resampling is empty, indicating that there is no audio resampling that can be called, so that the audio resampling does not need to perform a corresponding action. In some embodiments, resampling is not empty, indicating that there is an audio resampling that can be called, so that the audio resampling needs to perform a corresponding action.
[0197] S1514: The audio resampler is sent to the upmixer buffer provider to obtain the next buffer.
[0198] For example, the audio resampler sends provider->getNextBuffer(&mBuffenR) to the upmixer buffer provider.
[0199] S1515. The upmixer buffer provider sends a request to obtain the next buffer to the reformatter buffer provider.
[0200] For example, the upmixer buffer provider sends mTrackBufferProvider->getNextBuffer(pBuffer) to the reformatter buffer provider.
[0201] S1516. The reformatting buffer provider sends a request to the resampling buffer provider to obtain the next buffer.
[0202] For example, the reformatting buffer provider sends mTrackBufferProvider->getNextBuffer(pBuffer) to the resampling buffer provider.
[0203] S1517. The resampling buffer provider copies the audio frames (copyFrames()) and returns the result (corresponding to the fifth audio data mentioned above) to the reformatting buffer provider.
[0204] S1518: The reformatting buffer provider copies the audio frames (copyFrames()), and returns the result (corresponding to the first audio data mentioned above) to the upmixer buffer provider.
[0205] S1519. The upmixer buffer provider copies the audio frames (copyFrames()) and returns the result (corresponding to the second audio data mentioned above) to the audio resampler.
[0206] S1520 , the audio is resampled and a resample process (resampleProcess()) is performed to obtain fourth audio data.
[0207] S1521 . The audio resampler sends a released buffer to the upmixer buffer provider (eg, provider->getNextBuffer(&mBuffer)).
[0208] S1522: The upmixer buffer provider sends a release buffer to the reformatting buffer provider (eg, mTrackBufferProvider->releaseBuffer(&mBuffer)).
[0209] S1523. The reformatting buffer provider sends a release buffer (eg, mTrackBufferProvider->releaseBuffer(&mBuffer)) to the resampling buffer provider, and receives a return result.
[0210] After S1523, the reformatting buffer provider may return a result to the upmixer buffer provider, the upmixer buffer provider may return a result to the audio resampler, and the audio resampler may return a result to the recording buffer converter.
[0211] S1524, the recording buffer converter performs conversion without resampling (convertNoResampler()), and performs storage copy by audio format (memcpy_by_audio_format()).
[0212] S1525. The recording buffer converter resets (reset()) the upmixer buffer provider and receives a return result.
[0213] S1526. The recording buffer converter resets (reset()) the reformatting buffer provider and receives the return result.
[0214] Fig.16 A flow chart of a data processing method of an upmixer buffer provider provided in an embodiment of the present application is shown in FIG. Fig.16 As shown, the method includes:
[0215] S1601: The upmixer buffer provider obtains copy audio frames (copyFrames).
[0216] In some embodiments, the format of copying audio frames can be copyFrames(void*dst, const void*src, size_tframes).
[0217] S1602. The upmixer buffer provider sets external data (source data) (setExternalData(src)) to the effect factory Hal interface, and receives a return result.
[0218] In some embodiments, the external data (source data) may correspond to an address of an input buffer.
[0219] S1603. The upmixer buffer provider sets the number of frames (setFrameCount(frames)) to the effect factory Hal interface and receives the return result.
[0220] In some embodiments, the frame number in S1603 may be the frame number corresponding to the source data.
[0221] S1604. The upmixer buffer provider updates the input frame length (update(minFrameSize*frames)) to the effect factory Hal interface, and receives the return result.
[0222] In some embodiments, the frame length in S1604 may be the frame length corresponding to the source data.
[0223] S1605. The upmixer buffer provider sets external data (destination data) (setExternalData(dst)) to the effect Hal interface, and receives a return result.
[0224] In some embodiments, the target data in S1605 may be data returned by the effect Hal interface.
[0225] S1606. The upmixer buffer provider sets the number of frames (setFrameCount(frames)) to the effect Hal interface and receives the return result.
[0226] In some embodiments, the frame number in S1606 may be the frame number corresponding to the target data.
[0227] S1607. If the source data is not equal to the target data (if (dst!=src)), the upmixer buffer provider updates the output frame length (update (mOutFrameSize*frames)) to the effect Hal interface, and receives the return result.
[0228] S1608. The upmixer buffer provider sends a process (process()) to the effect factory Hal interface and receives a return result.
[0229] S1609. The upmixer buffer provider submits the output frame length (update(mOutFrameSize*frames)) to the effect Hal interface, and receives the return result.
[0230] Through S1609, the effect Hal interface outputs data according to the output frame length.
[0231] Fig.17 A schematic diagram of the file location corresponding to an audio processing flow provided in an embodiment of the present application, such as Fig.17 As shown, the file locations corresponding to the audio processing flow include the following: / system / lib64 / libaudioprocessing.so, / system / lib64 / libaudioflinger.so, and / system / lib64 / soundfx / liboplusupmixeffect.so.
[0232] Among them, / system / lib64 / libaudioprocessing.so corresponds to the recording thread, recording tracking, recording buffer converter, and the part where the upmix effect provider interacts with the recording buffer converter in the audio service module. / system / lib64 / libaudioflinger.so corresponds to the part where the upmix effect provider interacts with the upmix effect. / system / lib64 / soundfx / liboplusupmixeffect.so corresponds to the upmix effect.
[0233] This application adds the option of spatial audio recording mode during shooting, which can generate 5.1-channel video, and cooperate with the dynamic playback of spatial audio during playback to achieve an immersive effect of turning your head, with integrated recording and playback. A typical scenario is a travel video log (Video LOG, VLOG): recording the human sounds and the sounds of nature in the city during the trip, and obtaining a more immersive video experience during playback. In this way, the recording effect is more obvious in scenes with sound movement trajectories. For example, in some embodiments, the user can open the camera application and select video shooting. The indicator control for spatial audio recording can be displayed on the video shooting interface. When the user triggers the indicator control, or triggers the indicator control and triggers the shooting button, the spatial video is recorded, so that the audio service module is used to obtain the third audio data from the recording hardware abstraction layer.
[0234] This application adds upmix (2 to 6) sound effects in the RecordBufferConverter of AudioServer. The sound effect library is easy to expand, and both spatial audio recording and playback are implemented using the sound effect library. This architecture design is more symmetrical. Various Audio HAL modules (corresponding to the above-mentioned recording hardware abstraction layer) are universal and have no specific dependencies.
[0235] This application adds a sound effect library processing (implemented by a sound effect unit) for upmixing and spatial rendering in the recording path of AudioFlinger, complies with the sound effect design architecture of Android, and has good scalability. This solution is completely decoupled from the other functions of Audio HAL and AudioFlinger, and the sound effect library (for example, the sound effect unit) is replaceable. From the verification results, both binaural recording and mobile phone MIC recording can record smooth 5.1-channel audio, and no performance problems are found to cause freezes, which is in line with design expectations.
[0236] In other embodiments of the present application, channel conversion and spatial rendering can also be performed in the post-processing of the APP recording two-channel audio, but it is only valid within a single application and the scope of application is too small. For example, the sound effect unit is connected to the recording hardware abstraction layer, and channel conversion and spatial rendering can also be performed in the pre-processing of the Audio HAL (for example, after the recording hardware abstraction layer obtains the third audio data, the first audio data is determined according to the third audio data, and the sound effect conversion is performed through the sound effect unit to obtain the second audio data, and then the recording hardware abstraction layer determines the fourth audio data according to the second audio data). In this way, AudioFlinger can directly record 5.1-channel audio data, but it is only applicable to a single Audio HAL.
[0237] The present application provides a method for a system to record spatial audio for a terminal device. First, the audio manager AudioFlinger records the two-channel audio of the binaural mic (microphone) through headphones (wired or wireless), and then adds local sound effect processing to convert the two-channel into 5.1 channels (5.1 channels refer to the central channel, front left and right channels, rear left and right surround channels, and the so-called 0.1 channel subwoofer channel. A system can connect a total of 6 speakers), and perform spatial algorithm rendering. Finally, the upper-level application obtains the processed spatial audio through the audio recording (AudioRecord) or media recording (MediaRecord) interface, encodes it with audio formats such as AAC_HE, and generates an audio file or video file.
[0238] Fig.18 A schematic diagram of the structure of a terminal device provided in an embodiment of the present application is shown in FIG. Fig.18 As shown, the terminal device 20 includes the audio processing device 10 in any of the above embodiments. In some embodiments, the terminal device 20 may also include one or more of the following: a display device, a communication device, an audio playback device, and the like.
[0239] Each module, each unit, each hardware abstraction layer or processor in the above-mentioned audio processing device may include any one or more of the following integrations: a general-purpose processor, an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a graphics processing unit (GPU), an embedded neural network processor (neural-network processing units, NPU), a controller, a microcontroller, a microprocessor, a programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. It can be understood that the electronic device that implements the above-mentioned processor function can also be other, and the embodiments of the present application are not specifically limited. Each module, each unit, each hardware abstraction layer or processor in the audio processing device can implement or execute the disclosed methods, steps and logic block diagrams in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in a memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.
[0240] It is understood that the cache, memory or computer storage medium in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM) or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus random access memory (DR RAM). It should be noted that the memory of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0241] It should be understood that the "one embodiment" or "an embodiment" or "an embodiment of the present application" or "the aforementioned embodiment" or "some implementations" or "some embodiments" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" or "an embodiment of the present application" or "the aforementioned embodiment" or "some implementations" or "some embodiments" appearing throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the size of the sequence number of the above-mentioned processes does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present application. The above-mentioned sequence numbers of the embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.
[0242] The terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this application, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0243] In the description of this application, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection, an electrical connection, or mutual communication; it can be a direct connection, or an indirect connection through an intermediate medium, it can be the internal connection of two elements or the interaction relationship between two elements. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0244] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0245] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0246] In addition, all functional units in the embodiments of the present application may be integrated into one processing unit, or each unit may be a separate unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0247] The features disclosed in several product embodiments provided in this application can be combined arbitrarily without conflict to obtain new product embodiments. The features disclosed in the device embodiments provided in this application can be combined arbitrarily without conflict to obtain new device embodiments.
[0248] If the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer storage medium. Based on such an understanding, the technical solution of the embodiment of the present application can be essentially or partly embodied in the form of a software product that contributes to the relevant technology. The computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.
[0249] In the embodiments of the present application, the descriptions of the same steps and the same contents in different embodiments can refer to each other. In the embodiments of the present application, the term "and" does not affect the order of the steps. For example, the terminal device executes A and executes B, which means that the terminal device executes A first and then executes B, or the terminal device executes B first and then executes A, or the terminal device executes A and executes B at the same time.
[0250] It is worth noting that the drawings in the embodiments of the present application are only for illustrating the schematic positions of various components on the terminal device and do not represent the actual positions in the terminal device. The actual positions of various components or areas may be changed or offset accordingly according to actual conditions (for example, the structure of the terminal device), and the proportions of different parts of the terminal device in the drawings do not represent the actual proportions.
[0251] As used in the embodiments of the present application and the appended claims, the singular forms "a," "an," "said," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.
[0252] It should be understood that the term "and / or" used in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.
[0253] It should be noted that in each embodiment involved in the present application, all steps may be executed or part of the steps may be executed as long as a complete technical solution can be formed.
[0254] The above is only an implementation method of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. An audio processing device, characterized in that: The audio processing device comprises: an audio service module and an audio effect hardware abstraction layer, wherein the audio effect hardware abstraction layer comprises an audio effect unit; The audio service module is used to: determine first audio data, and send the first audio data to the sound effect unit; The audio effect unit is used to: convert the first audio data into second audio data, and send the second audio data to the audio service module; wherein the number of channels corresponding to the first audio data is smaller than the number of channels corresponding to the second audio data.
2. The audio processing device according to claim 1, characterized in that The audio processing device further comprises: a recording hardware abstraction layer; The audio service module is further used to: obtain third audio data from the recording hardware abstraction layer, determine the first audio data according to the third audio data; determine fourth audio data according to the second audio data; The number of channels corresponding to the first audio data is the same as the number of channels corresponding to the third audio data, and the number of channels corresponding to the second audio data is the same as the number of channels corresponding to the fourth audio data.
3. The audio processing device according to claim 2, characterized in that: The audio processing device further includes: a media service module and an application module; The audio service module is further used to: send the fourth audio data to the media service module; The media service module is used to: perform audio encoding on the fourth audio data to determine target spatial audio data; The application module is used to obtain the target spatial audio data from the media service module and generate a target audio file or a target video file corresponding to the target spatial audio data.
4. The audio processing device according to claim 2 or 3, characterized in that: The audio service module includes a sound effect provider and a converter; The converter is used to: obtain the first audio data, and send the first audio data to the sound effect provider; The sound effect provider is used to: receive the first audio data sent by the converter, and send the first audio data to the sound effect unit; receive the second audio data sent by the sound effect unit, and send the second audio data to the converter; The converter is used for receiving the second audio data sent by the sound effect provider.
5. The audio processing device according to claim 4, characterized in that: The audio service module also includes a conversion provider; the converter is used to: Acquire the third audio data; when the first sampling rate of the third audio data is different from the second sampling rate supported by the sound effect unit and / or the first quantization bit number is different from the second quantization bit number supported by the sound effect unit, call the conversion provider to resample the third audio data and / or convert the quantization bit number to obtain the first audio data of the second sampling rate and / or the second quantization bit number; or, when the first sampling rate of the third audio data is the same as the second sampling rate supported by the sound effect unit, and the first quantization bit number is the same as the second quantization bit number supported by the sound effect unit, determine the third audio data as the first audio data; and / or, When the second sampling rate of the second audio data is different from the third sampling rate required for recording and / or the second quantization bit number is different from the third quantization bit number required for recording, the conversion provider is called to resample the second audio data and / or convert the quantization bit number to obtain the fourth audio data of the third sampling rate and / or the third quantization bit number; or, when the second sampling rate of the second audio data is the same as the third sampling rate required for recording, and the second quantization bit number is the same as the third quantization bit number required for recording, the second audio data is determined as the fourth audio data.
6. The audio processing device according to claim 5, characterized in that: The audio service module also includes: an audio manager; The audio manager is used to create a recording thread, the recording thread is used to create a recording track configured corresponding to the recording thread, and the recording track is used to configure at least one of the following to the converter: a first content, a second content, a third content, a fourth content, and a fifth content; the first content includes the first sampling rate and the second sampling rate, the second content includes the first quantization bit number and the second quantization bit number, the third content includes the first channel number and the second channel number, the first channel number is the channel number corresponding to the first audio data, the second channel number is the channel number corresponding to the second audio data, the fourth content includes the second sampling rate and the third sampling rate, and the fifth content includes the second quantization bit number and the third quantization bit number; The converter is used for at least one of the following: when the first sampling rate and the second sampling rate are different, configuring the first content to the conversion provider; when the first quantization bit number is different from the second quantization bit number, configuring the second content to the conversion provider; in the case of upmix recording, configuring the third content to the sound effect provider so that the sound effect provider configures the third content to the sound effect unit; when the second sampling rate and the third sampling rate are different, configuring the fourth content to the conversion provider; when the second quantization bit number is different from the third quantization bit number, configuring the fifth content to the conversion provider.
7. The audio processing device according to claim 5, characterized in that: The conversion provider includes: a resampling buffer provider and a reformatting buffer provider; The resampling buffer provider is used to: receive the third audio data sent by the converter; if the first sampling rate of the third audio data is different from the second sampling rate, resample the third audio data to obtain fifth audio data, or if the first sampling rate of the third audio data is the same as the second sampling rate, determine the third audio data as fifth audio data; and send the fifth audio data to the reformatting buffer provider; The reformatting buffer provider is used to: receive the fifth audio data sent by the resampling buffer provider; when the first quantization bit number of the fifth audio data is different from the second quantization bit number, convert the quantization bit number of the fifth audio data to obtain the first audio data of the second quantization bit number, or when the first quantization bit number of the fifth audio data is the same as the second quantization bit number, determine the fifth audio data as the first audio data; and send the first audio data to the converter.
8. The audio processing device according to claim 5, characterized in that: The conversion provider includes: a resampling buffer provider and a reformatting buffer provider; The resampling buffer provider is used to: receive the second audio data sent by the converter; if the second sampling rate of the second audio data is different from the third sampling rate, resample the second audio data to obtain sixth audio data, or if the second sampling rate of the second audio data is the same as the third sampling rate, determine the second audio data as sixth audio data; and send the sixth audio data to the reformatting buffer provider; The reformatting buffer provider is used to: receive the sixth audio data sent by the resampling buffer provider; when the second quantization bit number of the sixth audio data is different from the third quantization bit number, convert the quantization bit number of the sixth audio data to obtain the fourth audio data of the third quantization bit number, or when the second quantization bit number of the sixth audio data is the same as the third quantization bit number, determine the sixth audio data as the fourth audio data; and send the fourth audio data to the converter.
9. The audio processing device according to claim 7 or 8, characterized in that: The audio service module also includes: an audio manager; The audio manager is used to create a recording thread, the recording thread is used to create a recording track configured corresponding to the recording thread, and the recording track is used to configure at least one of the following to the converter: a first content, a second content, a third content, a fourth content, and a fifth content; the first content includes the first sampling rate and the second sampling rate, the second content includes the first quantization bit number and the second quantization bit number, the third content includes the first channel number and the second channel number, the first channel number is the channel number corresponding to the first audio data, the second channel number is the channel number corresponding to the second audio data, the fourth content includes the second sampling rate and the third sampling rate, and the fifth content includes the second quantization bit number and the third quantization bit number; The converter is used for at least one of the following: when the first sampling rate and the second sampling rate are different, configuring the first content to the resampling buffer provider; when the first quantization bit number is different from the second quantization bit number, configuring the second content to the reformatting buffer provider; in the case of upmix recording, configuring the third content to the sound effect provider so that the sound effect provider configures the third content to the sound effect unit; when the second sampling rate and the third sampling rate are different, configuring the fourth content to the resampling buffer provider; when the second quantization bit number is different from the third quantization bit number, configuring the fifth content to the reformatting buffer provider.
10. The audio processing device according to claim 4, characterized in that: The sound effect provider is further used to send the following content to the sound effect unit: an input buffer corresponding to the first audio data, an output buffer corresponding to the second audio data, and configuration information of the upmixed sound effect; The sound effect unit is further used to: read the first audio data from the input buffer, store the second audio data in the output buffer, and convert the first audio data into the second audio data according to the configuration information of the upmixed sound effect.
11. The audio processing device according to claim 6 or 9, characterized in that: The recording thread is further used to delete the recording track, and the recording track is further used to delete at least one of the following configured to the converter: the first content, the second content, the third content, the fourth content, and the fifth content; The converter is further configured to at least one of the following: delete the first content and / or the fourth content configured to the resampling buffer provider; delete the second content and / or the fifth content configured to the reformatting buffer provider; delete the third content configured to the sound effect provider, so that the sound effect provider deletes the third content configured to the sound effect unit; The sound effect provider is further used to: close the interface corresponding to the sound effect unit.
12. The audio processing device according to claim 6 or 9, characterized in that: The recording thread is also used to send instruction information for instructing the converter to convert the data; The converter is further used to: send a first buffer to the sound effect provider; send a second buffer to the reformatted buffer provider, so that the reformatted buffer provider sends a third buffer to the resampled buffer provider; The resampling buffer provider is further used to: store the fifth audio data and / or the sixth audio data into the third buffer; The reformatting buffer provider is used to store the first audio data and / or the fourth audio data into the second buffer; the sound effect provider is used to store the second audio data into the first buffer; The sound effect provider is further used to: send target content to the sound effect unit; the target content includes a buffer address corresponding to the output data, an audio frame rate corresponding to the output data, and a data length corresponding to the output data; The sound effect unit is further used to send the second audio data to the sound effect provider according to the target content.
13. A terminal device, characterized in that: The audio processing device comprises the audio processing device according to any one of claims 1 to 12.
Citation Information
Cited By
Audio data stream processing device and method and electronic equipment
CN121478221A