Audio processing method, device, terminal equipment and storage medium
By adjusting the audio gain according to the zoom range and noise level during video recording, the poor audio quality caused by excessive noise in noisy environments is solved, and the auditory effect and quality are improved while amplifying the audio.
Patent Information
- Application Number
- CN202111073008.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-14
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2041-09-14
AI Technical Summary
When recording maximum gain in noisy environments, excessive noise is amplified to cause poor audio quality.
By recording a video clip, the audio gain is determined based on the zoom range of the video clip and the noise level of the initial audio clip, and the audio clip is adjusted based on that gain, amplifying the audio moderately to reduce the negative impact of noise amplification on audio quality.
While amplifying the audio, moderately amplify the ambient noise, improve the auditory effect and quality of the audio, ensure that the audio gain matches the noise level, and avoid excessive gain changes affecting the auditory experience.
Smart Images

Figure CN115811591B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of audio processing technology, and in particular to an audio processing method, apparatus, terminal device, and storage medium. Background Art
[0002] With the rapid development of terminal technology, audio and video recording has become an important application in terminal devices such as mobile phones and tablets, and users' requirements for audio effects in videos are getting higher and higher.
[0003] Currently, to create a closer-up effect, the audio volume is sometimes increased as the zoom factor increases during video recording. As the zoom factor increases, the audio volume is amplified using the gain control, allowing the audio volume to increase with the zoom factor.
[0004] However, the relationship between audio gain and video zoom range is fixed: the greater the zoom, the louder the recording volume. To maximize the amplification effect, the maximum audio gain can exceed 12dB. Therefore, if you record at maximum gain in a very noisy place, the noise in the recording file will be very impactful (noise is also amplified at maximum gain), resulting in poor audio quality. Summary of the Invention
[0005] The embodiments of the present application provide an audio processing method, apparatus, terminal device, and storage medium to solve the problem of excessive noise amplification resulting in poor audio quality when performing maximum gain recording in a noisy environment.
[0006] In a first aspect of an embodiment of the present application, an audio processing method is provided, which includes: recording a first video clip, the first video clip including a first initial audio clip; determining a first gain based on a first zoom range of the first video clip and a first noise level of the first initial audio clip; and adjusting the first initial audio clip based on the first gain to obtain a first audio clip.
[0007] According to a second aspect of an embodiment of the present application, an audio processing device is provided, which includes: a recording module, a determination module and an adjustment module; the recording module is used to record a first video clip, which includes a first initial audio clip; the determination module is used to determine a first gain based on a first zoom range of the first video clip recorded by the recording module and a first noise level of the first initial audio clip recorded by the recording module; the adjustment module is used to adjust the first initial audio clip based on the first gain determined by the determination module to obtain a first audio clip.
[0008] According to a third aspect of an embodiment of the present application, a terminal device is provided, which includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the audio processing method described in the first aspect are implemented.
[0009] According to a fourth aspect of an embodiment of the present application, a readable storage medium is provided, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the audio processing method described in the first aspect are implemented.
[0010] In a fifth aspect of an embodiment of the present application, a chip is provided, which includes a processor and a communication interface, the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the audio processing method described in the first aspect.
[0011] In an embodiment of the present application, a first video clip can be recorded, the first video clip including a first initial audio clip; a first gain can be determined based on a first zoom range of the first video clip and a first noise level of the first initial audio clip; and the first initial audio clip can be adjusted based on the first gain to obtain a first audio clip. In this solution, during the video recording process, the audio gain (hereinafter referred to as audio gain) can be determined based on the zoom range of the video clip and the noise level of the initial audio clip in the video clip. In this way, the audio gain can be determined based on the zoom range and the noise level, thereby obtaining a gain that moderately amplifies the audio based on the ambient noise level. While amplifying the audio, the ambient noise can also be moderately amplified, thereby improving the auditory effect of the audio and the quality of the audio. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments and the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application, and other drawings can be obtained based on these drawings.
[0013] Figure 1 A schematic diagram of the architecture of a possible Android operating system provided in an embodiment of the present application;
[0014] Figure 2 This is a flowchart of an audio processing method according to an embodiment of the present application;
[0015] Figure 3 The second flowchart of the audio processing method provided in the embodiment of the present application;
[0016] Figure 4 The third flowchart of the audio processing method provided in the embodiment of the present application;
[0017] Figure 5 Flowchart 4 of the audio processing method provided in the embodiment of the present application;
[0018] Figure 6 A structural block diagram of an audio processing device provided in an embodiment of the present application;
[0019] Figure 7 A schematic diagram of the hardware structure of a terminal device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0020] The following will be combined with the accompanying drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.
[0021] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.
[0022] First, some nouns or terms involved in the claims and description of the present invention are explained below.
[0023] Typically, when recording video on a terminal device, the recording volume increases with increasing zoom (and vice versa). This is becoming a popular practice in the industry. The current method for controlling the audio volume during video recording is to establish a relationship between audio gain and zoom range, then adjust the audio gain accordingly as the zoom factor changes. As shown in Table 1, a larger zoom factor generally results in a greater gain, which creates a closer-up effect.
[0024] Table 1
[0025]
[0026] As shown in Table 1, the relationship between audio gain and video zoom range is fixed: the greater the zoom, the louder the audio volume. To maximize the amplification effect, the maximum audio gain can exceed 12dB. However, this can lead to a problem: if you record at the maximum zoom in a very noisy place, the noise in the recorded file will be very noticeable.
[0027] In order to solve the above technical problems, in an embodiment of the present application, a first video clip can be recorded, the first video clip including a first initial audio clip; a first gain is determined based on a first zoom range of the first video clip and a first noise level of the first initial audio clip; and the first initial audio clip is adjusted based on the first gain to obtain a first audio clip. In this solution, during the video recording process, the audio gain (hereinafter referred to as audio gain) is determined based on the zoom range of the video clip and the noise level of the initial audio clip in the video clip, so that the audio gain can be jointly determined based on the zoom range and the noise level. In this way, a gain for moderately amplifying the audio determined based on the ambient noise level can be obtained. While amplifying the audio, the ambient noise is moderately amplified, thereby improving the auditory effect of the audio and improving the quality of the audio.
[0028] The terminal device in the embodiment of the present invention may be a terminal device having an operating system. The operating system may be an Android operating system, an iOS operating system, or a Hongmeng operating system, or may be other possible operating systems, which are not specifically limited in the embodiment of the present invention.
[0029] The following uses the Android operating system as an example to introduce the software environment used by the audio processing method provided in the embodiment of the present invention.
[0030] like Figure 1 As shown in FIG, a possible architecture diagram of an Android operating system provided by an embodiment of the present invention. Figure 1 The architecture of the Android operating system includes four layers: application layer, application framework layer, system runtime layer and kernel layer (specifically, the Linux kernel layer).
[0031] The application layer includes various applications in the Android operating system (including system applications and third-party applications).
[0032] The application framework layer is the framework of the application. Developers can develop some applications based on the application framework layer while complying with the development principles of the application framework.
[0033] The system runtime layer includes libraries (also called system libraries) and the Android operating system runtime environment. Libraries primarily provide the various resources required by the Android operating system. The Android operating system runtime environment provides the software environment for the Android operating system.
[0034] The kernel layer is the operating system layer of the Android operating system and is the lowest level of the Android operating system software hierarchy. Based on the Linux kernel, the kernel layer provides core system services and hardware-related drivers for the Android operating system.
[0035] Taking the Android operating system as an example, in the embodiment of the present invention, developers can Figure 1 The system architecture of the Android operating system shown in FIG. 1 is used to develop a software program for implementing the audio processing method provided by the embodiment of the present invention, so that the audio processing method can be based on the following example: Figure 1 The Android operating system shown is running. That is, the processor or terminal device can implement the audio processing method provided by the embodiment of the present invention by running the software program in the Android operating system.
[0036] The terminal devices in the embodiments of the present application may be mobile terminal devices or non-mobile terminal devices. Mobile terminal devices may be mobile phones, tablet computers, laptop computers, PDAs, in-vehicle terminal devices, wearable devices, ultra-mobile personal computers (UMPCs), netbooks, or personal digital assistants (PDAs); non-mobile terminal devices may be personal computers (PCs), televisions (TVs), ATMs, or self-service kiosks, etc.; these are not specifically limited in the embodiments of the present application.
[0037] The execution subject of the audio processing method provided in the embodiment of the present application can be the above-mentioned terminal device (including mobile terminal devices and non-mobile terminal devices), or it can be a functional module and / or functional entity in the terminal device that can implement the audio processing method. The specific execution subject can be determined according to actual usage requirements and is not limited by the embodiment of the present application.
[0038] The audio processing method provided in the embodiment of the present application is described in detail below through specific embodiments and their application scenarios in conjunction with the accompanying drawings.
[0039] like Figure 2 As shown, the embodiment of the present application provides an audio processing method. The following takes the execution subject as a terminal device as an example to exemplify the audio processing method provided by the embodiment of the present application. The method may include the following steps 201 to 203.
[0040] 201. The terminal device records a first video clip.
[0041] The first video segment includes a first initial audio segment.
[0042] Among them, the first video clip includes a first initial audio clip and N frames of video images. The terminal device records the N frames of video images through a camera and records the first initial audio clip through a microphone. The first initial audio clip is an audio clip that has not undergone any audio processing.
[0043] It can be understood that the first video clip is any video clip in the video recording process. The terminal device processes the initial audio clip in each video clip in the video recording process in the same way as it processes the first initial audio clip. For details, please refer to the following description.
[0044] 202. The terminal device determines a first gain according to a first zoom range of the first video clip and a first noise level of the first initial audio clip.
[0045] The first zoom range is the zoom range of the camera when recording the first video clip.
[0046] The first noise level (NoiseLevel) is the noise level of the ambient noise in the first initial audio clip. The noise level is obtained by classifying noise according to noise loudness, that is, different noise levels correspond to different noise loudness ranges.
[0047] Optionally, the terminal device may determine the first gain based on the first zoom range, the first noise level, and a first list, wherein the first list is a mapping relationship table between the zoom range, the noise level and the gain; the terminal device may also determine the first gain based on the first zoom range, the first noise level, and a first function, wherein the first function is a mapping function between the zoom range, the noise level and the gain; the terminal device may also determine the first gain based on the first zoom range and the first noise level in other feasible ways, which is not limited in the embodiments of the present application.
[0048] 203. The terminal device adjusts the first initial audio segment based on the first gain to obtain a first audio segment.
[0049] It can be understood that the terminal device can adjust the first initial audio segment according to the first gain to obtain the first audio segment; it can also obtain other gains based on the first gain, and then adjust the first initial audio segment according to the other gains to obtain the first audio segment; it can also adjust the first initial audio segment based on the first gain through other feasible means to obtain the first audio segment. The specific method can be determined based on actual usage requirements and is not limited in the embodiments of the present application.
[0050] It can be understood that after obtaining the first audio segment, the terminal device refers to the first audio segment and N frames of video images as a new first video segment, and the terminal device can also play the new first video images.
[0051] In an embodiment of the present application, during the process of recording a video, the audio gain (hereinafter referred to as audio gain) is determined based on the zoom range of the video clip and the noise level of the initial audio clip in the video clip. In this way, the audio gain can be jointly determined based on the zoom range and the noise level, so that a gain for moderately amplifying the audio determined according to the ambient noise level can be obtained. While amplifying the audio, the ambient noise is moderately amplified to improve the auditory effect of the audio.
[0052] Optionally, the above step 203 can be specifically implemented through the following step 203a.
[0053] 203a. When the absolute value of the difference between the first gain and the second gain is less than or equal to the gain threshold, the terminal device adjusts the first initial audio segment according to the first gain to obtain a first audio segment.
[0054] The second gain is an adjustment gain corresponding to a second initial audio segment, and the second initial audio segment belongs to a second video segment recorded before the first video segment.
[0055] The gain threshold can be determined according to actual use requirements and is not limited in the present embodiment. The second gain can be greater than the first gain, or less than the first gain, or equal to the first gain, and is not limited in the present embodiment.
[0056] The second video segment may be a video segment adjacent to the first video segment, or may be a video segment spaced a certain distance apart from the first video segment, which is not limited in this embodiment of the present application.
[0057] Optionally, the adjustment gain can be a gain actually used to adjust the loudness of the second initial audio segment, or it can be determined based on the second zoom range (of the second video segment to which the second initial audio segment belongs) and the second noise level (of the second initial audio segment). The specific adjustment gain can be determined based on actual usage requirements and is not limited in the embodiments of the present application.
[0058] It can be understood that the first gain is compared with the gain of the previous video clip (second gain). If the absolute value of the difference between the two is less than or equal to the gain threshold, the first initial audio clip is adjusted according to the first gain to obtain the first audio clip.
[0059] In an embodiment of the present application, when the absolute value of the difference between the first gain and the second gain is less than or equal to the gain threshold, the terminal device adjusts the first initial audio segment according to the first gain to obtain the first audio segment, which can ensure that the gain of the audio does not change suddenly, thereby avoiding the loudness of the audio segment (compared to the loudness of the previous audio segment) changing too much due to excessive gain change and affecting the auditory effect.
[0060] Optionally, the above step 203 can be specifically implemented through the following step 203b.
[0061] 203b: When the absolute value of the difference between the first gain and the second gain is greater than the gain threshold, the terminal device adjusts the first initial audio segment according to the third gain to obtain a first audio segment.
[0062] The third gain is greater than the first value and less than the second value; the first value is the smaller value of the first gain and the second gain, and the second value is the larger value of the first gain and the second gain.
[0063] Optionally, an absolute value of a difference between the third gain and the first gain is less than or equal to the gain threshold.
[0064] It can be understood that if the first gain is smaller than the second gain, the third gain is larger than the first gain and smaller than the second gain; if the first gain is larger than the second gain, the third gain is larger than the second gain and smaller than the first gain.
[0065] It can be understood that the first gain is compared with the gain of the previous video clip (the second gain). If the absolute value of the difference between the two is greater than the gain threshold, a third gain (between the first gain and the second gain) is determined based on the first gain and the second gain, and then the first initial audio clip is adjusted based on the third gain to obtain the first audio clip.
[0066] Among them, the preset gain interval can be increased or decreased on the basis of the second gain to obtain a third gain; or the gain interval corresponding to the first difference (half, one third or one quarter of the first difference, etc., the gain interval is less than the gain threshold) can be increased or decreased on the basis of the second gain to obtain a third gain; the specific gain can be determined according to actual usage requirements, and the embodiments of this application are not limited thereto.
[0067] In an embodiment of the present application, when the absolute value of the difference between the first gain and the second gain is greater than the gain threshold, the terminal device adjusts the first initial audio segment according to the third gain between the first gain and the second gain to obtain the first audio segment (that is, through smoothing processing, the first gain is reached with a delay), which can ensure that the audio gain does not change suddenly, and thus avoid the loudness of the audio segment (compared to the loudness of the previous audio segment) changing too much due to excessive gain change, thereby affecting the auditory effect.
[0068] In the embodiment of the present application, when the difference between the first gain and the second gain is too large,
[0069] Optionally, in the embodiment of the present application, the unit of gain (Gain) is decibel (db), a gain of 0db indicates that the corresponding audio segment does not need to be adjusted, and a gain not of 0db indicates that the corresponding audio segment needs to be adjusted.
[0070] It should be noted that if the first gain is 0db, the first initial audio segment does not need to be adjusted; alternatively, if the first gain is 0db and the gain of the adjacent initial audio segment before the first initial audio segment is also 0db or the absolute value is less than or equal to the gain threshold, the first initial audio segment does not need to be adjusted.
[0071] Optionally, the above step 202 can be specifically implemented through the following step 202a.
[0072] 202a. When the target condition is met, the terminal device determines a first gain according to the first zoom range and the first noise level.
[0073] The target condition includes at least one of the following: the first zoom range is different from the second zoom range, and the first noise level is different from the second noise level.
[0074] The second zoom range is: the zoom range corresponding to the second video clip recorded before the first video clip; the second noise level is: the noise level corresponding to the second initial audio clip included in the second video clip.
[0075] The description of the second zoom range may refer to the relevant description of the first zoom range in the above step 202, and is not limited in the embodiment of the present application.
[0076] The description of the second noise level may refer to the description of the first noise level in step 202, and is not limited in this embodiment of the present application.
[0077] The second video segment may be a video segment adjacent to the first video segment, or may be a video segment spaced a certain distance apart from the first video segment, which is not limited in this embodiment of the present application.
[0078] It will be appreciated that, if at least one of the first zoom range and the second zoom range are different, and the first noise level and the second noise level are different, the terminal device determines a first gain based on the first zoom range and the first noise level, and then adjusts the first initial audio segment based on the first gain to obtain the first audio segment. If the first zoom range and the second zoom range are the same, and the first noise level and the second noise level are the same, the terminal device does not need to determine the first gain and can directly adjust the first initial audio segment based on the second gain to obtain the first audio segment.
[0079] In the embodiment of the present application, whether the first gain needs to be determined can be determined based on whether the first zoom range is the same as the second zoom range and whether the first noise level is the same as the second noise level, thereby improving audio processing efficiency.
[0080] Optionally, the above step 202 can be specifically implemented through the following steps 202b to 202c.
[0081] 202b. The terminal device determines a first lookup table corresponding to the first zoom range from multiple lookup tables.
[0082] The first lookup table is a mapping table of noise level and gain.
[0083] Different zoom ranges correspond to different lookup tables, and different lookup tables are mapping tables of different noise levels and gains.
[0084] 202c. The terminal device determines a first gain corresponding to the first noise level according to the first lookup table.
[0085] Exemplarily, as shown in Table 2 and Table 3, different zoom ranges correspond to different lookup tables, and different lookup tables are mapping tables of different noise levels and gains.
[0086] Table 2
[0087]
[0088] Table 3
[0089]
[0090] Optionally, the terminal device may also determine a first functional relationship corresponding to the first zoom range from a plurality of functional relationships, wherein the first functional relationship is a function with noise level as an independent variable and gain as a dependent variable; different zoom ranges correspond to different functional relationships, and different functional relationships are functional relationships between different noise levels and gains.
[0091] Optionally, the above step 202 can be specifically implemented through the following steps 202d to 202e.
[0092] 202d. The terminal device determines a first mapping table corresponding to the first noise level from the multiple mapping tables.
[0093] The first mapping table is a mapping table of zoom range and gain.
[0094] Different mapping tables correspond to different zoom ranges, and different mapping tables are mapping tables for different noise levels and gains.
[0095] 202e. The terminal device determines a first gain corresponding to the first zoom range according to the first mapping table.
[0096] Among them, multiple mapping tables can refer to Table 2 and Table 3, which are not described here in detail.
[0097] Optionally, the terminal device may also determine a second functional relationship corresponding to the first noise level from a plurality of functional relationships. The second functional relationship is a function of the zoom range as an independent variable and the recording gain as a dependent variable; different noise levels correspond to different functional relationships, and the different functional relationships are functional relationships between different noise levels and gains.
[0098] In an embodiment of the present application, multiple schemes for determining the first gain based on the first zoom range and the first noise level are provided, so that a suitable scheme can be determined according to actual usage requirements, thereby improving audio processing efficiency.
[0099] Optionally, when the noise level is less than or equal to a first threshold, the terminal device may determine the gain of the corresponding initial audio segment based on the zoom range. Alternatively, when the noise level is less than or equal to a second threshold, regardless of the zoom range, the terminal device may determine that the corresponding initial audio segment does not require gain adjustment. This improves the efficiency of determining the first gain and audio processing efficiency.
[0100] Among them, the first level threshold and the second level threshold can be the same or different, that is, the first level threshold is less than or equal to the second level threshold. The first level threshold and the second level threshold can be determined according to actual usage requirements, and the embodiments of this application do not limit it.
[0101] Optionally, the above step 203 can be specifically implemented through the following steps 203c to 203d.
[0102] 203c. The terminal device pre-processes the first initial audio segment to obtain a processed initial audio segment.
[0103] 203d. The terminal device adjusts the processed initial audio segment based on the first gain to obtain a first audio segment.
[0104] The pre-processing includes at least one of the following: equalization (EQ) processing and noise reduction processing.
[0105] The fundamental function of EQ is to adjust the sound's timbre by applying gain or attenuation to one or more frequency bands. EQ typically involves three parameters: Frequency, which sets the frequency point to be adjusted; Gain, which adjusts the gain or attenuation at a set F value; and Quantize, which sets the "width" of the frequency band to be gained or attenuated. It's important to note that a smaller Q value results in a wider frequency band, while a larger Q value results in a narrower frequency band.
[0106] Among them, the specific EQ processing technology and noise reduction processing technology can refer to the existing related technologies, and the embodiments of this application are not limited.
[0107] In the embodiment of the present application, preprocessing the first initial audio segment can make the audio effect of the final first audio segment better and improve the audio quality.
[0108] Before the above step 202, the audio processing method provided in the embodiment of the present application may further include the following step 204.
[0109] 204. The terminal device performs noise analysis on the first initial audio segment to obtain a first noise level.
[0110] The noise analysis process includes any one of the following: Gaussian mixture model processing based on a neural network, Mel-frequency cepstral coefficient processing, and noise recognition processing based on a convolutional neural network.
[0111] Among them, Gaussian mixture model processing, Mel-frequency cepstral coefficient processing, and noise recognition processing based on convolutional neural network can refer to existing related technologies and are not limited in the embodiments of this application.
[0112] Among them, noise includes steady-state noise and non-steady-state noise. Steady-state noise refers to noise whose frequency components and amplitudes are basically stable, such as air conditioning sound, white noise, pink noise, wind noise, etc. Non-steady-state noise refers to noise with poor time continuity, such as the whistling sound of cars crossing the road. Current acoustic analysis has made great progress, and noise recognition has become more and more accurate. In the embodiments of this application, the noise analysis and processing method can refer to the existing relevant technologies, and the embodiments of this application do not limit the use of any noise recognition method.
[0113] In the embodiments of the present application, a variety of noise analysis and processing methods are provided, which can be determined according to actual usage requirements, thereby improving audio processing efficiency and improving audio quality.
[0114] In the embodiment of the present application, it is possible to ensure that the audio gain during video recording zoom is automatically adjusted according to the noise level of the ambient noise (the noise loudness level rather than the overall audio loudness level of the audio clip), so that the recording gain is higher in a low-noise environment and lower in a high-noise environment, taking into account both the amplification effect and the subjective listening experience. In this way, a reasonable audio loudness can be output in different noise scenes.
[0115] like Figure 3 As shown, the embodiment of the present application provides an audio processing method. The following takes the execution subject as a terminal device as an example to exemplify the audio processing method provided by the embodiment of the present application. The method may include the following steps 301 to 306.
[0116] 301. The terminal device records a first video clip.
[0117] 302. The terminal device determines a first gain according to the first zoom range and the first noise level.
[0118] The detailed description of steps 301 to 302 may refer to the relevant description of steps 201 to 202, which will not be repeated here.
[0119] 303. The terminal device determines whether the absolute value of the difference between the first gain and the second gain is less than or equal to a gain threshold.
[0120] It can be understood that if the terminal device determines that the absolute value of the difference between the first gain and the second gain is less than or equal to the gain threshold, the following step 304 is executed; if the terminal device determines that the absolute value of the difference between the first gain and the second gain is greater than the gain threshold, the following steps 305 to 306 are executed.
[0121] 304. The terminal device adjusts the first initial audio segment according to the first gain to obtain a first audio segment.
[0122] The detailed description of steps 303 to 304 can refer to the relevant description of step 203a, which will not be repeated here.
[0123] 305. The terminal device determines a third gain based on the first gain and the second gain.
[0124] 306. The terminal device adjusts the first initial audio segment according to the third gain to obtain a first audio segment.
[0125] For the detailed description of the above steps 303, 305 and 306, please refer to the relevant description of the above step 203b, which will not be repeated here.
[0126] like Figure 4 As shown, the embodiment of the present application provides an audio processing method. The following takes the execution subject as a terminal device as an example to exemplify the audio processing method provided by the embodiment of the present application. The method may include the following steps 401 to 407.
[0127] 401. The terminal device records a first video clip.
[0128] The detailed description of step 401 can refer to the relevant description of step 201, which will not be repeated here.
[0129] 402. The terminal device determines a first lookup table corresponding to a first zoom range from multiple lookup tables.
[0130] 403. The terminal device determines a first gain corresponding to the first noise level according to the first lookup table.
[0131] For the detailed description of steps 402 to 403, reference may be made to the relevant description of steps 202b to 202c, which will not be repeated here.
[0132] 404. The terminal device determines whether the absolute value of the difference between the first gain and the second gain is less than or equal to a gain threshold.
[0133] It can be understood that if the terminal device determines that the absolute value of the difference between the first gain and the second gain is less than or equal to the gain threshold, the following step 405 is executed; if the terminal device determines that the absolute value of the difference between the first gain and the second gain is greater than the gain threshold, the following steps 406 to 407 are executed.
[0134] 405. The terminal device adjusts the first initial audio segment according to the first gain to obtain a first audio segment.
[0135] The detailed description of steps 404 to 405 can refer to the relevant description of step 203a, which will not be repeated here.
[0136] 406. The terminal device determines a third gain based on the first gain and the second gain.
[0137] 407. The terminal device adjusts the first initial audio segment according to the third gain to obtain a first audio segment.
[0138] For the detailed description of the above steps 404, 406 and 407, please refer to the relevant description of the above step 203b, which will not be repeated here.
[0139] like Figure 5 As shown, the embodiment of the present application provides an audio processing method. The following takes the execution subject as a terminal device as an example to exemplify the audio processing method provided by the embodiment of the present application. The method may include the following steps 501 to 509.
[0140] 501. The terminal device records a first video clip.
[0141] For the detailed description of step 501 , reference may be made to the relevant description of step 201 , which will not be repeated here.
[0142] 502. The terminal device determines whether the target condition is met.
[0143] It can be understood that if the terminal device does not meet the target condition (i.e., the first zoom range and the second zoom range are the same, and the first noise level and the second noise level are the same), the following step 503 is executed; if the terminal device meets the target condition (i.e., the first zoom range and the second zoom range are different, and at least one of the first noise level and the second noise level is different), the following steps 504 to 509 are executed.
[0144] 503. The terminal device adjusts the first initial audio segment according to the second gain to obtain a first audio segment.
[0145] It can be understood that when the first zoom range is the same as the second zoom range, and the first noise level is the same as the second noise level, there is no need to redetermine the gain corresponding to the first initial audio segment. The first initial audio segment can be adjusted according to the second gain to obtain the first audio segment.
[0146] 504. The terminal device determines a first mapping table corresponding to the first noise level from the multiple mapping tables.
[0147] 505. The terminal device determines a first gain corresponding to the first zoom range according to the first mapping table.
[0148] For the detailed description of steps 502 to 505, reference may be made to the relevant description of step 202a and steps 202d to 202e, which will not be repeated here.
[0149] 506. The terminal device determines whether the absolute value of the difference between the first gain and the second gain is less than or equal to a gain threshold.
[0150] It can be understood that if the terminal device determines that the absolute value of the difference between the first gain and the second gain is less than or equal to the gain threshold, the following step 507 is executed; if the terminal device determines that the absolute value of the difference between the first gain and the second gain is greater than the gain threshold, the following steps 508 to 509 are executed.
[0151] 507. The terminal device adjusts the first initial audio segment according to the first gain to obtain a first audio segment.
[0152] 508. The terminal device determines a third gain based on the first gain and the second gain.
[0153] 509. The terminal device adjusts the first initial audio segment according to the third gain to obtain a first audio segment.
[0154] For the detailed description of steps 506 to 509 , reference may be made to the relevant description of steps 203 a to 203 b , which will not be repeated here.
[0155] Figure 6 This is a structural block diagram of an audio processing device shown in an embodiment of the present application. Figure 6 As shown, it includes: a recording module 601, a determination module 602 and an adjustment module 603; the recording module 601 is used to record a first video clip, which includes a first initial audio clip; the determination module 602 is used to determine a first gain according to a first zoom range of the first video clip recorded by the recording module 601 and a first noise level of the first initial audio clip recorded by the recording module 601; the adjustment module 603 is used to adjust the first initial audio clip based on the first gain determined by the determination module 602 to obtain a first audio clip.
[0156] Optionally, the adjustment module 603 is specifically used to adjust the first initial audio segment according to the first gain to obtain the first audio segment when the absolute value of the difference between the first gain and the second gain is less than or equal to the gain threshold; wherein the second gain is the adjustment gain corresponding to the second initial audio segment, and the second initial audio segment belongs to the second video segment recorded before the first video segment.
[0157] Optionally, the adjustment module 603 is specifically configured to adjust the first initial audio segment according to a third gain to obtain a first audio segment when the absolute value of the difference between the first gain and the second gain is greater than the gain threshold; wherein the third gain is greater than the first value and less than the second value; the first value is the smaller value of the first gain and the second gain, and the second value is the larger value of the first gain and the second gain.
[0158] Optionally, the determination module 602 is specifically used to determine the first gain based on the first zoom range and the first noise level when the target condition is met; wherein the target condition includes at least one of the following: the first zoom range is different from the second zoom range, and the first noise level is different from the second noise level; the second zoom range is: the zoom range corresponding to the second video clip recorded before the first video clip; the second noise level is: the noise level corresponding to the second initial audio clip included in the second video clip.
[0159] Optionally, the determination module 602 is specifically configured to determine a first lookup table corresponding to the first zoom range from a plurality of lookup tables, where the first lookup table is a mapping table of noise level and gain; and determine a first gain corresponding to the first noise level according to the first lookup table.
[0160] Optionally, the determination module 602 is specifically configured to determine a first mapping table corresponding to the first noise level from a plurality of mapping tables, where the first mapping table is a mapping table of zoom range and gain; and determine a first gain corresponding to the first zoom range according to the first mapping table.
[0161] Optionally, the adjustment module 603 is specifically configured to preprocess the first initial audio segment to obtain a processed initial audio segment; and adjust the processed initial audio segment based on the first gain to obtain a first audio segment; wherein the preprocessing includes at least one of the following: EQ processing and noise reduction processing.
[0162] It should be noted that in the embodiment of the present application, the audio processing device can be the terminal device in the above method embodiment, or it can be a functional module and / or functional entity in the terminal device in the above method embodiment that can realize the functions of the above device embodiment, and the embodiment of the present application does not limit this.
[0163] In the embodiment of the present application, each module can implement the audio processing method provided by the above method embodiment and achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0164] Figure 7 A schematic diagram of the hardware structure of a terminal device for implementing each embodiment of the present application is shown as follows: Figure 7 As shown, the terminal device includes but is not limited to: radio frequency (RF) circuit 701, memory 702, input unit 703, display unit 704, sensor 705, audio circuit 706, wireless communication (wireless fidelity, WiFi) module 707, processor 708, power supply 709, and camera 710. Among them, the RF circuit 701 includes a receiver 7011 and a transmitter 7012. Those skilled in the art will understand that Figure 7 The terminal device structure shown in the figure does not constitute a limitation on the terminal device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0165] The RF circuit 701 can be used to receive and send signals during information transmission or calls. In particular, after receiving the downlink information from the base station, it is sent to the processor 708 for processing; in addition, the designed uplink data is sent to the base station. Generally, the RF circuit 701 includes but is not limited to an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the RF circuit 701 can also communicate with the network and other devices through wireless communication. The above-mentioned wireless communication can use any communication standard or protocol, including but not limited to the global system of mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), long term evolution (LTE), email, short messaging service (SMS), etc.
[0166] The memory 702 can be used to store software programs and modules. The processor 708 executes the various functional applications and data processing of the terminal device by running the software programs and modules stored in the memory 702. The memory 702 may mainly include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function (such as a sound playback function, an image playback function, etc.); the data storage area may store data created based on the use of the terminal device (such as audio data, a phone book, etc.). In addition, the memory 702 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device.
[0167] The input unit 703 can be used to receive input digital or character information, and to generate key signal input related to the user settings and function control of the terminal device. Specifically, the input unit 703 may include a touch panel 7031 and other input devices 7032. The touch panel 7031, also known as a touch screen, can collect user touch operations on or near it (such as operations performed by the user using any suitable object or accessory such as a finger, stylus, etc. on or near the touch panel 7031) and drive the corresponding connection device according to a pre-set program. Optionally, the touch panel 7031 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the user's touch direction and detects the signal caused by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device and converts it into touch point coordinates, which are then sent to the processor 708. It can also receive commands sent by the processor 708 and execute them. In addition, the touch panel 7031 can be implemented using a variety of methods such as resistive, capacitive, infrared, and surface acoustic waves. In addition to the touch panel 7031, the input unit 703 may further include other input devices 7032. Specifically, the other input devices 7032 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick.
[0168] The display unit 704 can be used to display information input by the user or information provided to the user and various menus of the terminal device. The display unit 704 may include a display panel 7041. Optionally, the display panel 7041 may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. Further, the touch panel 7031 may cover the display panel 7041. When the touch panel 7031 detects a touch operation on or near it, it is transmitted to the processor 708 to determine the touch event. Subsequently, the processor 708 provides corresponding visual output on the display panel 7041 according to the touch event. Although in Figure 7 In the embodiment, the touch panel 7031 and the display panel 7041 are used as two independent components to realize the input and output functions of the terminal device, but in some embodiments, the touch panel 7031 and the display panel 7041 can be integrated to realize the input and output functions of the terminal device.
[0169] The terminal device may also include at least one sensor 705, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor, wherein the ambient light sensor may adjust the brightness of the display panel 7041 according to the brightness of the ambient light, and the proximity sensor may exit the display panel 7041 and / or the backlight when the terminal device is moved to the ear. As a type of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes), and can detect the magnitude and direction of gravity when stationary. It can be used for applications that identify the posture of the terminal device (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors that the terminal device can also be configured with, such as gyroscopes, geomagnetic sensors, barometers, hygrometers, thermometers, infrared sensors, etc., they will not be described in detail here. In the embodiment of the present application, the terminal device may include an accelerometer, a depth sensor, or a distance sensor, etc.
[0170] Audio circuit 706, speaker 7061, and microphone 7062 provide an audio interface between the user and the terminal device. Audio circuit 706 converts received audio data into electrical signals and transmits them to speaker 7061, which then converts them into sound signals for output. Microphone 7062, on the other hand, converts collected sound signals into electrical signals, which are then received by audio circuit 706 and converted into audio data. The audio data is then processed by processor 708 and transmitted via RF circuit 701 to, for example, another terminal device, or the audio data is output to memory 702 for further processing.
[0171] WiFi is a short-range wireless transmission technology. The terminal device can help users send and receive emails, browse web pages and access streaming media through the WiFi module 707. It provides users with wireless broadband Internet access. Figure 7 A WiFi module 707 is shown, but it is understandable that it is not an essential component of the terminal device and can be omitted as needed without changing the essence of the invention.
[0172] Processor 708 is the control center of the terminal device. It connects the various components of the entire terminal device using various interfaces and lines. By running or executing software programs and / or modules stored in memory 702 and accessing data stored in memory 702, it performs various terminal device functions and processes data, thereby providing overall monitoring of the terminal device. Optionally, processor 708 may include one or more processing units. Preferably, processor 708 may integrate an application processor and a modem processor. The application processor primarily processes the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 708.
[0173] The terminal device also includes a power supply 709 (e.g., a battery) for supplying power to the various components. Preferably, the power supply can be logically connected to the processor 708 via a power management system, thereby enabling the power management system to manage functions such as charging, discharging, and power consumption. The terminal device also includes a camera 710 for recording video images in the video clip. Although not shown, the terminal device may also include a Bluetooth module, etc., which will not be described in detail here.
[0174] In an embodiment of the present application, the processor 708 is used to record a first video clip, where the first video clip includes a first initial audio clip; determine a first gain based on a first zoom range of the first video clip and a first noise level of the first initial audio clip; and adjust the first initial audio clip based on the first gain to obtain a first audio clip.
[0175] Optionally, the processor 708 is specifically used to adjust the first initial audio segment according to the first gain to obtain the first audio segment when the absolute value of the difference between the first gain and the second gain is less than or equal to the gain threshold; wherein the second gain is the adjustment gain corresponding to the second initial audio segment, and the second initial audio segment belongs to the second video segment recorded before the first video segment.
[0176] Optionally, the processor 708 is specifically configured to adjust the first initial audio segment according to a third gain to obtain a first audio segment when the absolute value of the difference between the first gain and the second gain is greater than the gain threshold; wherein the third gain is greater than the first value and less than the second value; the first value is the smaller value of the first gain and the second gain, and the second value is the larger value of the first gain and the second gain.
[0177] Optionally, the processor 708 is specifically used to determine a first gain based on the first zoom range and the first noise level when a target condition is met; wherein the target condition includes at least one of the following: the first zoom range is different from the second zoom range, and the first noise level is different from the second noise level; the second zoom range is: the zoom range corresponding to the second video clip recorded before the first video clip; the second noise level is: the noise level corresponding to the second initial audio clip included in the second video clip.
[0178] Optionally, the processor 708 is specifically configured to determine a first lookup table corresponding to the first zoom range from a plurality of lookup tables, where the first lookup table is a mapping table of noise level and gain; and determine a first gain corresponding to the first noise level according to the first lookup table.
[0179] Optionally, the processor 708 is specifically configured to determine a first mapping table corresponding to the first noise level from a plurality of mapping tables, where the first mapping table is a mapping table of zoom range and gain; and determine a first gain corresponding to the first zoom range according to the first mapping table.
[0180] Optionally, the processor 708 is specifically used to preprocess the first initial audio segment to obtain a processed initial audio segment; adjust the processed initial audio segment based on the first gain to obtain a first audio segment; wherein the preprocessing includes at least one of the following: EQ processing, noise reduction processing.
[0181] The beneficial effects of various implementation methods in this embodiment can be specifically referred to the beneficial effects of the corresponding implementation methods in the above-mentioned audio processing method embodiment. To avoid repetition, they will not be described here.
[0182] An embodiment of the present application also provides a terminal device, which may include: a processor, a memory, and a program or instruction stored in the memory and runnable on the processor. When the program or instruction is executed by the processor, the various processes of the audio processing method provided in the above method embodiment can be implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0183] An embodiment of the present application provides a readable storage medium, which stores a program or instruction. When the program or instruction is executed by a processor, the various processes of the audio processing method provided in the above method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0184] An embodiment of the present application also provides a computer program product, wherein the computer program product includes computer instructions. When the computer program product runs on a processor, the processor executes the computer instructions to implement the various processes of the audio processing method provided in the above method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0185] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned audio processing method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0186] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.
[0187] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, servers and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0188] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0189] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0190] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0191] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. An audio processing method, characterized in that: The method comprises: Recording a first video clip, wherein the first video clip includes a first initial audio clip; determining a first gain based on a first zoom range of the first video clip and a first noise level of the first initial audio clip; when the first zoom range remains unchanged, the smaller the value corresponding to the first noise level, the larger the first gain; and the larger the value corresponding to the first noise level, the smaller the first gain; The first initial audio segment is adjusted based on the first gain to obtain a first audio segment.
2. The method according to claim 1, characterized in that The step of adjusting the first initial audio segment based on the first gain to obtain a first audio segment includes: When an absolute value of a difference between the first gain and the second gain is less than or equal to a gain threshold, adjusting the first initial audio segment according to the first gain to obtain the first audio segment; The second gain is an adjustment gain corresponding to a second initial audio segment, and the second initial audio segment belongs to a second video segment recorded before the first video segment.
3. The method according to claim 2, characterized in that The adjusting the first initial recording data based on the first gain to obtain the first recording data includes: When the absolute value of the difference between the first gain and the second gain is greater than the gain threshold, adjusting the first initial audio segment according to a third gain to obtain the first audio segment; The third gain is greater than the first value and less than the second value; the first value is the smaller value between the first gain and the second gain, and the second value is the larger value between the first gain and the second gain.
4. The method according to claim 1, wherein The determining a first gain according to a first zoom range of the first video segment and a first noise level of the first initial audio segment includes: determining the first gain according to the first zoom range and the first noise level when a target condition is met; The target condition includes at least one of the following: the first zoom range is different from the second zoom range, and the first noise level is different from the second noise level; The second zoom range is: the zoom range corresponding to the second video clip recorded before the first video clip; the second noise level is: the noise level corresponding to the second initial audio clip included in the second video clip.
5. The method according to claim 1, wherein The determining a first gain according to a first zoom range of the first video segment and a first noise level of the first initial audio segment includes: Determining a first lookup table corresponding to the first zoom range from a plurality of lookup tables, wherein the first lookup table is a mapping table of noise level and gain; A first gain corresponding to the first noise level is determined according to the first lookup table.
6. The method according to claim 1, characterized in that The determining a first gain according to a first zoom range of the first video segment and a first noise level of the first initial audio segment includes: Determining a first mapping table corresponding to the first noise level from a plurality of mapping tables, wherein the first mapping table is a mapping table of zoom range and gain; A first gain corresponding to the first zoom range is determined according to the first mapping table.
7. The method according to any one of claims 1 to 6, characterized in that The step of adjusting the first initial audio segment based on the first gain to obtain a first audio segment includes: Preprocessing the first initial audio segment to obtain a processed initial audio segment; adjusting the processed initial audio segment based on the first gain to obtain the first audio segment; The pre-processing includes at least one of the following: equalization (EQ) processing and noise reduction processing.
8. An audio processing device, characterized in that: The device includes: a recording module, a determination module and an adjustment module; The recording module is configured to record a first video clip, wherein the first video clip includes a first initial audio clip; The determining module is configured to determine a first gain based on a first zoom range of the first video clip recorded by the recording module and a first noise level of the first initial audio clip recorded by the recording module; when the first zoom range remains unchanged, the smaller the value corresponding to the first noise level, the larger the first gain; and the larger the value corresponding to the first noise level, the smaller the first gain; The adjustment module is configured to adjust the first initial audio segment based on the first gain determined by the determination module to obtain a first audio segment.
9. A terminal device, characterized in that: The method comprises a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the audio processing method according to any one of claims 1 to 7.
10. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by a processor, the steps of the audio processing method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Volume adjustment method and device and equipment
CN107124149A
Sound processing method, device and equipment
CN110970057A