Audio enhancement method and device, equipment and storage medium

By obtaining the audio sampling rate and performing delay and superimposition according to the energy difference, the problem of audio lacking spatial and immersion in consumer scenarios is solved, and the audio enhancement effect is achieved.

CN120343483APending Publication Date: 2025-07-18BEIJING ZITIAO NETWORK TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410064654.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-16
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, the audio in the consumer scenario is usually mono or pseudo stereo, lacking spatial and immersion. Traditional enhancement methods have problems such as auditory fatigue, poor versatility and performance impact.

Method used

By obtaining the sampling rate of the target audio, determine the energy difference between the left and right channels. If the threshold is not exceeded, the audio is delayed and superimposed according to the sampling rate to enhance the left and right channels audio, and improve the sense of space and immersion.

Benefits of technology

It improves the spatial and immersion of the audio, avoids the auditory fatigue and performance impact of traditional methods, and enhances the overall effect of the audio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343483A_ABST
    Figure CN120343483A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an audio enhancement method and device, equipment and a storage medium. Acquiring the sampling rate of the target audio; wherein the target audio comprises a left channel audio and a right channel audio; determining an energy difference between the left channel audio and the right channel audio; judging whether the energy difference exceeds a set threshold value or not; and if the energy difference does not exceed a set threshold value, performing enhancement processing on the left channel audio and / or the right channel audio according to the sampling rate to obtain an enhanced target audio. According to the audio enhancement method provided by the embodiment of the invention, when the energy difference between the left channel audio and the right channel audio does not exceed the set threshold value, enhancement processing is performed on the left channel audio and / or the right channel audio according to the sampling rate of the target audio, so that the sense of space and immersion of the target audio can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of audio processing technologies, and in particular, to an audio enhancement method, apparatus, device, and storage medium. Background Art

[0002] The sense of space or immersion of sound is affected by two factors: the sense of direction of the sound source and the reverberation effect generated by phenomena such as reflection and diffraction of sound waves in the environment. Currently, in consumer scenarios such as short video on-demand or live broadcast scenarios, many audio are monophonic, pseudo-stereo, or audio with insignificant left and right channel differences. Although in encoding or transmission, a monophonic audio may be encoded into a stereo format file, in fact, it is just to copy the monophonic audio into two identical audios to form a pseudo-stereo, and in terms of listening experience, it still does not give users a sense of space and immersion. Therefore, it is particularly important to perform enhancement processing on audio. Summary of the Invention

[0003] Embodiments of the present disclosure provide an audio enhancement method, apparatus, device, and storage medium, which can improve the sense of space and immersion of audio.

[0004] In a first aspect, embodiments of the present disclosure provide an audio enhancement method, including:

[0005] Obtaining a sampling rate of a target audio; wherein, the target audio includes a left-channel audio and a right-channel audio;

[0006] Determining an energy difference between the left-channel audio and the right-channel audio;

[0007] Judging whether the energy difference exceeds a set threshold;

[0008] If the energy difference does not exceed the set threshold, enhancing the left-channel audio and / or the right-channel audio according to the sampling rate to obtain an enhanced target audio.

[0009] In a second aspect, embodiments of the present disclosure further provide an audio enhancement apparatus, including:

[0010] A sampling rate acquisition module, configured to obtain a sampling rate of a target audio; wherein, the target audio includes a left-channel audio and a right-channel audio;

[0011] An energy difference determination module, configured to determine an energy difference between the left-channel audio and the right-channel audio;

[0012] A judgment module, configured to judge whether the energy difference exceeds a set threshold;

[0013] An audio enhancement module, configured to perform enhancement processing on the left-channel audio and / or the right-channel audio according to the sampling rate to obtain enhanced target audio when the energy difference does not exceed a set threshold.

[0014] In a third aspect, an embodiment of the present disclosure further provides an electronic device, where the electronic device includes:

[0015] One or more processors;

[0016] A storage device, configured to store one or more programs,

[0017] When the one or more programs are executed by the one or more processors, the one or more processors implement the audio enhancement method as described in the embodiments of the present disclosure.

[0018] In a fourth aspect, an embodiment of the present disclosure further provides a storage medium containing computer-executable instructions, where the computer-executable instructions are used to execute the audio enhancement method as described in the embodiments of the present disclosure when executed by a computer processor.

[0019] Embodiments of the present disclosure disclose an audio enhancement method, apparatus, device, and storage medium. The sampling rate of target audio is obtained, where the target audio includes left-channel audio and right-channel audio. The energy difference between the left-channel audio and the right-channel audio is determined. It is judged whether the energy difference exceeds a set threshold. If the energy difference exceeds the set threshold, enhancement processing is performed on the left-channel audio and / or the right-channel audio according to the sampling rate to obtain enhanced target audio. In the audio enhancement method provided by the embodiments of the present disclosure, when the energy difference between the left-channel audio and the right-channel audio does not exceed the set threshold, enhancement processing is performed on the left-channel audio and / or the right-channel audio according to the sampling rate of the target audio, which can improve the spatial sense and immersion of the target audio. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the original elements and elements are not necessarily drawn to scale.

[0021] Figure 1 is a flowchart of an audio enhancement method provided by an embodiment of the present disclosure;

[0022] Figure 2 is an example diagram of performing enhancement processing on target audio provided by an embodiment of the present disclosure;

[0023] Figure 3 is a flowchart of an audio enhancement method provided by an embodiment of the present disclosure;

[0024] Figure 4 is a schematic structural diagram of an audio enhancement device provided by an embodiment of the present disclosure;

[0025] Figure 5 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Specific embodiments

[0026] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not used to limit the protection scope of the present disclosure.

[0027] It should be understood that the various steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0028] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0029] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependent relationships.

[0030] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".

[0031] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0032] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0033] For example, when a user's active request is received, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that performs the operations of the present disclosure's technical solution based on the prompt message.

[0034] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving the user's active request may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0035] It can be understood that the above notification and user authorization process is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0036] It can be understood that the data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of corresponding laws, regulations and related provisions.

[0037] A traditional audio enhancement method is to increase the loudness difference or delay difference between the left and right channels of the pseudo stereo to obtain an audio that does not sound in the exact middle but rather sounds to the left or right to enhance the sense of space of the audio. This method has the following defects: it overall changes the direction of the sound source, and the sound always being to the left or right will cause auditory fatigue to people; it does not increase reverberation, and the audio that originally sounded in a smaller space still does not change its perception of the size of the space; it will damage the audio with a relatively good original effect.

[0038] Another traditional method is to convolve the original audio with a sampled reverberation or use a reverberator to adjust the reverberation. The defects of this method are: it needs to be adjusted separately according to the reverberation of the original audio, and the versatility is poor; the convolution reverberation calculation is large, which will affect the system performance.

[0039] Figure 1 It is a flowchart of an audio enhancement method provided by an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to the situation of enhancing audio. This method can be executed by an audio enhancement device, and this device can be implemented in the form of software and / or hardware. Optionally, it is implemented by an electronic device, and this electronic device can be a mobile terminal, a PC terminal or a server, etc.

[0040] As Figure 1 shown, the method includes:

[0041] S110, obtain the sampling rate of the target audio.

[0042] Among them, the target audio includes the left-channel audio and the right-channel audio. The left-channel audio and the right-channel audio can be understood as such. The target audio can be the audio data obtained by splitting the real-time audio stream according to a set duration, and the set duration can be pre-set. In one application scenario, the real-time audio stream can be the audio stream generated during a live broadcast.

[0043] In this embodiment, the target audio is represented in the form of a digital signal. Therefore, the sampling rate can be understood as the audio sampling points included in the target audio per unit time. For example: the audio sampling points included in 1 second. The sampling rate is the attribute information of the target audio and can be directly obtained. For example, the description information or configuration information in the target audio file can be read.

[0044] S120, determine the energy difference between the left-channel audio and the right-channel audio.

[0045] Among them, the energy difference can be understood as the difference between the energy of the left-channel audio and the energy of the right-channel audio. The process of determining the energy difference between the left-channel audio and the right-channel audio can be: first, determine the energy of the left-channel audio and the energy of the right-channel audio respectively, and then subtract the energy of the right-channel audio from the energy of the left-channel audio to obtain the energy difference.

[0046] Optionally, the way to determine the energy difference between the left-channel audio and the right-channel audio can be: obtain the loudness of each sampling point of the left-channel audio and the right-channel audio respectively; determine the energy of the left-channel audio and the right-channel audio respectively based on the loudness; determine the energy difference between the left-channel audio and the right-channel audio based on the energy.

[0047] Among them, the loudness is characterized by the amplitude of each sampling point. The way to determine the energy of the left-channel audio and the right-channel audio respectively based on the loudness can be: determine the root mean square of the amplitudes of each sampling point of the left-channel audio as the energy of the left-channel audio, and determine the root mean square of the amplitudes of each sampling point of the right-channel audio as the energy of the right-channel audio. After obtaining the energies of the two channels, subtract the two energies to obtain the energy difference. Exemplarily, assume that the amplitudes of each sampling point of the left-channel audio are represented as L i , and the amplitudes of each sampling point of the right-channel audio are represented as R i , then the calculation formula of the energy difference can be expressed as: Among them, D represents the calculated energy difference, and N represents the number of sampling points included in the two channels.

[0048] S130, determine whether the energy difference exceeds the set threshold; if the energy difference does not exceed the set threshold, then execute S140.

[0049] Among them, the set threshold value can be a value obtained in advance through a listening experiment. For example, it can be set to any value between 0.3 and 0.5.

[0050] In this embodiment, if the energy difference between the left-channel audio and the right-channel audio exceeds the set threshold value, it indicates that the difference between the left-channel audio and the right-channel audio is relatively large, and the sense of space of the target audio is relatively good. At this time, there is no need to process the target audio and it can be directly output. If the energy difference between the left-channel audio and the right-channel audio does not exceed the set threshold value, it indicates that the sense of space of the target audio is weak, and at this time, the target audio needs to be enhanced.

[0051] S140, perform enhancement processing on the left-channel audio and / or the right-channel audio according to the sampling rate to obtain the enhanced target audio.

[0052] Among them, the method of performing enhancement processing on the left-channel audio and / or the right-channel audio according to the sampling rate can be: perform delay processing on the left-channel audio and / or the right-channel audio respectively according to the sampling rate, and then perform enhancement processing on the left-channel audio based on the right-channel audio after delay processing, and / or, perform enhancement processing on the right-channel audio based on the left-channel audio after delay processing.

[0053] Specifically, the method of performing delay processing on the left-channel audio and / or the right-channel audio respectively according to the sampling rate can be: determine the first delay information and / or the second delay information respectively according to the sampling rate; perform delay processing on the left-channel audio based on the first delay information, and / or, perform delay processing on the right-channel audio based on the second delay information.

[0054] Among them, the first delay information and the second delay information can be characterized by the number of sampled points of the delay.

[0055] Specifically, the process of determining the first delay information according to the sampling rate can be: obtain the reference sampling rate and the first set value, and then determine the first delay information according to the sampling rate of the target audio, the reference sampling rate and the first set value. The method of determining the first delay information according to the sampling rate of the target audio, the reference sampling rate and the first set value can be: divide the sampling rate of the target audio by the reference sampling rate, and then multiply the quotient result by the first set value to obtain the first delay information, that is, the number of sampled points by which the left-channel audio is delayed backward. The formula can be expressed as: d1 = (f / F)·A1, where f represents the sampling rate of the target audio (i.e., the left-channel audio), F is the reference sampling rate, and A1 is the first set value. The reference sampling rate can be set by the user. For example: 44100, and the first set value can be any value between 400 and 500. For example: 496.

[0056] Specifically, the process of determining the second delay information according to the sampling rate can be as follows: obtain the reference sampling rate and the second setting value, and then determine the second delay information according to the sampling rate of the target audio, the reference sampling rate, and the second setting value. Among them, the first setting value and the second setting value can be the same or different. The method of determining the second delay information according to the sampling rate of the target audio, the reference sampling rate, and the second setting value can be: divide the sampling rate of the target audio by the reference sampling rate, and then multiply the quotient by the second setting value to obtain the second delay information, that is, the number of sampling points by which the right-channel audio is delayed backward. The formula can be expressed as: d2 = (f / F)·A2, where f represents the sampling rate of the target audio (i.e., the left-channel audio), F is the reference sampling rate, and A2 is the second setting value. The reference sampling rate can be set by the user, for example: 44100, and the second setting value can be any value between 400 and 500, for example: 478.

[0057] In this embodiment, the method of delaying the left-channel audio based on the first delay information can be: delaying the left-channel audio backward by the number of sampling points corresponding to the first delay information. The method of delaying the right-channel audio based on the second delay information can be: delaying the right-channel audio backward by the number of sampling points corresponding to the second delay information.

[0058] Specifically, the method of enhancing the left-channel audio based on the delayed right-channel audio can be: superimposing the delayed right-channel audio and the left-channel audio to obtain the enhanced left-channel audio.

[0059] In this embodiment, the process of superimposing the delayed right-channel audio and the left-channel audio can be: first, fuse the delayed right-channel audio with the setting coefficient, and then superimpose it with the left-channel audio to obtain the enhanced left-channel audio. Among them, the method of fusing the delayed right-channel audio with the setting coefficient can be: multiplying the right-channel audio by the setting coefficient. Among them, the setting coefficient is set in advance, for example, it can be any value between 0.4 and 0.6, for example: 0.5.

[0060] Specifically, the method of enhancing the right-channel audio based on the delayed left-channel audio can be: superimposing the delayed left-channel audio and the right-channel audio to obtain the enhanced right-channel audio.

[0061] In this embodiment, the process of superimposing the left-channel audio after delay processing and the right-channel audio may be as follows: First, the left-channel audio after delay processing is fused with a set coefficient, and then superimposed with the right-channel audio to obtain the enhanced right-channel audio. Among them, the way of fusing the left-channel audio after delay processing and the set coefficient may be: multiplying the left-channel audio by the set coefficient. Among them, the set coefficient is set in advance, for example, it can be any value between 0.4 and 0.6, for example: 0.5.

[0062] Exemplarily, Figure 2 is an example diagram of enhancing the target audio in the embodiments of the present disclosure. As Figure 2 shown, the left-channel audio is delayed by d1 sampling points, then multiplied by the set coefficient, and then superimposed with the right-channel audio to obtain the enhanced right-channel audio; the right-channel audio is delayed by d2 sampling points, then multiplied by the set coefficient, and then superimposed with the left-channel audio to obtain the enhanced left-channel audio.

[0063] In this embodiment, after the target audio is enhanced, digital signal overflow may occur. Therefore, in order to prevent unnecessary changes in the overall volume of the target audio, it is necessary to adjust the loudness of the enhanced target audio. The overflow here refers to the digital signal of the target audio exceeding the range that can be represented by the data type. For example, for an int16 type digital signal, its value range needs to be between -32768 and 32768.

[0064] Specifically, after obtaining the enhanced target audio, the following steps are further included: obtaining the loudness range; performing audio compression processing and / or dynamic range control processing on the enhanced target audio based on the loudness range to obtain the target audio with adjusted loudness.

[0065] Among them, the loudness range can be set in advance and is used to limit the loudness of the target audio. The enhanced target audio includes the enhanced left-channel audio and the enhanced right-channel audio, that is, the loudness range of the enhanced left-channel audio and the enhanced right-channel audio is adjusted respectively based on the loudness range to obtain the left-channel audio with adjusted loudness and the right-channel audio with adjusted loudness.

[0066] In this embodiment, the adjustment of the loudness range of the audio can be implemented by using an audio compressor / limiter or other dynamic range control algorithms. In this embodiment, both the audio compressor and the dynamic range control algorithm can adjust the loudness range of the target audio.

[0067] Optionally, Figure 3 is a schematic flowchart of an audio enhancement method provided by the embodiments of the present disclosure. As Figure 3 shown, after obtaining the enhanced target audio, the following steps are further included:

[0068] S150. Determine the gain factor according to the loudness of each sampling point in the enhanced target audio.

[0069] S160. Adjust the loudness of each sampling point in the enhanced target audio based on the gain factor to obtain the target audio with adjusted loudness.

[0070] In this embodiment, determine the gain factor corresponding to the left-channel audio according to the loudness of each sampling point in the enhanced left-channel audio; determine the gain factor corresponding to the right-channel audio according to the loudness of each sampling point in the enhanced right-channel audio.

[0071] Among them, the method for determining the gain factor corresponding to the left-channel audio according to the loudness of each sampling point in the enhanced left-channel audio can be: extract the maximum amplitude of the sampling points in the enhanced left-channel audio, and determine the reciprocal of the maximum amplitude as the gain factor corresponding to the left-channel audio. The method for determining the gain factor corresponding to the right-channel audio according to the loudness of each sampling point in the enhanced right-channel audio can be: extract the maximum amplitude of the sampling points in the enhanced right-channel audio, and determine the reciprocal of the maximum amplitude as the gain factor corresponding to the right-channel audio.

[0072] After obtaining the gain factor, multiply the gain factor by the values of each sampling point in the enhanced target audio to obtain the target audio with adjusted loudness. This can prevent popping sounds.

[0073] Optionally, a limiter can also be used to process the enhanced target audio. This can prevent popping sounds.

[0074] The technical solution of the embodiment of the present disclosure: obtain the sampling rate of the target audio; where the target audio includes left-channel audio and right-channel audio; determine the energy difference between the left-channel audio and the right-channel audio; determine whether the energy difference exceeds a set threshold; if the energy difference exceeds the set threshold, perform enhancement processing on the left-channel audio and / or the right-channel audio according to the sampling rate to obtain the enhanced target audio. The audio enhancement method provided by the embodiment of the present disclosure can improve the spatial sense and immersion of the target audio when the energy difference between the left-channel audio and the right-channel audio does not exceed the set threshold by performing enhancement processing on the left-channel audio and / or the right-channel audio according to the sampling rate of the target audio.

[0075] Figure 4 As shown in the structural schematic diagram of an audio enhancement device provided by the embodiment of the present disclosure, Figure 4 The device includes:

[0076] A sampling rate acquisition module 310, configured to acquire the sampling rate of the target audio; where the target audio includes left-channel audio and right-channel audio;

[0077] An energy difference determination module 320, configured to determine an energy difference between a left-channel audio and a right-channel audio;

[0078] A determination module 330, configured to determine whether the energy difference exceeds a set threshold;

[0079] An audio enhancement module 340, configured to, when the energy difference does not exceed the set threshold, perform enhancement processing on the left-channel audio and / or the right-channel audio according to a sampling rate to obtain an enhanced target audio.

[0080] Optionally, the energy difference determination module 320 is further configured to:

[0081] Obtain the loudness of each sampling point of the left-channel audio and the right-channel audio respectively;

[0082] Determine the energy of the left-channel audio and the right-channel audio respectively based on the loudness;

[0083] Determine the energy difference between the left-channel audio and the right-channel audio based on the energy.

[0084] Optionally, the audio enhancement module 340 is further configured to:

[0085] Perform delay processing on the left-channel audio and / or the right-channel audio respectively according to the sampling rate;

[0086] Perform enhancement processing on the left-channel audio based on the delay-processed right-channel audio, and / or perform enhancement processing on the right-channel audio based on the delay-processed left-channel audio.

[0087] Optionally, the audio enhancement module 340 is further configured to:

[0088] Determine first delay information and / or second delay information respectively according to the sampling rate;

[0089] Perform delay processing on the left-channel audio based on the first delay information, and / or perform delay processing on the right-channel audio based on the second delay information.

[0090] Optionally, the audio enhancement module 340 is further configured to:

[0091] Superimpose the delay-processed right-channel audio and the left-channel audio to obtain enhanced left-channel audio; and / or,

[0092] Superimpose the delay-processed left-channel audio and the right-channel audio to obtain enhanced right-channel audio.

[0093] Optionally, it further includes a first loudness adjustment module, configured to:

[0094] Obtain a loudness range;

[0095] Perform audio compression processing and / or dynamic range control processing on the enhanced target audio based on the loudness range to obtain the target audio with adjusted loudness.

[0096] Optionally, it further includes a second loudness adjustment module for:

[0097] Determine the gain factor according to the loudness of each sampling point in the enhanced target audio;

[0098] Adjust the loudness of each sampling point in the enhanced target audio based on the gain factor to obtain the target audio with adjusted loudness.

[0099] The audio enhancement device provided by the embodiments of the present disclosure can execute the audio enhancement method provided by any embodiment of the present disclosure, and has corresponding functional modules and beneficial effects for executing the method.

[0100] It should be noted that the various units and modules included in the above device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the embodiments of the present disclosure.

[0101] Figure 5 It is a schematic structural diagram of an electronic device provided by the embodiments of the present disclosure. The following refers to Figure 5 , which shows a schematic structural diagram of an electronic device 500 suitable for implementing the embodiments of the present disclosure (such as Figure 5 the terminal device or server in). The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The electronic device shown is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present disclosure.

[0102] As Figure 5 shown, the electronic device 500 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 502 or the program loaded from the storage device 508 into the random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. The editing / output (I / O) interface 505 is also connected to the bus 504.

[0103] Typically, the following devices can be connected to the I / O interface 505: input devices 506 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. The communication device 509 can allow the electronic device 500 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 5 the electronic device 500 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices can be implemented or had.

[0104] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above functions defined in the methods of the embodiments of the present disclosure are executed.

[0105] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0106] The electronic device provided by the embodiment of the present disclosure and the audio enhancement method provided by the above embodiment belong to the same inventive concept. Technical details not described in detail in this embodiment can be referred to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.

[0107] The embodiment of the present disclosure provides a computer storage medium, on which a computer program is stored, and when the program is executed by a processor, the audio enhancement method provided by the above embodiment is implemented.

[0108] It should be noted that the above-mentioned computer-readable medium in the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0109] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0110] The above-mentioned computer-readable medium may be included in the above-mentioned electronic device; or it may exist separately without being assembled into the electronic device.

[0111] The above-mentioned computer-readable medium carries one or more programs, and when the above one or more programs are executed by the electronic device, the electronic device is caused to:

[0112] The above computer-readable medium carries one or more programs which, when executed by the electronic device, cause the electronic device to: obtain the sampling rate of a target audio; wherein the target audio includes a left-channel audio and a right-channel audio; determine the energy difference between the left-channel audio and the right-channel audio; determine whether the energy difference exceeds a set threshold; and if the energy difference does not exceed the set threshold, enhance the left-channel audio and / or the right-channel audio according to the sampling rate to obtain an enhanced target audio.

[0113] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0114] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0115] The units described in the embodiments of the present disclosure may be implemented in software or in hardware. Wherein, the name of the unit does not, in some cases, constitute a limitation on the unit itself. For example, the first acquisition unit may also be described as "the unit for acquiring at least two Internet protocol addresses".

[0116] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0117] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a Random Access Memory (RAM), a Read Only Memory (ROM), an Erasable Programmable Read Only Memory (EPROM or Flash memory), an optical fiber, a portable Compact Disc Read Only Memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0118] The above description is only a preferred embodiment of this disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the disclosure involved in this disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in this disclosure.

[0119] Furthermore, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although a number of specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of separate embodiments can also be implemented combinatorially in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0120] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. An audio enhancement method, characterized in that, including: obtaining a sampling rate of a target audio; wherein, the target audio includes a left-channel audio and a right-channel audio; determining an energy difference between the left-channel audio and the right-channel audio; judging whether the energy difference exceeds a set threshold; if the energy difference does not exceed the set threshold, enhancing the left-channel audio and / or the right-channel audio according to the sampling rate to obtain an enhanced target audio.

2. The method according to claim 1, wherein Determining the energy difference between the left-channel audio and the right-channel audio includes: respectively obtaining the loudness of each sampling point of the left-channel audio and the right-channel audio; respectively determining the energy of the left-channel audio and the right-channel audio based on the loudness; determining the energy difference between the left-channel audio and the right-channel audio based on the energy.

3. The method according to claim 1, wherein Enhancing the left-channel audio and / or the right-channel audio according to the sampling rate includes: respectively performing delay processing on the left-channel audio and / or the right-channel audio according to the sampling rate; enhancing the left-channel audio based on the delay-processed right-channel audio, and / or enhancing the right-channel audio based on the delay-processed left-channel audio.

4. The method according to claim 3, wherein Respectively performing delay processing on the left-channel audio and / or the right-channel audio according to the sampling rate includes: respectively determining first delay information and / or second delay information according to the sampling rate; performing delay processing on the left-channel audio based on the first delay information, and / or performing delay processing on the right-channel audio based on the second delay information.

5. The method according to claim 3, characterized in that Enhancing the left-channel audio based on the delay-processed right-channel audio, and / or enhancing the right-channel audio based on the delay-processed left-channel audio includes: superposing the delay-processed right-channel audio and the left-channel audio to obtain an enhanced left-channel audio; and / or, superposing the delay-processed left-channel audio and the right-channel audio to obtain an enhanced right-channel audio.

6. The method according to claim 1, wherein After obtaining the enhanced target audio, it further includes: obtaining a loudness range; performing audio compression processing and / or dynamic range control processing on the enhanced target audio based on the loudness range to obtain a target audio with adjusted loudness.

7. The method according to claim 1, characterized in that, After obtaining the enhanced target audio, it further includes: determining a gain factor according to the loudness of each sampling point in the enhanced target audio; adjusting the loudness of each sampling point in the enhanced target audio based on the gain factor to obtain a target audio with adjusted loudness.

8. An audio enhancement device, characterized in that, including: a sampling rate acquisition module for obtaining a sampling rate of a target audio; wherein, the target audio includes a left-channel audio and a right-channel audio; an energy difference determination module for determining an energy difference between the left-channel audio and the right-channel audio; a judgment module for judging whether the energy difference exceeds a set threshold; an audio enhancement module for, when the energy difference does not exceed the set threshold, enhancing the left-channel audio and / or the right-channel audio according to the sampling rate to obtain an enhanced target audio.

9. An electronic device, characterized in that, The electronic device includes: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the audio enhancement method according to any one of claims 1-7.

10. A storage medium containing computer-executable instructions, which are used to execute the audio enhancement method according to any one of claims 1-7 when executed by a computer processor.