Audio playing tail POP sound optimization method and electronic equipment
By presetting breakpoints in the audio file and generating replacement audio, the problem of pop noise at the end of audio playback is solved, and the pop noise during audio playback is eliminated or reduced, thus improving the user experience.
Patent Information
- Application Number
- CN202510758402.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-11-04
AI Technical Summary
Existing technologies struggle to effectively eliminate pop noise at the end of audio playback, especially since the stop time cannot be accurately determined when audio or video playback ends, making it difficult to handle pop noise.
By pre-setting the breakpoint position of the audio file to be played, a replacement audio corresponding to the breakpoint position is generated. The replacement audio is replaced or inserted within a specific time period before and after the breakpoint position, gradually reducing the amplitude of the audio signal to generate the final playback audio, thus weakening or eliminating pop noise.
It enables the elimination or reduction of pop noise during audio playback, improving the user experience.
Smart Images

Figure CN120897145A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of audio playing, more particularly to an audio playing tail POP sound optimization method and electronic device. BACKGROUND
[0002] POP sound generally occurs at the moment of starting or stopping of audio amplifier, or at the starting, pausing or terminating moment of audio file playing. At these moments, the amplitude of the electrical signal received by the transducer component (speaker) will change dramatically, and the dramatic change of the electrical signal is converted into sound by the transducer component (speaker), which is the "POP sound" heard by people. The dramatic change of the amplitude of the audio electrical signal at the specific moment is the source of the POP sound. How to eliminate the POP sound is a difficult problem in the process of audio playing.
[0003] Especially in smart devices such as Android, at the beginning or controlled interruption of audio (or video) playing, because the playing processor knows when the starting point or interruption point of audio playing is, it can process at the beginning or interruption of audio to weaken the influence of "POP sound"; but at the end or artificial interruption of audio (or video) playing, because the playing processor cannot accurately know when the stopping or interruption time of audio is, it is difficult to process the timely POP sound. SUMMARY
[0004] The technical problem to be solved by the present application is to provide an audio playing tail POP sound optimization method and electronic device to solve the above technical defects of the prior art.
[0005] The technical solution adopted by the present application to solve the technical problem is: constructing an audio playing tail POP sound optimization method, comprising the following steps:
[0006] S1, presetting a breakpoint position of a to-be-played audio file;
[0007] S2, generating a replacement audio corresponding to the breakpoint position, and replacing the to-be-played audio file with the replacement audio within a first preset time length ending at the breakpoint position to generate a final playing audio, and / or inserting the replacement audio within a second preset time length starting at the breakpoint position to generate a final playing audio.
[0008] Preferably, in an embodiment of the audio playing tail POP sound optimization method of the present application, it comprises:
[0009] S21A, obtaining an audio data segment of the first preset time length in the to-be-played audio file with the breakpoint position as the ending point;
[0010] S22A, gradually reduce the amplitude of the audio signal corresponding to the audio data segment, so that the amplitude of the audio signal of the audio data segment at the breakpoint position is equal to a first preset value, to obtain a first replacement audio corresponding to the audio data segment;
[0011] S23A, replacing the corresponding audio data segment with the first replacement audio to obtain the final playback audio file.
[0012] Preferably, in an embodiment of the audio playback tail POP sound optimization method of the present application, the first preset value is less than the human hearing range or less than the original audio signal amplitude value.
[0013] Preferably, in an embodiment of the audio playback tail POP sound optimization method of the present application, the first preset time is less than 5 seconds and greater than or equal to 0.05 seconds.
[0014] Preferably, in an embodiment of the audio playback tail POP sound optimization method of the present application, it further comprises:
[0015] S21B, obtaining the amplitude of the audio signal corresponding to the breakpoint position;
[0016] S22B, generating a second replacement audio with gradually decreasing amplitude and lasting a second preset time from the audio signal amplitude as the starting point, wherein the tail audio of the second replacement audio is equal to a second preset value;
[0017] S23B, inserting the second replacement audio to generate the final playback audio within the second preset time after the breakpoint position.
[0018] Preferably, in an embodiment of the audio playback tail POP sound optimization method of the present application, the second preset value is less than the human hearing range or less than the original audio signal amplitude value.
[0019] Preferably, in an embodiment of the audio playback tail POP sound optimization method of the present application, the second preset time is less than 1 second and greater than or equal to 0.05 seconds.
[0020] Preferably, in an embodiment of the audio playback tail POP sound optimization method of the present application, it further comprises:
[0021] Obtaining the audio spectrum corresponding to the breakpoint position, and generating the second replacement audio from the audio spectrum.
[0022] Preferably, in an embodiment of the audio playback tail POP sound optimization method of the present application, the breakpoint position comprises the actual end position of the audio, the segmented end position of the audio and / or the preset pause position of the playback.
[0023] The application also discloses an electronic device, comprising a memory and a processor, wherein the memory is used for storing a computer program; and the processor is used for executing the computer program to realize the method.
[0024] The application discloses an audio playback tail POP sound optimization method and an electronic device. BRIEF DESCRIPTION OF DRAWINGS
[0025] The application will be further described below in combination with the drawings and embodiments, wherein:
[0026] Figure 1 is a program flowchart of the audio playback tail POP sound optimization method of the embodiment of the application;
[0027] Figure 2 is a program flowchart of another embodiment of the audio playback tail POP sound optimization method of the application;
[0028] Figure 3 is a program flowchart of another embodiment of the audio playback tail POP sound optimization method of the application. DETAILED DESCRIPTION
[0029] In order to have a clearer understanding of the technical features, purposes and effects of the application, the specific embodiments of the application will be described in detail with reference to the drawings.
[0030] As shown in FIG. 1, Figure 1 As shown in FIG. 1, Figure 1 In the embodiment shown in FIG. 1, the audio playback tail POP sound optimization method of the application comprises the following steps: S1, presetting a breakpoint position of a to-be-played audio file; S2, generating a replacement audio corresponding to the breakpoint position, and replacing the to-be-played audio file with the replacement audio within a first preset time length ending at the breakpoint position to generate a final playback audio, and / or inserting the replacement audio within a second preset time length starting at the breakpoint position to generate the final playback audio.
[0031] Based on step S1, the breakpoint position of the audio file to be played can be preset first. The presetting process of the breakpoint position can be based on the audio file itself, such as the end points of each segment in the audio file or the end point of the entire audio file, i.e. the actual end position of audio playing. The setting of the segments of the audio file can be obtained according to the specific content of the audio file, such as segmenting based on complete sentences or complete paragraphs in the audio file, for example, taking the time point corresponding to the end of a sentence in a large section of speech as a segment point. It can also be segmented based on natural pauses or theme changes in the audio content. It can also be directly segmented with special notes in the audio file as nodes, such as auxiliary phonemes such as "ah" and "hey" that have little effect on playing. In another embodiment, the user's playing habits for the audio file or the user's usage habits for the audio playing device can also be predicted, such as using AI or manual analysis of application scenarios, sound sources or semantic meanings in advance to obtain possible pause positions and then obtain the prediction result, i.e. the predicted position of playing pause. For example, by setting the audio playing device to play a complete audio file, the preset breakpoint position can be obtained based on the end point of the audio document. If the audio playing device is set to play at a fixed time, the preset breakpoint position can be obtained based on the end point of the fixed time.
[0032] In an embodiment, the breakpoint position can be preset before each playing of the audio file to be played. For example, before playing the audio file to be played, the breakpoint position can be preset according to the structure of the audio file, the settings of the audio playing device for playing the audio file, or the playing habits of the audio file or the usage habits of the audio playing device. The playing habits of the audio file can be the user's tendency to pause or stop at which positions of the audio file. The usage habits of the audio playing device can be the user's usage time of the audio playing device, such as the current time of the audio playing device being a first time point, and the user usually turns off the audio playing device at a second time point, so the breakpoint position of the audio file can be set based on the first time point and the second time point.
[0033] In another embodiment, the source file of the audio file to be played can also be permanently processed. The specific process can be to divide the audio file according to the file structure of the audio file. The setting of the segments of the audio file can be obtained according to the specific content of the audio file. For example, it can be segmented based on complete sentences or complete paragraphs in the audio file, it can also take the time point corresponding to the end of a sentence in a large section of speech as a segment point, it can also be segmented based on natural pauses or theme changes in the audio content, and it can also be directly segmented with special notes in the audio file as nodes, such as auxiliary phonemes such as "ah" and "hey" that have little effect on playing, to finally preset the breakpoint position of the audio file.
[0034] Based on step S2, and based on the preset breakpoint position, a replacement audio corresponding to that breakpoint position can be generated. In one embodiment, the duration of the replacement audio can be a first preset duration. Taking the breakpoint position as the end point, the first preset duration before the breakpoint position is obtained. Within this first preset duration, the replacement audio file replaces the audio file to be played, resulting in the final playback audio. Thus, during the playback of the audio to be played, when the audio reaches the position corresponding to the first preset duration before the breakpoint position, the replacement audio begins to play, and the pop sound is weakened or eliminated when the audio playback ends at the breakpoint position. In another embodiment, the duration of the replacement audio can be a second preset duration. Taking the breakpoint position as the starting point, the replacement audio is inserted within the second preset duration after the breakpoint position. When the audio to be played reaches the breakpoint position, the replacement audio begins to play within the second preset duration, so that the pop sound is weakened or eliminated when the audio ends. In one embodiment, a first preset duration and a second preset duration, along with their corresponding replacement audio, can be set simultaneously and reasonably, so that during audio playback, the replacement audio is played within both the first and second preset durations, so that no pop sound appears when the audio playback finally ends.
[0035] This process ensures that when a pause action is triggered within a preset time period before or after the breakpoint, i.e., when the pause or stop button is pressed within a specific time period before or after the breakpoint, the audio file will actually play up to that point during playback. Ultimately, it presets possible pause points and then processes the audio before and after these points; when the user stops or pauses playback, it doesn't stop immediately, but rather at the nearest preset time point.
[0036] Optional, such as Figure 2 As shown, in one embodiment of the audio playback tail pop sound optimization method of the present invention, the specific steps may further include: S21A, obtaining an audio data segment of a first preset duration with the breakpoint position as the end point; S22A, gradually reducing the audio signal amplitude corresponding to the audio data segment so that the audio signal amplitude of the audio data segment at the breakpoint position is equal to the first preset value, so as to obtain the first replacement audio corresponding to the audio data segment; S23A, replacing the corresponding audio data segment with the first replacement audio to obtain the final playback audio file.
[0037] After the preset breakpoint position is obtained based on step S1, an audio data segment with a first preset time length in the audio file to be played can be obtained with the preset breakpoint position as the end point. The amplitude of the audio signal corresponding to the audio data segment is gradually reduced from the start point of the audio data segment, so that the audio amplitude of the audio data segment at the breakpoint position is reduced to equal the first preset value. Finally, the first replacement audio corresponding to the audio data segment is obtained, which can also be understood as the replacement audio corresponding to the breakpoint position. The first replacement audio is used to replace the audio data segment to obtain the final playing audio. In the actual audio playing process, when the audio playing interruption occurs at a certain breakpoint position, the replacement audio is played before the breakpoint position. When the interruption occurs at the breakpoint position, a large POP sound will not be produced due to the playing interruption, so as to achieve the purpose of eliminating or weakening the POP sound in the audio playing process. When the actual interruption action occurs before or after the preset time at the breakpoint position, the audio playing interruption is considered to occur at the breakpoint position. It can also be understood that when the interruption does not occur at the corresponding breakpoint position in the audio playing process, the replacement audio corresponding to the breakpoint position is actually played, but this process does not affect the playing process of the entire audio file.
[0038] In an embodiment, the first preset value is less than the human hearing range or less than the original audio signal amplitude value. That is, when the amplitude of the audio signal of the audio data segment is reduced, the amplitude of the audio signal of the first replacement audio at the breakpoint position is reduced to less than the human hearing range, for example, less than 10 decibels, so that the human is in a near-silent state, and the POP sound is finally eliminated at the end of the playing. In an embodiment, it can also be achieved to reduce the audio signal amplitude as much as possible to be lower than the original audio signal amplitude value to weaken the POP sound.
[0039] In an embodiment, the first preset time length is less than 5 seconds and greater than or equal to 0.05 seconds. That is, when the audio data segment corresponding to the breakpoint position is selected, an audio data segment with a length between 0.05 seconds and 5 seconds can be selected. The first replacement audio with the same length is generated for replacement. Because the replacement audio is also normally played when no interruption occurs, when the replacement audio file is generated, the replacement audio should be as little distorted as possible to affect the playing process of the entire audio file. The length of the replacement audio can also be controlled to reduce the impact on the playing process of the entire audio file, thereby not affecting the user's use effect. The length can be set as short as possible while eliminating or weakening the POP sound.
[0040] Optionally, as Figure 3In an embodiment of the method for optimizing the POP sound at the end of audio playback, the specific steps can further include: S21B, obtaining the audio signal amplitude corresponding to the breakpoint position; S22B, generating a second replacement audio with gradually decreasing amplitude and a second preset duration from the audio signal amplitude as the starting point, wherein the tail audio of the second replacement audio is equal to a second preset value; and S23B, inserting the second replacement audio in the second preset duration after the breakpoint position to generate the final playback audio.
[0041] After obtaining the preset breakpoint position based on step S1, the audio signal amplitude corresponding to the breakpoint position is taken as the starting point to gradually reduce the amplitude to obtain a second replacement audio with a second preset duration. Meanwhile, the tail audio of the second replacement audio is gradually reduced to equal to a second preset value. The second replacement audio is inserted in the second preset duration after the breakpoint position to gradually reduce the audio sound to less than the original audio signal amplitude value after the breakpoint position in the audio playback process, and in an embodiment, the audio sound is finally reduced to silence to avoid generating a large POP sound. In the actual audio playback process, when the audio playback interruption occurs at a breakpoint position, the playback process of the replacement audio will be performed after the breakpoint position. When the interruption occurs at the breakpoint position, a large POP sound will not be generated due to the playback interruption, thereby achieving the purpose of finally eliminating or weakening the POP sound in the audio playback process. In addition, it can be understood that when no interruption occurs at the corresponding breakpoint position in the audio playback process, the replacement audio corresponding to the breakpoint position will actually be played, but this process does not affect the playback process of the entire audio file.
[0042] In an embodiment, the second preset value is less than the human hearing range or less than the original audio signal amplitude value. That is, when the audio signal amplitude of the audio data segment is reduced, the audio signal amplitude at the end of the second replacement audio is reduced as much as possible, and can even be reduced to less than the human hearing range, for example, less than 10 decibels, so that the human is in a near-silent state, and the POP sound at the end of the broadcast is finally realized. In an embodiment, the audio signal amplitude can also be reduced as much as possible to be lower than the original audio signal amplitude value to achieve the effect of weakening the POP sound.
[0043] In an embodiment, the second preset duration is less than 1 second and greater than or equal to 0.05 seconds. That is, when the second replacement audio is produced, an audio data segment with a length between 0.05 seconds and 1 second can be produced.
[0044] In an embodiment, in the embodiment of the tail POP sound optimization method of the audio playback, the specific steps further comprise: obtaining an audio spectrum corresponding to the breakpoint position, and generating the second replacement audio based on the audio spectrum. That is, when the second replacement audio is generated, the audio spectrum of the second replacement audio is the same as the audio spectrum of the breakpoint position. Because the replacement audio is to be normally played when no interruption occurs, when the replacement audio file is generated, the replacement audio should be prevented from being excessively distorted to affect the playback process of the entire audio file. The duration of the replacement audio can also be controlled to reduce the impact on the playback process of the entire audio file, thereby not affecting the use effect of the user. The duration can be set to be as short as possible while eliminating the POP sound.
[0045] Through the above embodiments, the final audio file of the to-be-played audio file can be finally obtained, and the tail POP sound of the to-be-played audio file in the audio playback process is eliminated or weakened by using the optimization method, thereby improving the user experience.
[0046] In addition, the electronic device of the present application can include a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program to realize the above method. Specifically, according to the embodiments of the present application, the processes described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments of the present application include a computer program product comprising a computer program carried on a computer readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such embodiments, the computer program can be downloaded and installed by the electronic device and executed to perform the above functions defined in the method of the embodiments of the present application. The electronic device in the present application can be a notebook, desktop, tablet computer, smart phone, etc. terminal, or a server.
[0047] That is, it can be understood that the above processes are executed by the electronic device to obtain the final audio file of the to-be-played audio file, and the POP sound in the playback process of the to-be-played audio file is eliminated.
[0048] It can be understood that the above embodiments only express the preferred embodiments of the present application, which are described in detail, but cannot be understood as a limitation on the scope of the patent of the present application; it should be pointed out that for ordinary skilled in the art, the above technical features can be freely combined without departing from the concept of the present application, and some modifications and improvements can be made, which all belong to the protection scope of the present application; therefore, any equivalent transformation and modification within the scope of the claims of the present application should belong to the scope of the claims of the present application.
Claims
1. A method for optimizing the pop sound at the end of audio playback, characterized in that, Includes the following steps: S1: Preset the breakpoint position of the audio file to be played; S2. Generate a replacement audio corresponding to the breakpoint position, and replace the audio file to be played with the replacement audio within a first preset duration starting from the breakpoint position to generate the final playback audio, and / or insert the replacement audio within a second preset duration starting from the breakpoint position to generate the final playback audio.
2. The method for optimizing the POP sound at the end of audio playback according to claim 1, characterized in that, The method includes: S21A. Using the breakpoint position as the end point, obtain the audio data segment of the first preset duration in the audio file to be played; S22A. Gradually reduce the amplitude of the audio signal corresponding to the audio data segment, so that the amplitude of the audio signal at the breakpoint of the audio data segment is equal to the first preset value, so as to obtain the first replacement audio corresponding to the audio data segment. S23A: Replace the corresponding audio data segment with the first replacement audio to obtain the final playback audio file.
3. The method for optimizing the POP sound at the end of audio playback according to claim 2, characterized in that, The first preset value is less than the range of human hearing or less than the amplitude value of the original audio signal.
4. The method for optimizing the POP sound at the end of audio playback according to claim 2, characterized in that, The first preset duration is less than 5 seconds and greater than or equal to 0.05 seconds.
5. The method for optimizing the POP sound at the end of audio playback according to claim 1, characterized in that, The method further includes: S21B. Obtain the amplitude of the audio signal corresponding to the breakpoint position; S22B: Generate a second replacement audio with a gradually decreasing amplitude and a duration of a second preset duration, starting from the amplitude of the audio signal, wherein the tail audio of the second replacement audio is equal to a second preset value; S23B, Insert the second replacement audio within a second preset duration after the breakpoint position to generate the final playback audio.
6. The method for optimizing the POP sound at the end of audio playback according to claim 5, characterized in that, The second preset value is less than the range of human hearing or less than the amplitude value of the original audio signal.
7. The method for optimizing the POP sound at the end of audio playback according to claim 5, characterized in that, The second preset duration is less than 1 second and greater than or equal to 0.05 seconds.
8. The method for optimizing the POP sound at the end of audio playback according to claim 5, characterized in that, The method further includes: Obtain the audio spectrum corresponding to the breakpoint position, and generate the second replacement audio using the audio spectrum.
9. The method for optimizing the POP sound at the end of audio playback according to claim 1, characterized in that, The breakpoint locations include the actual end point of the audio, the end point of a segment of the audio, and / or the predicted point where playback pauses.
10. An electronic device, characterized in that, Including memory and processor; The memory is used to store computer programs; The processor is used to execute the computer program to implement the method as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Method and device for realizing smooth rise and fall of sound volume
CN104683920A
Audio tail POP sound processing method and device
CN108182953A
Audio processing method and device, electronic equipment and storage medium
CN115914937A
Audio playing processing method and device, equipment, medium and product
CN120034787A
Apparatus for processing framed audio data for fade-in / fade-out effects
US20050234714A1