An audio processing method, device, intelligent device and storage medium

By performing beat analysis and filtering of audio files, deletion of inaccurate beat time points, and setting audio special effect data in the target audio data, the problem of inconsistent beats in the existing technology is solved, and convenient addition of special effect audio data is achieved, improving synthesis effect and user experience.

CN111862935BActive Publication Date: 2025-07-18TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202010732034.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-27
Publication Date
2025-07-18
Estimated Expiration
2040-07-27

AI Technical Summary

Technical Problem

When setting special effects audio data for music accompaniment, the direct splicing method cannot ensure that the beat of the audio file and the special effects audio data is consistent, resulting in poor synthesis effect, and users need to manually set the playback timeline, which is complex and inefficient.

Method used

By performing beat analysis on the audio file, filter out appropriate beat time points, delete inaccurate beat time points, and set audio special effect data in the target audio data to achieve convenient addition of special effect audio data.

Benefits of technology

It improves the rhythm and layering of synthetic audio files, simplifies the process of adding special-effect audio data, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111862935B_ABST
    Figure CN111862935B_ABST
Patent Text Reader

Abstract

An embodiment of the present application discloses an audio processing method, apparatus, intelligent device, and computer-readable storage medium. The method includes: performing beat analysis processing on the acquired target audio data to obtain a first beat time point sequence, screening the first beat time point sequence to obtain a second beat time point sequence, and setting audio effect data at at least one beat time point indicated by the second beat time point sequence in the target audio data to obtain special effect audio data corresponding to the target audio data. It can be seen that by screening the first beat time point sequence, inaccurate beat time points in the first beat time point sequence can be removed, and by setting audio effect data at at least one beat time point indicated by the second beat time point sequence in the target audio data, special effect audio data can be conveniently added to the audio file to be processed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and particularly to an audio processing method, device, intelligent device, and computer-readable storage medium. Background Art

[0002] With the continuous development of computer technology, many application scenarios involve the audio processing process. For example, when setting special effect audio data for a music accompaniment, the audio processing process will be involved. Currently, when setting special effect audio data for a music accompaniment, the direct splicing method is mostly used, that is, the audio file and the special effect audio data are directly merged.

[0003] It is found through practice that the synthesized audio file obtained by the direct splicing method cannot ensure the consistency of the beats between the audio file and the special effect audio data, and the effect is not good. In order to improve the synthesis effect of the synthesized audio file, the user needs to manually set the playback timeline of the special effect audio data, and the implementation method is complex and the efficiency is low. Summary of the Invention

[0004] Embodiments of the present invention provide an audio processing method, device, intelligent device, and computer-readable storage medium, which can conveniently add special effect audio data to an audio file.

[0005] On the one hand, an embodiment of the present application provides an audio processing method, which includes:

[0006] Performing beat analysis processing on the target audio data of the audio file to be processed to obtain a first beat time point sequence, where the first beat time point sequence includes beat time points corresponding to N beats, and N is a positive integer greater than or equal to 2;

[0007] Calculating the beat interval values between adjacent beats in the first beat time point sequence to obtain a beat interval sequence including a plurality of beat interval values; and screening out the beat interval values to be optimized from the beat interval sequence;

[0008] Deleting the beat time points corresponding to the beat interval values to be optimized in the first beat time point sequence to obtain a second beat time point sequence, where the second beat time point sequence includes M beat time points, M is a positive integer, and M≤N;

[0009] Setting audio special effect data at P beat time points indicated by the second beat time point sequence in the target audio data to obtain the special effect audio data corresponding to the target audio data, where P is a positive integer and P≤M.

[0010] On the one hand, the present application provides an audio processing device, and the processing device includes:

[0011] An acquisition unit for acquiring an audio file to be processed;

[0012] A processing unit for performing beat analysis processing on target audio data of the audio file to be processed to obtain a first beat time point sequence, the first beat time point sequence including beat time points corresponding to N beats, where N is a positive integer greater than or equal to 2; calculating beat interval values between adjacent beats in the first beat time point sequence to obtain a beat interval sequence including a plurality of beat interval values; and screening out beat interval values to be optimized from the beat interval sequence; deleting the beat time points corresponding to the beat interval values to be optimized in the first beat time point sequence to obtain a second beat time point sequence, the second beat time point sequence including M beat time points, where M is a positive integer and M ≤ N; setting audio effect data at P beat time points indicated by the second beat time point sequence in the target audio data to obtain special effect audio data corresponding to the target audio data, where P is a positive integer and P ≤ M.

[0013] On the one hand, the present application provides an intelligent device, including a processor, a memory, and a communication interface, the processor, the memory, and the communication interface are interconnected, wherein the memory is used for storing a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute the above audio processing method.

[0014] On the one hand, the present application provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above audio processing method is implemented.

[0015] On the one hand, the present application provides a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above audio processing method.

[0016] In the embodiments of the present application, first, the obtained target audio data is subjected to beat analysis processing to obtain the first beat time point sequence. Then, the time intervals between adjacent beat time points are calculated to obtain the beat interval sequence of the target audio data. And according to this beat interval sequence, appropriate beat time points are selected from the first beat sequence to obtain the second beat time point sequence. Next, audio effect data is set at at least one beat time point indicated by the second beat time point sequence in the target audio data to obtain the special effect audio data corresponding to the target audio data. It can be seen that by screening the first beat time point sequence, appropriate beat time points in the first beat time point sequence can be obtained. By setting audio effect data at at least one beat time point indicated by the second beat time point sequence in the target audio data, the special effect audio data can be conveniently added to the audio file to be processed. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.

[0018] Figure 1 It is a scene architecture diagram of audio processing provided by the embodiments of the present application;

[0019] Figure 2 It is a flowchart of an audio processing method provided by the embodiments of the present application;

[0020] Figure 3 It is a schematic diagram of a beat interval sequence provided by the embodiments of the present application;

[0021] Figure 4 It is a flowchart of another audio processing method provided by the embodiments of the present application;

[0022] Figure 5a It is a schematic diagram of the spectrum of a target audio provided by the embodiments of the present application;

[0023] Figure 5b It is a schematic diagram of the spectrum of another target audio provided by the embodiments of the present application;

[0024] Figure 5c It is a schematic diagram of splicing video data and splicing audio data provided by the embodiments of the present application;

[0025] Figure 6 It is a schematic diagram of the structure of an audio processing device provided by the embodiments of the present application;

[0026] Figure 7A schematic structural diagram of an intelligent device provided by an embodiment of the present application. Detailed implementation manners

[0027] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention.

[0028] The embodiments of the present application relate to artificial intelligence (AI), machine learning (ML), and cloud technology. Among them, AI uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.

[0029] AI technology is an interdisciplinary subject that involves a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, processing technologies for large application programs, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning. During the audio processing process, based on AI technology, the needs of the user can be obtained, and then the audio file to be processed or the video file including music can be determined, so as to achieve the purpose of adding special effects.

[0030] ML is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. ML is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. ML and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. In some embodiments of the present application, based on the processing method of deep learning, the beats of the audio can be tracked to obtain the sequence of beat time points of the audio.

[0031] The audio processing method of the present application can provide services for different users through cloud technology. Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data computing, storage, processing, and sharing. Cloud technology is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model, which can form a resource pool, be used as needed, and be flexible and convenient. Cloud computing technology will become an important support. The background services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites, and more portal websites. With the high development and application of the Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the background system for logical processing. Data at different levels will be processed separately, and various types of industry data require the support of a powerful system background, which can only be achieved through cloud computing.

[0032] Please refer to Figure 1 , Figure 1 which is a scenario architecture diagram for audio processing provided by an embodiment of the present application. As Figure 1 shown, the scenario architecture diagram includes terminal devices 101, terminal device 103, and terminal device 104, and server 102. Among them, terminal devices 101, 103, and 104 are devices used by users. Terminal devices may include, but are not limited to: smartphones (such as Android phones, iOS phones, etc.), tablets, portable personal computers, mobile Internet devices (abbreviated as MID), wired / wireless headphones, smart speakers, and other devices capable of outputting audio signals. The embodiments of the present invention are not limited thereto. Server 102 refers to a background device that can provide technical support for audio services to terminal devices 101, 103, and 104. Server 102 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. Terminal devices 101, 103, 104, and server 102 can be directly or indirectly connected through wired or wireless communication methods, and the present application does not limit this here.

[0033] Figure 1In the audio processing scenario shown, the audio processing process mainly includes: (1) obtaining target audio data and performing beat analysis processing on the target audio data (such as through a deep learning-based beat tracking method) to obtain a first beat time point sequence, where the first beat time point sequence includes at least two beat time points; (2) calculating the beat interval values of adjacent beats in the first beat time point sequence to obtain a beat interval sequence including multiple beat interval values, and screening out the beat interval values to be optimized from the beat interval sequence (for example, screening out the beat interval values that do not belong to the threshold interval); (3) deleting the beat time points in the first beat time point sequence corresponding to the beat interval values to be optimized (such as the beat time point 3 corresponding to the beat interval value 2) to obtain a second beat time point sequence; (4) setting audio effect data (such as adjusting the pitch, adding drum sounds, etc.) at at least one beat time point indicated by the second beat time point sequence in the target audio data to obtain the special effect audio data corresponding to the target audio data.

[0034] In the embodiment of the present application, when the terminal device recognizes through the user interface or AI that the user has a need to add special effects to an audio file or a video file, it sends the target audio data to the server 102. The server 102 performs beat analysis processing on the obtained target audio data to obtain a first beat time point sequence, which includes multiple time values, and calculates the beat interval sequence of the target audio data through the first beat time point sequence. The time interval sequence includes multiple values, such as the difference between adjacent time points. Then, the beat interval values to be optimized are screened out from the beat interval sequence, the beat time points in the first beat sequence corresponding to the beat interval values to be optimized are deleted to obtain a second beat time point sequence, and audio effect data is set at at least one beat time point indicated by the second beat time point sequence in the target audio data to obtain the special effect audio data corresponding to the target audio data. Finally, the server 102 can send the obtained special effect audio data to the corresponding terminal device. For example, it is sent to the terminal device 101, and the terminal device 101 can play the music with added special effects. For example, if the special effect data is drum sounds, the intelligent device sequentially adds drum sounds at the beat time points indicated by the second beat time point sequence, which can enhance the rhythm of the target audio data. It can be seen that by screening the first beat time point sequence, inaccurate beat time points in the first beat time point sequence can be removed. By setting audio effect data at at least one beat time point indicated by the second beat time point sequence in the target audio data, special effect audio data can be conveniently added to the audio file to be processed.

[0035] In some other architectures, when the terminal device recognizes through the user interface or AI that the user has a need to add special effects to an audio file or a video file, the target audio data can also be processed through the above audio processing process to obtain the special effect audio data corresponding to the target audio data.

[0036] Please refer to Figure 2 , Figure 2 which is a flowchart of an audio processing method provided by an embodiment of this application. This method can be executed by an intelligent device, which can specifically be Figure 1 the terminal device 101, terminal device 103, terminal device 104 or server 102 shown in

[0037] S201: The intelligent device performs beat analysis processing on the target audio data in the audio file to be processed, and obtains the first beat time point sequence. The first beat time point sequence includes the time points corresponding to N beats, and N is a positive integer greater than or equal to 2.

[0038] In one implementation, the intelligent device obtains the audio file to be processed, and performs beat analysis processing on the target audio data in the audio file to be processed, and obtains the first beat time point sequence. Among them, the audio file to be processed can be obtained from a terminal device or a server, or can be recorded by an audio recording device through an audio processing device, or can also be read from the memory of the audio processing device. The target audio data can be all the audio data in the audio file to be processed, or can be part of the audio data in the audio file to be processed. The beat analysis processing specifically refers to the audio beat tracking method, and this method can be a method based on dynamic programming, or a method based on frequency domain features, or a method based on deep learning. The first beat time point sequence is used to record the time point positions of each beat after the target audio data undergoes beat analysis processing; for example, assuming that after the audio data 1 undergoes beat analysis processing, beats 1 - 3 are obtained, the time point when beat 1 appears in the audio data 1 is 0 minutes and 11 seconds, the time point when beat 2 appears in the audio data 1 is 0 minutes and 15 seconds, and the time point when beat 3 appears in the audio data 1 is 0 minutes and 23 seconds, then the beat time point sequence of the audio data 1 is: {0:11, 0:15, 0:23}.

[0039] S202: The intelligent device calculates the beat interval values of each adjacent beat in the first beat time point sequence, obtains a beat interval sequence including multiple beat interval values, and screens out the beat interval values to be optimized from the beat interval sequence.

[0040] Figure 3 which is a schematic diagram of a beat interval sequence provided by an embodiment of this application. As Figure 3As shown, the selected segment of the target audio data includes beats 1 - 5, and the beat time point sequence of the target audio data is: {0:40, 0:42, 0:44, 0:46, 0:49}. Since beat 1 and beat 2 are adjacent beats, the beat interval value 1 is: 42 seconds - 40 seconds = 2 seconds; similarly, the beat interval value 2 can be calculated to be 2 seconds, the beat interval value 3 is 2 seconds, and the beat interval value 4 is 3 seconds. Therefore, the beat interval sequence of the target audio data is: {2 seconds, 2 seconds, 2 seconds, 3 seconds}. It can be understood that the above-mentioned beat time points are only for illustrative purposes. When the beat frequency is relatively high, the specific beat time points may be in milliseconds, and the beat interval value between two beats may be only a few hundred milliseconds or even dozens of milliseconds.

[0041] In one implementation, the intelligent device filters out the beat interval values from the beat interval sequence that do not belong to the beat interval value threshold range; for example, assuming the beat interval sequence 1 is: {9 seconds, 11 seconds, 8 seconds, 10 seconds, 13 seconds}, and the beat interval value threshold range is [9, 11] seconds, then the beat interval values to be optimized filtered by the intelligent device from the beat interval sequence 1 are 8 seconds and 13 seconds.

[0042] In another implementation, the value refers to the value after converting each beat interval value in the beat interval sequence to a standard unit. The intelligent device counts the number of times each value appears in the beat interval sequence and filters out the beat interval values with the number of occurrences less than the number threshold; for example, assuming the beat interval sequence 1 is: {9 seconds, 10 seconds, 9 seconds, 10 seconds, 11 seconds, 10 seconds, 13 seconds, 11 seconds}, and the number threshold is 2. Since the number of occurrences of 13 seconds is 1 < 2, the beat interval value to be optimized filtered by the intelligent device from the beat interval sequence 1 is 13 seconds.

[0043] S203: The intelligent device deletes the beat time points corresponding to the beat interval values to be optimized in the first beat time point sequence to obtain the second beat time point sequence. The second beat time point sequence includes M beat time points, where M is a positive integer and M ≤ N. It can be understood that when there are no beat time points corresponding to the beat interval values to be optimized in the first beat time point sequence, M = N; when there are beat time points corresponding to the beat interval values to be optimized in the first beat time point sequence, M < N.

[0044] In one implementation, the beat interval sequence includes N - 1 beat interval values. The i-th beat interval value in the beat interval sequence corresponds to the (i + 1)-th beat time point in the beat time point sequence, where i is a positive integer and i ≤ N - 1. For example, assuming the beat time point sequence includes 10 beat time points, then the beat interval sequence includes 10 - 1 = 9 beat interval values. The 3rd beat interval value in the beat interval sequence corresponds to the 4th beat time point in the beat time point sequence. For example, assuming the first beat time point sequence is {0:16, 0:18, 0:30, 0:45, 0:59}, and the beat time point corresponding to the beat interval value to be optimized is 0:18, then after deleting this beat time point, the obtained second beat time point sequence is {0:16, 0:30, 0:45, 0:59}.

[0045] S204: The intelligent device sets audio effect data at the P beat time points indicated by the second beat time point sequence in the target audio data to obtain the special effect audio data corresponding to the target audio data. P is a positive integer and P ≤ M. Setting the audio effect data includes: adding audio effect data, for example, adding drum beats, adding electronic sound effects (sound effects produced using electronic musical instruments and electronic music technology), adding environmental sound effects (such as adding echo effects), etc.; adjusting the audio effects in the target audio data, for example, raising / lowering the pitch, etc. By setting the audio effect data, the rhythm sense of the target audio data can be enhanced (such as adding drum beats), and the music style can also be transformed (such as adding electronic music effects), thereby improving the layering and listenability of the target audio data.

[0046] In one implementation, the intelligent device sets audio effect data at each time point indicated by the second beat time point sequence in the target audio data, that is, P = M, to obtain the special effect audio data corresponding to the target audio data. In another implementation, the intelligent device sets audio effect data at some of the time points indicated by the second beat time point sequence in the target audio data, that is, P < M, to obtain the special effect audio data corresponding to the target audio data; for example, the intelligent device sets audio effect data at the start time point and the end time point indicated by the second beat time point sequence in the target audio data, or for another example, the intelligent device randomly selects 1 out of every 2 time points indicated by the second beat time point sequence in the target audio data to set the audio effect data.

[0047] In the embodiments of the present application, first, the obtained target audio data is subjected to beat analysis processing to obtain a first beat time point sequence. Then, the time intervals between adjacent beat time points are calculated to obtain a beat interval sequence of the target audio data, and appropriate beat time points are selected from the first beat sequence according to the beat interval sequence to obtain a second beat time point sequence. Next, audio effect data is set at at least one beat time point indicated by the second beat time point sequence in the target audio data to obtain special effect audio data corresponding to the target audio data. It can be seen that by screening the first beat time point sequence, appropriate beat time points in the first beat time point sequence can be obtained. By setting audio effect data at at least one beat time point indicated by the second beat time point sequence in the target audio data, special effect audio data can be conveniently added to the audio file to be processed.

[0048] Please refer to Figure 4 , Figure 4 which is a flowchart of another audio processing method provided by the embodiments of the present application. This method can be executed by an intelligent device, which can specifically be Figure 1 the terminal device 101, terminal device 103, terminal device 104 or server 102 shown in

[0049] S401: The intelligent device obtains the audio file to be processed, and performs beat analysis processing on the target audio data in the audio file to be processed to obtain a first beat time point sequence.

[0050] S402: The intelligent device calculates the beat interval values of adjacent beats in the first beat time point sequence to obtain a beat interval sequence including multiple beat interval values.

[0051] For the specific implementation manners of step S401 and step S402, reference can be made to Figure 2 the implementation manners of step S201 and step S202 in

[0052] S403: The intelligent device counts the number of occurrences of each value in the beat interval sequence, and determines the value with the most occurrences in the beat interval sequence as the reference beat interval value of the target audio data. The value refers to the value after converting each beat interval value in the beat interval sequence into a standard unit. The beat interval values in the beat interval sequence can be the same or different. For example, assume that the beat interval sequence 1 includes beat interval values from beat interval value 1 to beat interval value 3. The unit of the value is seconds. The beat interval value 1 is 28 seconds, the beat interval value 2 is 1 minute, and the beat interval value 3 is 60 seconds. Then the value 1 corresponding to the beat interval value 1 is 28, the value 2 corresponding to the beat interval value 2 is 60, and the value 3 corresponding to the beat interval value 3 is 60. It can be seen that the value 2 is the same as the value 3, and 60 appears the most times in the beat interval sequence 1. Therefore, 60 seconds is determined as the reference beat interval value of the target audio data.

[0053] S404: The intelligent device filters out the beat interval values to be optimized from the beat interval sequence according to the reference beat interval value. The intelligent device obtains a beat filtering threshold, which can be a preset empirical value or a value defined by the user. The intelligent device performs arithmetic processing on the i-th beat interval value and the reference beat interval value to obtain an interval error value, where i is a positive integer and i ≤ N - 1. The specific calculation formula for the interval error value is:

[0054] Formula 1: Interval error value 1 =

[0055] Formula 2: Interval error value 2 =

[0056] Formula 3: Interval error value 3 =

[0057] In the above Formulas 1 - 3, represents the reference beat interval value, represents the i-th beat interval value.

[0058] In one implementation, if the interval error values 1 - 3 calculated through Formulas 1 - 3 are all greater than the beat filtering threshold, the intelligent device determines the i-th beat interval value as the beat interval value to be optimized. For example, assume that the beat filtering threshold is 2 seconds, the reference beat interval value is 10 seconds, the first beat interval value is 5 seconds, the second beat interval value is 21 seconds, and the third beat interval value is 17 seconds. Since the interval error value 1 calculated through Formula 1 for the third beat interval value is 7 seconds > 2 seconds, the interval error value 2 calculated through Formula 2 is 3 seconds > 2 seconds, and the interval error value 3 calculated through Formula 3 is 12 seconds > 3 seconds, the intelligent device determines the third beat interval value as the beat interval value to be optimized.

[0059] S405: The intelligent device deletes the beat time points corresponding to the beat interval value to be optimized in the first beat time point sequence, obtaining a second beat time point sequence.

[0060] For the specific implementation manner of step S405, reference can be made to Figure 2 the implementation manner of step S203 in , which will not be elaborated here.

[0061] S406: The intelligent device calculates the number of beats X of the target audio per unit time according to the reference beat interval value of the target audio data. In one implementation manner, the unit time is minute, and the unit of the reference beat interval value is second, that is, the number of beats X is the number of beats of the target audio within 1 minute. The specific calculation formula of X is: X = 60 / .

[0062] S407: The intelligent device determines the type of audio effect data to be added according to the threshold interval to which X belongs. Among them, the type of audio effect data is determined according to sound parameters, and the sound parameters include one or more of the following: timbre, pitch, volume, and duration.

[0063] In one implementation manner, if X is less than or equal to the first threshold, the type of audio effect data to be added is determined as the first type; if X is greater than the first threshold and less than or equal to the second threshold, the type of audio effect data to be added is determined as the second type; if X is greater than the second threshold, the type of audio effect data to be added is determined as the third type. For example, the first threshold is 60, and the second threshold is 100. If X ≤ 60, the type of audio effect data to be added is determined as the slow type, and the gentle and soothing rhythm can make the listener calm down; if 60 < X ≤ 100, the type of audio effect data to be added is determined as the medium type, and the moderate rhythm can stimulate the listener's imagination; if 100 < X, the type of audio effect data to be added is determined as the fast type, and the fast and exciting rhythm can convey positive emotions to the listener.

[0064] Among them, the first threshold is less than or equal to the second threshold, and the sound parameters of the audio effect data of the first type, the second type, and the third type are different; for example, the pitch of the audio effect data of the first type is lower than the pitch of the audio effect data of the second type, and the pitch of the audio effect data of the second type is lower than the pitch of the audio effect data of the third type; another example is that the sound parameters of the audio effect data of the first type only include timbre, and the audio effect data of the first type is the audio effect data with a sharp timbre, the sound parameters of the audio effect data of the second type only include pitch, and the audio effect data of the second type is the audio effect data with a pitch lower than C key, and the sound parameters of the audio effect data of the third type include timbre and pitch, and the audio effect data of the third type is the audio effect data with a deep timbre and a pitch higher than C key.

[0065] S408: The intelligent device selects the audio effect data corresponding to the determined type of audio effect data to be added from the audio effect data set. In one implementation, the first type is the slow type, and the audio effect data corresponding to the slow type has a deep timbre, a small volume, and a long duration; for example, the audio effect data corresponding to the slow type is a bass drum. The second type is the medium speed type, and the audio effect data corresponding to the medium speed type has a moderate timbre, a moderate volume, and a certain duration; for example, the audio effect data corresponding to the medium speed type is a snare drum. The third type is the fast type, and the audio effect data corresponding to the fast type has a sharp timbre, a high pitch, a large volume, and an extremely short duration; for example, the audio effect data corresponding to the fast type is a tom drum.

[0066] S409: The intelligent device sets the selected audio effect data at the P beat time points indicated by the second beat time point sequence in the target audio data. P is a positive integer, and P ≤ M.

[0067] In one implementation, the intelligent device sequentially fills the audio effect data (such as adding drum sounds, electronic sounds, etc.) at the P beat time points indicated by the second beat time point sequence in the target audio data to obtain the special effect audio data corresponding to the target audio data.

[0068] In another implementation, setting the audio effect data means adjusting the target audio data (such as adjusting the pitch of the target audio data), that is, no new audio effect data is added. It can be understood that since no new audio effect data needs to be added to the target audio data, after performing step S407, step S409 can be directly executed. Specifically, the intelligent device sequentially adjusts the audio data at the P beat time points indicated by the second beat time point sequence in the target audio data to obtain the special effect audio data corresponding to the target audio data.

[0069] It can be seen that steps S406 - S409 divide the audio effect data from the rhythm dimension. The intelligent device determines the corresponding audio effect data according to the number of beats per unit time (i.e., the beat frequency) of the target audio data, thereby enhancing the rhythm sense of the target audio data.

[0070] S410: The intelligent device obtains the first special effect type corresponding to the human voice audio data and the second special effect type corresponding to the accompaniment audio data. In one implementation, the target audio data includes human voice audio data and accompaniment audio data. The second beat time point sequence includes j beat time points within the range of the human voice audio data and k beat time points not within the range of the human voice audio data, where j and k are positive integers, and j + k = M. The first special effect type corresponding to the human voice audio data and the second special effect type corresponding to the accompaniment audio data can be preset according to experience or customized by the user.

[0071] S411: The intelligent device selects the human voice audio special effect data corresponding to the first special effect type from the audio special effect data set, and selects the accompaniment audio special effect data corresponding to the second special effect type from the audio special effect data set. For example, the human voice audio special effect data corresponding to the first special effect type is the drum sound, and the accompaniment audio special effect data corresponding to the second special effect type is the electronic special effect sound.

[0072] S412: The intelligent device sets the human voice audio special effect data at at least one time point indicated by the j beat time points in the second beat time point sequence. In one implementation, the intelligent device sequentially fills the audio special effect data (such as adding drum sounds, electronic sounds, etc.) at at least one beat time point indicated by the j beat time points within the range of the human voice audio data in the target audio data.

[0073] In another implementation, setting the human voice audio special effect data means adjusting the human voice audio data (such as adjusting the pitch of the human voice audio data), that is, no new audio special effect data is added. It can be understood that since no new audio special effect data needs to be added to the human voice audio data, after step S410 is executed, step S412 can be directly executed. Specifically, the intelligent device sequentially adjusts the audio data at at least one beat time point indicated by the j beat time points within the range of the human voice audio data in the target audio data.

[0074] S413: The intelligent device sets the accompaniment audio special effect data at at least one time point indicated by the k beat time points in the second beat time point sequence to obtain the special effect audio data corresponding to the target audio data.

[0075] The specific implementation manner of step S413 is similar to that of step S412 and will not be elaborated here. It should be noted that step S413 can be executed before step S412 or can be executed simultaneously with step S412. After both step S412 and step S413 are executed, the special effect audio data corresponding to the target audio data can be obtained.

[0076] Figure 5a A spectrum schematic diagram of a target audio provided by an embodiment of this application, such asFigure 5a As shown, the spectrum schematic diagram includes a total of 77 beats, among which 43 beats are within the range of human voice audio data, and 34 beats are within the range of accompaniment audio data. Figure 5b This is another spectrum schematic diagram of the target audio provided by the embodiment of the present application. As Figure 5b shown, beats 1 - 11 are obtained after the intelligent device filters the 77 beats in Figure 5a (that is, deletes the beat time points corresponding to the interval values of the beats to be optimized in the first beat time point sequence). Among them, beats 3 - 5 and beats 8 - 10 are within the range of human voice audio data; beats 1, 2, 6, 7, and 11 are within the range of accompaniment audio data. The intelligent device sets the human voice audio special effect data at at least one time point indicated by beats 3 - 5 and beats 8 - 10, and sets the accompaniment audio special effect data at at least one time point indicated by beats 1, 2, 6, 7, and 11, to obtain the special effect audio data corresponding to the target audio data.

[0077] It can be seen that steps S410 - S413 divide the target audio data according to whether the target audio data contains human voice audio data. The intelligent device respectively sets the corresponding audio special effect data for the human voice audio data and the accompaniment audio data in the target audio data, thereby enhancing the layering of the target audio data.

[0078] S414: The intelligent device determines the target video data in the video file to be processed. The audio file to be processed is extracted from the video file to be processed. The intelligent device determines the target video data corresponding to the target audio data in the video file to be processed.

[0079] S415: The intelligent device sets the video special effect data in the video playback time intervals corresponding to the Q beat time points indicated by the second beat time point sequence in the target video data, to obtain the special effect video data corresponding to the target video data. The video playback time interval corresponding to the beat time point refers to a period of time interval including this beat time point; for example, if the beat time point 1 is 31 seconds, the video playback time interval corresponding to the beat time point 1 is 30 seconds 950 milliseconds - 31 seconds 50 milliseconds.

[0080] Setting the video special effect data includes: processing the target video data (such as rotating, adjusting the brightness of the target video data, etc.) and adding video special effect data (such as adding stickers, adding text, etc.). The specific processing method of the target video data can be preset or configured by the user. Similarly, the added video special effect data can be obtained from the video special effect material library or configured by the user.

[0081] It can be seen that by performing step S414 and step S415, the intelligent device can set video special effect data for the target video data corresponding to the target audio data according to the second beat time point sequence of the target audio data, thereby enhancing the user's viewing experience. It should be noted that step S414 and step S415 can be executed in parallel with step S406-step S409, step S410-step S413, and step S416-step S418, or can be executed after the above steps.

[0082] S416: The intelligent device obtains the splicing rule of the spliced audio data. The audio file to be processed includes spliced audio data, and the spliced audio data is obtained by splicing at least two target audio data. The splicing rule is used to indicate the splicing order and splicing time points of each target audio data; for example, splicing rule 1 is used to indicate that target audio data 2 is spliced after target audio data 1, and the splicing time point is 1 minute and 30 seconds.

[0083] S417: The intelligent device determines the splicing time points in the spliced audio data according to the splicing rule. In one implementation, the intelligent device obtains the splicing time points in the spliced audio data from the splicing rule.

[0084] S418: The intelligent device sets audio special effect data at at least one time point indicated by the splicing time point in the spliced audio data to obtain the special effect audio data corresponding to the spliced audio data.

[0085] The specific implementation manner of step S418 is similar to that of step S409 and will not be elaborated here.

[0086] Figure 5c It is a schematic diagram of spliced video data and spliced audio data provided by an embodiment of the present application. As Figure 5c shown, the spliced video data is spliced from target video data 1-target video data 3, and the spliced video data includes two splicing time points; the spliced audio data is spliced from target audio data 1-target audio data 3, and the spliced audio data includes two splicing time points, and target video data 1 corresponds to target audio data 1, target video data 2 corresponds to target audio data 2, and target video data 3 corresponds to target audio data 3. The intelligent device sets audio special effect data at at least one time point indicated by the splicing time point in the spliced audio data to obtain the special effect audio data corresponding to the spliced audio data. Correspondingly, the intelligent device sets video special effect data at at least one time point indicated by the splicing time point in the spliced video data to obtain the special effect audio data corresponding to the spliced video data.

[0087] It can be seen that by executing step S416-step S418, the intelligent device can determine the splicing time point of the spliced audio according to the splicing rule, and set the audio effect data for the spliced audio data according to the splicing time point, so that the transition of the spliced audio is more natural and the audibility of the spliced audio is improved.

[0088] In the embodiments of the present application, on the Figure 2 basis of the embodiment, the audio effect data is divided from the rhythm dimension to enhance the rhythm of the target audio data; the target audio data is divided according to whether the target audio data includes human voice audio data to enhance the layering of the target audio data; the video effect data is set for the target video data corresponding to the target audio data according to the second beat time point sequence of the target audio data to improve the user's viewing experience; the audio effect data is set for the spliced audio data according to the splicing time point, so that the transition of the spliced audio is more natural and the audibility of the spliced audio is improved.

[0089] The above details the method of the embodiments of the present application. To facilitate better implementation of the above solutions of the embodiments of the present application, correspondingly, the device of the embodiments of the present application is provided below.

[0090] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of an audio processing device provided by an embodiment of the present application. The device can be mounted on the intelligent device in the above method embodiment. The intelligent device can specifically be Figure 1 the terminal device 101, terminal device 103, terminal device 104 or server 102 shown in Figure 6 The shown audio processing device can be used to execute some or all of the functions described in the above Figure 2 and Figure 4 method embodiments. Among them, the detailed description of each unit is as follows:

[0091] The obtaining unit 601 is used to obtain the audio file to be processed;

[0092] The processing unit 602 is used to perform beat analysis processing on the target audio data of the audio file to be processed to obtain the first beat time point sequence, where the first beat time point sequence includes the beat time points corresponding to N beats, and N is a positive integer greater than or equal to 2;

[0093] Calculate the beat interval values between adjacent beats in the first beat time point sequence to obtain a beat interval sequence including multiple beat interval values; and screen out the beat interval values to be optimized from the beat interval sequence; delete the beat time points corresponding to the beat interval values to be optimized in the first beat time point sequence to obtain a second beat time point sequence, where the second beat time point sequence includes M beat time points, M is a positive integer, and M ≤ N; set audio effect data at the P beat time points indicated by the second beat time point sequence in the target audio data to obtain the special effect audio data corresponding to the target audio data, P is a positive integer, and P ≤ M.

[0094] In one embodiment, the processing unit 602 is specifically configured to: screen out the beat interval values to be optimized from the beat interval sequence;

[0095] Count the number of occurrences of each value in the beat interval sequence, and determine the value with the most occurrences in the beat interval sequence as the reference beat interval value of the target audio data;

[0096] Screen out the beat interval values to be optimized from the beat interval sequence according to the reference beat interval value.

[0097] In one embodiment, the processing unit 602 is specifically configured to: screen out the beat interval values to be optimized from the beat interval sequence according to the reference beat interval value;

[0098] Obtain a beat screening threshold;

[0099] Process the i-th beat interval value and the reference beat interval value to obtain an interval error value, where i is a positive integer and i is less than or equal to N - 1;

[0100] If the interval error value is greater than the screening threshold, then determine the i-th beat interval value as the beat interval value to be optimized.

[0101] In one embodiment, the processing unit 602 is specifically configured to: set audio effect data at the P beat time points indicated by the second beat time point sequence in the target audio data;

[0102] Calculate the number of beats X per unit time of the target audio according to the reference beat interval value of the target audio data, where X is a positive number;

[0103] Determine the type of the audio effect data to be added according to the threshold interval to which X belongs, and the type of the audio effect data is determined according to sound parameters;

[0104] Select the audio effect data corresponding to the determined type of the audio effect data to be added from the audio effect data set;

[0105] Set the selected audio effect data at the P beat time points indicated by the second beat time point sequence in the target audio data.

[0106] In one embodiment, the processing unit 602 is specifically configured to: determine the type of audio effect data to be added according to the threshold interval to which X belongs;

[0107] If X is less than or equal to the first threshold, determine the type of audio effect data to be added as the first type;

[0108] If X is greater than the first threshold and less than or equal to the second threshold, determine the type of audio effect data to be added as the second type;

[0109] If X is greater than the second threshold, determine the type of audio effect data to be added as the third type;

[0110] Wherein, the first threshold is less than or equal to the second threshold, and the sound parameters of the three types of audio effect data are different from each other. The sound parameters include one or more of the following: timbre, pitch, volume, and duration.

[0111] In one embodiment, the target audio data includes human voice audio data and accompaniment audio data. The second beat time point sequence includes j beat time points within the range of the human voice audio data and k beat time points not within the range of the human voice audio data. j and k are positive integers, and j + k = M. The processing unit 602 is specifically configured to: set audio effect data at the P beat time points indicated by the second beat time point sequence in the target audio data to obtain the special effect audio data corresponding to the target audio data;

[0112] Obtain the first special effect type corresponding to the human voice audio data and the second special effect type corresponding to the accompaniment audio data;

[0113] Select the human voice audio effect data corresponding to the first special effect type from the audio effect data set and select the accompaniment audio effect data corresponding to the second special effect type from the audio effect data set;

[0114] Set the human voice audio effect data at at least one time point indicated by the j beat time points in the second beat time point sequence;

[0115] Set the accompaniment audio effect data at at least one time point indicated by the k beat time points in the second beat time point sequence to obtain the special effect audio data corresponding to the target audio data.

[0116] In one embodiment, the audio file to be processed is extracted from a video file to be processed; the processing unit 602 is further configured to:

[0117] Determine target video data in the video file to be processed;

[0118] Set video special effect data in a video playback time interval corresponding to Q beat time points indicated by the second beat time point sequence in the target video data to obtain special effect video data corresponding to the target video data, where Q is a positive integer and Q is less than or equal to M.

[0119] In one embodiment, the audio file to be processed includes spliced audio data, and the spliced audio data is obtained by splicing at least two target audio data; the processing unit 602 is further configured to:

[0120] Obtain a splicing rule of the spliced audio data;

[0121] Determine splicing time points in the spliced audio data according to the splicing rule;

[0122] Set audio special effect data at at least one time point indicated by the splicing time point in the spliced audio data to obtain special effect audio data corresponding to the spliced audio data.

[0123] According to an embodiment of the present application, Figure 2 and Figure 4 Some steps involved in the audio processing method shown can be executed by each unit in the Figure 6 audio processing device shown. For example, Figure 2 step S201 shown in can be executed by the Figure 6 acquisition unit 601 shown, and steps S202 - S204 can be executed by the Figure 6 processing unit 602 shown. Figure 4 Step S401, step S410, and step S416 shown in can be executed by the Figure 6 acquisition unit 601 shown, and steps S402 - S409, steps S411 - S415, step S417, and step S418 can be executed by the Figure 6 processing unit 602 shown. Figure 6Each unit in the audio processing device shown can be separately or all combined into one or several other units to form, or a certain (some) unit among them can also be further split into multiple smaller units with more specific functions to form. This can achieve the same operations without affecting the realization of the technical effects of the embodiments of this application. The above units are divided based on logical functions. In practical applications, the function of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of this application, the audio processing device can also include other units. In practical applications, these functions can also be assisted by other units and can be realized through the cooperation of multiple units.

[0124] According to another embodiment of this application, it can be achieved by running a computer program (including program code) that can execute the respective steps involved in the corresponding methods shown in Figure 2 and Figure 4 on a general computing device such as a computer that includes processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read-only storage medium (ROM), to construct an audio processing device as shown in Figure 6 and to implement the audio processing method of the embodiments of this application. The computer program can be recorded on, for example, a computer-readable recording medium, loaded into the above computing device through the computer-readable recording medium, and run therein.

[0125] Based on the same inventive concept, the principle of problem-solving and the beneficial effects of the audio processing device provided in the embodiments of this application are similar to those of the audio processing device in the method embodiments of this application. For the principle and beneficial effects of the method implementation, reference can be made thereto. For the sake of brevity, they will not be elaborated here.

[0126] Please refer to Figure 7 , Figure 7The figure shows a schematic structural diagram of an intelligent device provided by an embodiment of the present application. The intelligent device at least includes a processor 701, a communication interface 702, and a memory 703. Among them, the processor 701, the communication interface 702, and the memory 703 can be connected through a bus or other means. Among them, the processor 701 (or Central Processing Unit, CPU) is the computing core and control core of the terminal. It can parse various instructions in the terminal and process various data of the terminal. For example, the CPU can be used to parse the power-on and power-off instructions sent by the user to the terminal and control the terminal to perform power-on and power-off operations. Another example is that the CPU can transmit various interactive data between the internal structures of the terminal, and so on. The communication interface 702 may optionally include a standard wired interface, a wireless interface (such as WI-FI, a mobile communication interface, etc.), and under the control of the processor 701, it can be used to send and receive data; the communication interface 702 can also be used for the transmission and interaction of internal data of the terminal. The memory 703 (Memory) is the memory device in the terminal, used to store programs and data. It can be understood that the memory 703 here can include both the built-in memory of the terminal and, of course, the extended memory supported by the terminal. The memory 703 provides a storage space, and this storage space stores the operating system of the terminal, which may include but is not limited to: Android system, iOS system, Windows Phone system, etc. The present application does not make any limitations in this regard.

[0127] In the embodiment of the present application, the processor 701 is configured to perform the following operations by running the executable program code in the memory 703:

[0128] Obtain an audio file to be processed through the communication interface 702;

[0129] Perform beat analysis processing on the target audio data of the audio file to be processed to obtain a first beat time point sequence. The first beat time point sequence includes beat time points corresponding to N beats, and N is a positive integer greater than or equal to 2;

[0130] Calculate the beat interval values of adjacent beats in the first beat time point sequence to obtain a beat interval sequence including multiple beat interval values; and screen out the beat interval values to be optimized from the beat interval sequence;

[0131] Delete the beat time points corresponding to the beat interval values to be optimized in the first beat time point sequence to obtain a second beat time point sequence. The second beat time point sequence includes M beat time points, M is a positive integer, and M ≤ N;

[0132] Set audio effect data at P beat time points indicated by the second beat time point sequence in the target audio data to obtain the special effect audio data corresponding to the target audio data, where P is a positive integer and P ≤ M.

[0133] As an optional embodiment, the processor 701 is configured to:

[0134] Count the number of occurrences of each value in the beat interval sequence, and determine the reference beat interval value of the target audio data as the value with the most occurrences in the beat interval sequence;

[0135] Filter out the beat interval values to be optimized from the beat interval sequence according to the reference beat interval value.

[0136] As an optional embodiment, a specific example of the processor 701 filtering out the beat interval values to be optimized from the beat interval sequence according to the reference beat interval value is:

[0137] Obtain a beat screening threshold;

[0138] Process the i-th beat interval value and the reference beat interval value to obtain an interval error value, where i is a positive integer and i ≤ N - 1;

[0139] If the interval error value is greater than the screening threshold, then determine the i-th beat interval value as the beat interval value to be optimized.

[0140] As an optional embodiment, the processor 701 is configured to:

[0141] Calculate the number of beats X per unit time of the target audio according to the reference beat interval value of the target audio data, where X is a positive number;

[0142] Determine the type of the audio effect data to be added according to the threshold interval to which X belongs, and the type of the audio effect data is determined according to the sound parameters;

[0143] Select the audio effect data corresponding to the determined type of the audio effect data to be added from the audio effect data set;

[0144] Set the selected audio effect data at P beat time points indicated by the second beat time point sequence in the target audio data.

[0145] As an optional embodiment, the processor 701 is configured to:

[0146] If X is less than or equal to the first threshold, then determine the type of the audio effect data to be added as the first type;

[0147] If X is greater than the first threshold and less than or equal to the second threshold, determine the type of the audio effect data to be added as the second type;

[0148] If X is greater than the second threshold, determine the type of the audio effect data to be added as the third type;

[0149] Wherein, the first threshold is less than or equal to the second threshold, and the sound parameters of the three types of audio effect data are different from each other. The sound parameters include one or more of the following: timbre, pitch, volume, and duration.

[0150] As an optional embodiment, the target audio data includes human voice audio data and accompaniment audio data. The second beat time point sequence includes j beat time points within the range of the human voice audio data and k beat time points not within the range of the human voice audio data, where j and k are positive integers, and j + k = M. The processor 701 is configured to:

[0151] Obtain a first special effect type corresponding to the human voice audio data and a second special effect type corresponding to the accompaniment audio data;

[0152] Select human voice audio effect data corresponding to the first special effect type from the audio effect data set, and select accompaniment audio effect data corresponding to the second special effect type from the audio effect data set;

[0153] Set the human voice audio effect data at at least one time point indicated by the j beat time points in the second beat time point sequence;

[0154] Set the accompaniment audio effect data at at least one time point indicated by the k beat time points in the second beat time point sequence to obtain special effect audio data corresponding to the target audio data.

[0155] As an optional embodiment, the audio file to be processed is extracted from a video file to be processed. The processor 701 is further configured to:

[0156] Determine target video data in the video file to be processed;

[0157] Set video special effect data in a video play time interval corresponding to Q beat time points indicated by the second beat time point sequence in the target video data to obtain special effect video data corresponding to the target video data, where Q is a positive integer and Q is less than or equal to M.

[0158] As an optional embodiment, the audio file to be processed includes spliced audio data, and the spliced audio data is obtained by splicing at least two target audio data. The processor 701 is further configured to:

[0159] Obtain the splicing rule of the spliced audio data;

[0160] Determine the splicing time points in the spliced audio data according to the splicing rule;

[0161] Set audio effect data at at least one time point indicated by the splicing time point in the spliced audio data to obtain the special effect audio data corresponding to the spliced audio data.

[0162] Based on the same inventive concept, the principle of problem-solving and the beneficial effects of the intelligent device provided in the embodiments of the present application are similar to the principle of problem-solving and the beneficial effects of the audio processing method in the method embodiments of the present application. For the principle and beneficial effects of the method implementation, refer to them. For the sake of concise description, they will not be elaborated here.

[0163] The embodiments of the present application further provide a computer-readable storage medium, in which one or more instructions are stored, and the one or more instructions are adapted to be loaded and executed by a processor to perform the audio processing method described in the above method embodiments.

[0164] The embodiments of the present application further provide a computer program product containing instructions, which, when running on a computer, causes the computer to execute the audio processing method described in the above method embodiments.

[0165] The embodiments of the present application further provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above audio processing method.

[0166] The steps in the method embodiments of the present application can be adjusted, combined, and deleted according to actual needs.

[0167] The modules in the device embodiments of the present application can be combined, divided, and deleted according to actual needs.

[0168] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program. The program can be stored in a computer-readable storage medium, and the readable storage medium can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0169] The above-disclosed is only a preferred embodiment of the present application. Of course, it cannot be used to limit the scope of rights of the present application. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present application still fall within the scope covered by the invention.

Claims

1. An audio processing method, characterized in that, The method includes: Performing beat analysis processing on the target audio data of the audio file to be processed, obtaining a first beat time point sequence, where the first beat time point sequence includes beat time points corresponding to N beats, and N is a positive integer greater than or equal to 2; the target audio data includes human voice audio data and accompaniment audio data; Calculating the beat interval values of adjacent beats in the first beat time point sequence, obtaining a beat interval sequence including multiple beat interval values; and screening out the beat interval values to be optimized from the beat interval sequence; Deleting the beat time points corresponding to the beat interval values to be optimized in the first beat time point sequence, obtaining a second beat time point sequence, where the second beat time point sequence includes M beat time points, M is a positive integer, and M≤N; the second beat time point sequence includes j beat time points within the range of the human voice audio data and k beat time points not within the range of the human voice audio data, j and k are positive integers, and j + k = M; Obtaining a first special effect type corresponding to the human voice audio data and a second special effect type corresponding to the accompaniment audio data; selecting the accompaniment audio special effect data corresponding to the second special effect type from the audio special effect data set; Adjusting the audio data corresponding to at least one time point indicated by the j beat time points in the second beat time point sequence, and setting the accompaniment audio special effect data at at least one time point indicated by the k beat time points in the second beat time point sequence, obtaining the special effect audio data corresponding to the target audio data, where the audio data adjustment performed includes adjusting the pitch; Among them, the screening out of the beat interval values to be optimized from the beat interval sequence includes: Counting the number of occurrences of each value in the beat interval sequence, and determining the reference beat interval value of the target audio data as the value with the most occurrences in the beat interval sequence; Obtaining a beat screening threshold; processing the i-th beat interval value and the reference beat interval value to obtain an interval error value, where i is a positive integer and i is less than or equal to N - 1; among them, the interval error value of the i-th beat interval value includes: the interval error value determined according to the absolute value of the difference between the reference beat interval value and the i-th beat interval value, the interval error value determined according to the absolute value of the difference between twice the reference beat interval value and the i-th beat interval value, and the interval error value determined according to the absolute value of the difference between half of the reference beat interval value and the i-th beat interval value; If the interval error values are all greater than the screening threshold, then determining the i-th beat interval value as the beat interval value to be optimized.

2. The method according to claim 1, wherein The audio file to be processed is extracted from the video file to be processed, and the method further includes: Determining the target video data in the video file to be processed; Set video special effect data on the video playback time intervals corresponding to the Q beat time points indicated by the second beat time point sequence in the target video data, to obtain the special effect video data corresponding to the target video data, where Q is a positive integer and Q ≤ M.

3. The method according to claim 1, characterized in that, The audio file to be processed includes spliced audio data, which is obtained by splicing at least two target audio data. The method further includes: Obtain the splicing rule of the spliced audio data; Determine the splicing time points in the spliced audio data according to the splicing rule; Set audio special effect data at at least one time point indicated by the splicing time points in the spliced audio data, to obtain the special effect audio data corresponding to the spliced audio data.

4. An audio processing device, characterized in that, including: An acquisition unit, configured to acquire an audio file to be processed; A processing unit, configured to perform beat analysis processing on the target audio data of the audio file to be processed, to obtain a first beat time point sequence, where the first beat time point sequence includes the beat time points corresponding to N beats, and N is a positive integer greater than or equal to 2; the target audio data includes human voice audio data and accompaniment audio data; calculate the beat interval values between adjacent beats in the first beat time point sequence, to obtain a beat interval sequence including a plurality of beat interval values; And screen out the beat interval values to be optimized from the beat interval sequence; delete the beat time points corresponding to the beat interval values to be optimized in the first beat time point sequence, to obtain a second beat time point sequence, where the second beat time point sequence includes M beat time points, and M is a positive integer and M ≤ N; the second beat time point sequence includes j beat time points within the range of the human voice audio data and k beat time points not within the range of the human voice audio data, j and k are positive integers, and j + k = M; The processing unit is further configured to obtain a first special effect type corresponding to the human voice audio data and a second special effect type corresponding to the accompaniment audio data; select the accompaniment audio special effect data corresponding to the second special effect type from the audio special effect data set; Adjust the audio data corresponding to at least one time point indicated by the j beat time points in the second beat time point sequence, and set the accompaniment audio special effect data at at least one time point indicated by the k beat time points in the second beat time point sequence, to obtain the special effect audio data corresponding to the target audio data, where the audio data adjustment performed includes adjusting the pitch; Among them, when the processing unit is used to screen out the beat interval values to be optimized from the beat interval sequence, it is used to count the number of times each value appears in the beat interval sequence, and determine the reference beat interval value of the target audio data as the value that appears the most times in the beat interval sequence; obtain a beat screening threshold; process the i-th beat interval value and the reference beat interval value to obtain an interval error value, where i is a positive integer and i is less than or equal to N - 1; among them, the interval error value of the i-th beat interval value includes: an interval error value determined according to the absolute value of the difference between the reference beat interval value and the i-th beat interval value, an interval error value determined according to the absolute value of the difference between twice the reference beat interval value and the i-th beat interval value, and an interval error value determined according to the absolute value of the difference between half of the reference beat interval value and the i-th beat interval value; if all the interval error values are greater than the screening threshold, the i-th beat interval value is determined as the beat interval value to be optimized.

5. An intelligent device, characterized in that, Including: a storage device and a processor; a computer program is stored in the storage device; The processor executes the computer program to implement the audio processing method according to any one of claims 1 - 3.

6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the audio processing method according to any one of claims 1 - 3 is implemented.

7. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the audio processing method according to any one of claims 1 - 3 is implemented.

Citation Information

Patent Citations

  • Karaoke sharing system based on set top box

    CN104079966A

  • Audio data processing method and device

    CN106970771A

  • Music classification method and beat point detection method, storage equipment and computer equipment

    CN108320730A

  • Detection method for main beat points in music, computer storage medium and terminal

    CN108335688A