An audio processing method, apparatus, electronic device and storage medium
By performing vocal detection and beat analysis on the audio, and automatically marking the starting time point of the chorus, the problem of high manpower and money costs in the existing technology is solved, and efficient and accurate marking of the time point of the chorus is achieved.
Patent Information
- Application Number
- CN202111571943.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-21
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2041-12-21
AI Technical Summary
In the prior art, marking the time point of the song chorus in the prior art requires a lot of manpower and money, and the accuracy is difficult to guarantee.
By performing vocal detection on the audio, extracting vocal clips and performing rhythm detection, sorting them according to timestamps, clustering them to determine the starting time point of the chorus.
It reduces manpower and money costs and improves the accuracy of marking the time points of the chorus.
Smart Images

Figure CN114512147B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of Internet technologies, and in particular, to an audio processing method, apparatus, electronic device, and storage medium. Background Art
[0002] With the development of multimedia technologies, music has become an essential part of people's lives, and many people obtain happiness and relieve stress by listening to music. The structure of music is an important part of a complete and excellent music work. Among them, the verse and the chorus are particularly important in the music structure. As the largest part of a song, the verse is usually the part before the chorus in a song, which is used to tell a story and advance the mood. And the chorus, as the most core part of the whole song, has a strong contrast with the verse in terms of melody and rhythm. In the chorus part, the emotion of the song is usually sublimated, and because of the uniqueness of its melody, it is often the memory point of the whole song. However, due to the diversity of song structures, it is very difficult to efficiently and accurately obtain the chorus time point of a song.
[0003] Currently, for the annotation of the chorus time point, it is usually to manually annotate the chorus time point of a song. Although the standard accuracy of this method can be guaranteed, it requires a high labor cost and financial cost. Summary of the Invention
[0004] The present disclosure provides an audio processing method, apparatus, electronic device, and storage medium. The technical solutions of the present disclosure are as follows:
[0005] According to a first aspect of an embodiment of the present disclosure, an audio processing method is provided, including:
[0006] Performing voice detection on the audio to obtain a voice segment;
[0007] Performing beat detection on the voice segment to obtain a plurality of bars corresponding to the voice segment; the plurality of bars are sorted according to timestamps;
[0008] Clustering the plurality of bars to divide the plurality of bars into a plurality of first clustering segments; each first clustering segment in the plurality of first clustering segments includes at least one bar;
[0009] Determining a first target clustering segment from the plurality of first clustering segments;
[0010] Determining the start time point of the first target clustering segment as the start time point of the chorus of the audio.
[0011] In some possible embodiments, performing beat detection on the voice segment to obtain a plurality of bars corresponding to the voice segment includes:
[0012] Perform beat detection on the vocal segment to obtain the timestamps corresponding to each measure in the vocal segment;
[0013] Segment the vocal segment according to the timestamps corresponding to each measure to obtain multiple measures corresponding to the vocal segment.
[0014] In some possible embodiments, performing beat detection on the vocal segment to obtain the timestamps corresponding to each measure in the vocal segment includes:
[0015] Extract the Mel-frequency cepstral coefficients of the audio as the feature information of the audio;
[0016] Perform beat detection on the vocal segment based on the feature information of the audio to obtain the timestamps corresponding to each measure in the vocal segment.
[0017] In some possible embodiments, the vocal segment includes at least one vocal sub-segment, and each vocal sub-segment carries a vocal start time point and a vocal end time point; the method further includes:
[0018] Determine the start time point and the end time point of each first clustering segment in the multiple first clustering segments;
[0019] Perform boundary adjustment on the multiple first clustering segments based on the start time point and the end time point of each first clustering segment and the vocal start time point and the vocal end time point carried by each vocal sub-segment to obtain the updated multiple first clustering segments.
[0020] In some possible embodiments, determining the first target clustering segment from the multiple first clustering segments includes:
[0021] Determine the first target clustering segment from the updated multiple first clustering segments.
[0022] In some possible embodiments, determining the first target clustering segment from the updated multiple first clustering segments includes:
[0023] Determine the short-time energy information of each first clustering segment in the updated multiple first clustering segments;
[0024] Determine the first target clustering segment from the updated multiple first clustering segments based on the short-time energy information of each first clustering segment.
[0025] In some possible embodiments, the method further includes:
[0026] When the duration of the vocal segment is less than the first preset duration, determine the audio to be processed based on the audio;
[0027] Perform beat detection on the audio to be processed to obtain multiple measures corresponding to the audio to be processed;
[0028] Determine similar segments from multiple segments corresponding to the audio to be processed, as well as the category information of each segment;
[0029] Cluster multiple segments corresponding to the audio to be processed, and divide the multiple segments into multiple second clustering segments corresponding to the audio to be processed;
[0030] Adjust multiple second clustering segments corresponding to the audio to be processed based on the similar segments and the category information of each segment, to obtain updated multiple second clustering segments;
[0031] Determine a second target clustering segment from multiple second clustering segments;
[0032] Determine the start time point of the second target clustering segment as the start time point of the chorus of the audio.
[0033] In some possible embodiments, performing beat detection on the audio to be processed, and obtaining multiple segments corresponding to the audio to be processed includes:
[0034] Performing beat detection on the audio to be processed, to obtain the time stamps corresponding to each segment in the audio to be processed;
[0035] Segment the audio to be processed according to the time stamps corresponding to each segment in the audio to be processed, to obtain multiple segments corresponding to the audio to be processed.
[0036] In some possible embodiments, performing beat detection on the audio to be processed, and obtaining the time stamps corresponding to each segment in the audio to be processed includes:
[0037] Extract the Mel-frequency cepstral coefficients of the audio to be processed as the feature information of the audio to be processed;
[0038] Perform beat detection on the audio to be processed based on the feature information of the audio to be processed, to obtain the time stamps corresponding to each segment in the audio to be processed.
[0039] In some possible embodiments, determining similar segments from multiple segments corresponding to the audio to be processed, as well as the category information of each segment includes:
[0040] Calculate the similarity information between any two segments among multiple segments corresponding to the audio to be processed;
[0041] Determine similar segments based on the similarity information;
[0042] Perform spectral clustering processing on the similarity information to determine the category information of each segment.
[0043] According to the second aspect of the embodiments of the present disclosure, there is provided an audio processing apparatus, including:
[0044] A voice detection module, configured to perform voice detection on the audio to obtain voice segments;
[0045] A beat detection module, configured to perform beat detection on the voice segments to obtain multiple bars corresponding to the voice segments; the multiple bars are sorted according to timestamps;
[0046] A clustering module, configured to perform clustering on the multiple bars to divide the multiple bars into multiple first clustering segments; each first clustering segment in the multiple first clustering segments includes at least one bar;
[0047] A segment determination module, configured to perform determining a first target clustering segment from the multiple first clustering segments;
[0048] A chorus start point determination module, configured to perform determining the start time point of the first target clustering segment as the chorus start time point of the audio.
[0049] In some possible embodiments, the beat detection module is configured to perform:
[0050] Perform beat detection on the voice segments to obtain the timestamps corresponding to each bar in the voice segments;
[0051] Segment the voice segments according to the timestamps corresponding to each bar to obtain multiple bars corresponding to the voice segments.
[0052] In some possible embodiments, the beat detection module is configured to perform:
[0053] Extract the Mel-frequency cepstral coefficients of the audio as the feature information of the audio;
[0054] Perform beat detection on the voice segments based on the feature information of the audio to obtain the timestamps corresponding to each bar in the voice segments.
[0055] In some possible embodiments, the voice segments include at least one voice sub-segment, and each voice sub-segment carries a voice start time point and a voice end time point; the apparatus further includes:
[0056] A time point determination module, configured to perform determining the start time point and the end time point of each first clustering segment in the multiple first clustering segments;
[0057] A segment update module, configured to perform boundary adjustment on the multiple first clustering segments based on the start time point and the end time point of each first clustering segment and the voice start time point and the voice end time point carried by each voice sub-segment to obtain the updated multiple first clustering segments.
[0058] In some possible embodiments, the segment determination module is configured to perform:
[0059] Determine a first target clustering segment from multiple updated first clustering segments.
[0060] In some possible embodiments, the segment determination module is configured to perform:
[0061] Determine the short-time energy information of each first clustering segment in the multiple updated first clustering segments;
[0062] Based on the short-time energy information of each first clustering segment, determine a first target clustering segment from the multiple updated first clustering segments.
[0063] In some possible embodiments, the apparatus further includes:
[0064] An audio to be processed determination module, configured to perform: when the duration of the human voice segment is less than a first preset duration, determine the audio to be processed based on the audio;
[0065] A beat detection module, configured to perform beat detection on the audio to be processed to obtain multiple bars corresponding to the audio to be processed;
[0066] A bar information determination module, configured to perform: determine similar bars and the category information of each bar from the multiple bars corresponding to the audio to be processed;
[0067] A clustering module, configured to perform clustering on the multiple bars corresponding to the audio to be processed, and divide the multiple bars into multiple second clustering segments corresponding to the audio to be processed;
[0068] A segment update module, configured to perform: adjust the multiple second clustering segments corresponding to the audio to be processed based on the similar bars and the category information of each bar to obtain multiple updated second clustering segments;
[0069] A segment determination module, configured to perform: determine a second target clustering segment from the multiple second clustering segments;
[0070] A chorus start point determination module, configured to perform: determine the start time point of the second target clustering segment as the chorus start time point of the audio.
[0071] In some possible embodiments, the beat detection module is configured to perform:
[0072] Perform beat detection on the audio to be processed to obtain the time stamp corresponding to each bar in the audio to be processed;
[0073] Segment the audio to be processed according to the time stamp corresponding to each bar in the audio to be processed to obtain multiple bars corresponding to the audio to be processed.
[0074] In some possible embodiments, a beat detection module is configured to perform:
[0075] Extract the Mel-frequency cepstral coefficients of the audio to be processed as the feature information of the audio to be processed;
[0076] Perform beat detection on the audio to be processed based on the feature information of the audio to be processed, and obtain the time stamps corresponding to each measure in the audio to be processed.
[0077] In some possible embodiments, a measure information determination module is configured to perform:
[0078] Calculate the similarity information between any two of the multiple measures corresponding to the audio to be processed;
[0079] Determine similar measures based on the similarity information;
[0080] Perform spectral clustering on the similarity information to determine the class information of each measure.
[0081] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to execute the instructions to implement the method according to any one of the first aspects as described above.
[0082] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device can execute the method according to any one of the first aspects of the embodiments of the present disclosure.
[0083] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product, the computer program product includes a computer program, the computer program is stored in a readable storage medium, and at least one processor of the computer device reads and executes the computer program from the readable storage medium, so that the computer device executes the method according to any one of the first aspects of the embodiments of the present disclosure.
[0084] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:
[0085] Perform voice detection on the audio to obtain voice segments, perform beat detection on the voice segments to obtain multiple measures corresponding to the voice segments, sort the multiple measures according to time stamps, perform clustering on the multiple measures, divide the multiple measures into multiple first clustering segments, each first clustering segment in the multiple first clustering segments includes at least one measure, determine a first target clustering segment from the multiple first clustering segments, and determine the start time point of the first target clustering segment as the start time point of the chorus of the audio. In this way, the start time point of the chorus of the audio can be determined by the device, reducing the labor cost and financial cost.
[0086] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0087] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an undue limitation to the present disclosure.
[0088] Figure 1 is a schematic diagram of an application environment shown according to an exemplary embodiment;
[0089] Figure 2 is a flowchart of an audio processing method shown according to an exemplary embodiment;
[0090] Figure 3 is a flowchart of a measure processing method shown according to an exemplary embodiment;
[0091] Figure 4 is a process of a clustering segment boundary adjustment method shown according to an exemplary embodiment;
[0092] Figure 5 is a flowchart of an audio processing method shown according to an exemplary embodiment;
[0093] Figure 6 is a flowchart of a measure processing method shown according to an exemplary embodiment;
[0094] Figure 7 is a block diagram of an audio processing device shown according to an exemplary embodiment;
[0095] Figure 8 is a block diagram of an electronic device for audio processing shown according to an exemplary embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0096] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0097] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present disclosure are used to distinguish similar first objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described here can be implemented in an order other than those illustrated or described here. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0098] All data related to users in this application are data after user authorization.
[0099] Please refer to Figure 1 , Figure 1 is a schematic diagram of the application environment of an audio processing method shown according to an exemplary embodiment. As Figure 1 shown, the application environment may include a server 01 and a client 02.
[0100] In the embodiment of the present application, the server 01 may obtain the audio sent by the client 02. Subsequently, the server 01 may perform voice detection on the audio to obtain a voice segment, the duration of the voice segment may be greater than or equal to a first preset duration, perform beat detection on the voice segment to obtain a plurality of bars corresponding to the voice segment, the plurality of bars are sorted according to timestamps, cluster the plurality of bars, divide the plurality of bars into a plurality of first cluster segments, each first cluster segment in the plurality of first cluster segments includes at least one bar, determine a first target cluster segment from the plurality of first cluster segments, and determine the start time point of the first target cluster segment as the start time point of the chorus of the audio.
[0101] The server 01 may include an independent physical server, or may be a server cluster or distributed system composed of multiple physical servers, or may also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The operating systems running on the server may include, but are not limited to, Android system, IOS system, linux, windows, Unix, etc.
[0102] In some possible embodiments, the above-mentioned client 02 sends audio to the server 01. The client 02 may include, but is not limited to, clients of types such as smart phones, desktop computers, tablet computers, laptop computers, smart speakers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, smart wearable devices, etc. It may also be software running on the above-mentioned clients, such as applications, applets, etc. Optionally, the operating systems running on the clients may include, but are not limited to, Android system, IOS system, Linux, Windows, Unix, etc.
[0103] Figure 2 is a flowchart of an audio processing method shown according to an exemplary embodiment. As Figure 2 shown, the audio processing method can be applied to the server or other node devices, including the following steps:
[0104] In step S201, perform voice detection on the audio to obtain a voice segment.
[0105] In the embodiments of the present application, the server can obtain an original audio, which can be a complete piece of music, an incomplete piece of music (such as a segment of a certain piece of music), a music containing voices, or a pure music mode music without voices.
[0106] In an optional embodiment, when the server obtains the original audio, it can perform voice detection on the audio to obtain a first segment and a second segment. Among them, the first segment can be the segment containing voices in the audio. The second segment can be the audio without voices.
[0107] Since in some possible embodiments, the processing of the audio may be adjusted based on the proportion of the first segment, therefore, the present application can determine the duration of the first segment. When the duration of the first segment meets the first preset duration, the server can define the first segment as the voice segment.
[0108] Optionally, the server can obtain the duration of the audio (such as 210 seconds), then the server can obtain a preset proportion value (such as 50%), and determine the first preset duration based on the duration of the audio and the preset proportion data. For example, the server can determine the first preset duration as 105 seconds based on the product of the duration of the audio and the preset proportion data. When the duration of the first segment is greater than or equal to 105 seconds, the first segment can be defined as the voice segment for subsequent processing.
[0109] Optionally, the server can directly obtain the first preset duration, which can be predefined. When the duration of the first segment is greater than or equal to the first preset duration, the first segment can be defined as a vocal segment for subsequent processing. Optionally, the first preset duration can be predefined according to statistical values or empirical values.
[0110] In the embodiments of the present application, a vocal segment can include multiple vocal sub-segments.
[0111] In some possible embodiments, each vocal sub-segment can correspond to a line of lyrics.
[0112] In other possible embodiments, for the convenience of subsequent analysis and processing, when the segment corresponding to a line of lyrics is greater than or equal to the cutting duration, the line of lyrics can be divided into two vocal sub-segments. Optionally, the cutting duration can be 10 seconds, or it can be adaptively adjusted according to the actual situation of the processing.
[0113] In the embodiments of the present application, the server can perform vocal detection on the audio using Voice Activity Detection (VAD) to obtain vocal sub-segments.
[0114] In step S203, beat detection is performed on the vocal segment to obtain multiple bars corresponding to the vocal segment; the multiple bars are sorted according to the timestamps.
[0115] In an alternative embodiment, the server can perform beat detection on the vocal segment to obtain multiple bars corresponding to the vocal segment. Specifically, the server can perform beat detection on each vocal sub-segment to obtain multiple bars corresponding to each vocal sub-segment, and then the multiple bars corresponding to each vocal sub-segment can be sorted according to the timestamps carried by each bar to obtain the bar sequence corresponding to the vocal segment. The bar sequence contains multiple bars corresponding to each vocal sub-segment.
[0116] Figure 3 is a flowchart of a bar processing method shown according to an exemplary embodiment, as Figure 3 shown, including the following steps:
[0117] In step S2031, beat detection is performed on the vocal segment to obtain the timestamps corresponding to each bar in the vocal segment.
[0118] Optionally, the server can use a beat detection algorithm to perform beat detection on each vocal sub-segment in the vocal segment to obtain the timestamps corresponding to each bar in each vocal sub-segment. The timestamp can be the time point of the start of each bar in the audio, or it can be the time point of the end of each bar in the audio.
[0119] In the embodiments of the present application, in order to make the segmentation of subsections more accurate, before detecting the beats of the vocal sub-fragments, the server can extract the Mel Frequency Cepstral Coefficients (MFCCs) of the audio to obtain the feature information of the audio, and then detect the beats of the vocal fragment based on the feature information of the audio to obtain the timestamps corresponding to each subsection in the vocal fragment.
[0120] Optionally, the server can also extract the Mel Frequency Cepstral Coefficients of each vocal sub-fragment to obtain the feature information of each vocal sub-fragment, and then detect the beats of the vocal fragment based on the feature information of each vocal sub-fragment to obtain the timestamps corresponding to each subsection in the vocal fragment.
[0121] In step S2033, the vocal fragment is segmented according to the timestamps corresponding to each subsection to obtain multiple subsections corresponding to the vocal fragment.
[0122] After determining the timestamps corresponding to each subsection in the vocal fragment, the server can segment each vocal sub-fragment in the vocal fragment according to the timestamps corresponding to each subsection to obtain multiple subsections corresponding to the vocal fragment. These multiple subsections are sorted according to the timestamps carried by each subsection.
[0123] In another alternative embodiment, while detecting the vocals in the audio, the beats of the audio can be detected to obtain multiple subsections corresponding to the audio, and each subsection can carry a timestamp. When the vocal fragment is determined from the audio, the start time point and end time point of the vocals of each vocal sub-fragment in the vocal fragment can be determined, and based on the start time point and end time point of the vocals, and the timestamps corresponding to each subsection, multiple subsections corresponding to the vocal fragment are determined from the multiple subsections corresponding to the audio, and all the subsections are sorted in the order of the timestamps.
[0124] In step S205, clustering is performed on the multiple subsections to divide the multiple subsections into multiple first clustering segments; each first clustering segment in the multiple first clustering segments includes at least one subsection.
[0125] Assume that the above-mentioned subsection sequence includes 100 subsections, and the timestamps of the subsections in the front are earlier than those of the subsections in the back. These 100 subsections can be located as subsection 1, subsection 2, subsection 3... subsection 100.
[0126] In the embodiments of the present application, the server can use constrained agglomerative hierarchical clustering to cluster the multiple subsections to divide the multiple subsections into multiple first clustering segments. Among them, the minimum number of subsections included in each first clustering segment is one. When there are multiple subsections included in the first clustering segment, the multiple subsections are sorted according to the timestamps, and the multiple subsections are all adjacent subsections.
[0127] For example, assuming it is divided into 5 first clustering segments, the first first clustering segment may include subsection 1, subsection 2, subsection 3... subsection 20; the second first clustering segment may include subsection 21, subsection 22, subsection 23... subsection 40; the third first clustering segment may include subsection 41, subsection 42, subsection 43... subsection 60; the fourth first clustering segment may include subsection 61, subsection 62, subsection 63... subsection 80; the fifth first clustering segment may include subsection 81, subsection 82, subsection 83... subsection 100. And there will be no situation where the first first clustering segment includes subsection 1, subsection 2... subsection 18, subsection 20, subsection 22, while the second first clustering segment includes subsection 19, subsection 21, subsection 23... subsection 40.
[0128] However, for the multiple first clustering segments obtained by the server, it is possible that the start time point of a certain first clustering segment is exactly in the middle of a certain lyric sentence, or the end time point is exactly in the middle of a certain lyric sentence. To solve the above problems, the server can adjust the boundaries of the first clustering segments based on each sub-segment in the vocal segment to obtain multiple updated first clustering segments.
[0129] Figure 4 It is a flowchart of a method for adjusting the boundaries of clustering segments shown according to an exemplary embodiment, as Figure 4 shown, and includes the following steps:
[0130] In step S2061, determine the start time point and end time point of each first clustering segment among the multiple first clustering segments.
[0131] In the embodiments of the present application, the server determines the start time point and end time point of each first clustering segment among the multiple first clustering segments. In addition, the server can also determine the start time point and end time point of each vocal sub-segment. To make a distinction from the start time point and end time point of the first clustering segments, the server can call the start time point of each vocal sub-segment the vocal start time point and the end time point the vocal end time point.
[0132] In step S2063, based on the start time point and end time point of each first clustering segment, and the vocal start time point and vocal end time point carried by each vocal sub-segment, adjust the boundaries of the multiple first clustering segments to obtain multiple updated first clustering segments.
[0133] The server can adjust the boundaries of the multiple first clustering segments based on the start time point and end time point of each first clustering segment and the vocal start time point and vocal end time point carried by each vocal sub-segment to obtain multiple updated first clustering segments.
[0134] For example, when the end time point of the first clustering segment is the 20th second and it contains 10 bars, with each bar of the audio lasting 2 seconds, there is a vocal sub-segment with a vocal start time point of the 18th second and a vocal end time point of the 22nd second. Then the server can adjust the end time point of the first clustering segment to the 18th second, which contains 9 bars, or adjust the end time point of the first clustering segment to the 22nd second, which contains 11 bars. It can be seen that the service can adjust the first clustering segment forward or backward.
[0135] In this way, the updated first first clustering segment can contain bar 1, bar 2, bar 3... bar 19; the second first clustering segment can contain bar 20, bar 22, bar 23... bar 42; the third first clustering segment can contain bar 43... bar 60; the fourth first clustering segment can contain bar 61, bar 62, bar 63... bar 78; the fifth first clustering segment can contain bar 79, bar 80, bar 81, bar 82, bar 83... bar 100.
[0136] In step S207, determine the first target clustering segment from multiple first clustering segments.
[0137] When the first clustering segment is not updated, optionally, the server can determine the short-time energy information of each first clustering segment in the updated multiple first clustering segments. And determine the first target clustering segment from the updated multiple first clustering segments based on the short-time energy information of each first clustering segment.
[0138] When the first clustering segment is updated, optionally, the server can determine the first target clustering segment from the updated multiple first clustering segments. Specifically, the server can determine the short-time energy information of each first clustering segment in the updated multiple first clustering segments, and determine the first target clustering segment from the updated multiple first clustering segments based on the short-time energy information of each first clustering segment.
[0139] Optionally, the server can determine the short-time energy information of each frame in each first clustering segment, and determine the short-time energy information of the first clustering segment based on the average value of the short-time energy information of each frame in the same first clustering segment. Subsequently, the server can determine the first clustering segment with the highest short-time energy information as the first target clustering segment. Optionally, the server can determine the first clustering segment with the second highest short-time energy information as the first target clustering segment.
[0140] In the embodiments of the present application, since the energy of the speech signal changes over time and the energy difference between voiceless and voiced sounds is quite significant. Therefore, analyzing the short-time energy can describe this characteristic change of the speech.
[0141] In step S209, determine the start time point of the first target clustering segment as the start time point of the chorus of the audio.
[0142] The server may determine the start time point of the first target clustering segment as the start time point of the chorus of the audio.
[0143] In the embodiments of the present application, Figure 5 is a flowchart of an audio processing method shown according to an exemplary embodiment, applied to a server, and includes the following steps:
[0144] In step S501, when the duration of the vocal segment is less than the first preset duration, determine the audio to be processed based on the audio.
[0145] How to determine the first preset duration has been mentioned above. When the duration of the first segment, that is, the vocal segment, is less than the first preset duration, the audio to be processed can be determined based on the audio.
[0146] In an alternative embodiment, the server may determine the duration of the audio and determine how to determine the audio to be processed based on the duration of the audio and the second preset duration.
[0147] Optionally, when the duration of the audio is greater than or equal to the second preset duration, a certain segment in the audio can be selected as the audio to be processed. The certain segment can be determined based on the empirical value of the possible appearance area of the chorus. For example, the 10%-50% area of the audio can be determined as the audio to be processed.
[0148] Optionally, when the duration of the audio is less than the second preset duration, the audio can be determined as the audio to be processed.
[0149] In step S503, perform beat detection on the audio to be processed to obtain multiple bars corresponding to the audio to be processed.
[0150] In an alternative embodiment, the server may perform beat detection on the audio to be processed to obtain multiple bars corresponding to the audio to be processed.
[0151] Figure 6 is a flowchart of a bar processing method shown according to an exemplary embodiment, as Figure 6 shown, and includes the following steps:
[0152] In step S5031, perform beat detection on the audio to be processed to obtain the time stamps corresponding to each bar in the audio to be processed.
[0153] Optionally, the server may use a beat detection algorithm to perform beat detection on the audio to be processed, and obtain the timestamps corresponding to the audio to be processed. The timestamp may be the time point of the beginning of each measure in the audio, or may be the time point of the end of each measure in the audio.
[0154] In the embodiments of the present application, in order to make the segmentation of measures more accurate, before performing beat detection on the audio to be processed, the server may extract the Mel-frequency cepstral coefficients of the audio to be processed to obtain the feature information of the audio to be processed, and then perform beat detection on the audio to be processed based on the feature information of the audio to be processed to obtain the timestamps corresponding to each measure in the audio to be processed.
[0155] In step S5033, the audio to be processed is segmented according to the timestamps corresponding to each measure in the audio to be processed, and multiple measures corresponding to the audio to be processed are obtained.
[0156] After determining the timestamps corresponding to each measure in the multiple measures corresponding to the audio to be processed, the server may segment the audio to be processed according to the timestamps corresponding to each measure, and obtain multiple measures corresponding to the audio to be processed.
[0157] In another alternative embodiment, while performing voice detection on the audio, beat detection may be performed on the audio to obtain multiple measures corresponding to the audio, and each measure may carry a timestamp. When the audio to be processed is determined from the audio, the start time point and end time point of the audio to be processed may be determined, and based on the start time point and end time point of the audio to be processed, and the timestamps corresponding to each measure, multiple measures corresponding to the audio to be processed are determined from the multiple measures corresponding to the audio.
[0158] In step S505, similar measures are determined from the multiple measures corresponding to the audio to be processed, and the category information of each measure is obtained.
[0159] Optionally, the server may calculate the similarity information between any two measures in the multiple measures corresponding to the audio to be processed, and determine the similar measures based on the similarity information. Spectral clustering processing is also performed on the similarity information to determine the category information of each measure.
[0160] Specifically, the server may calculate the cosine similarity between each measure to form a self-similarity matrix. For example, assume there are 100 measures. The server needs to calculate the cosine similarity between each measure and the remaining 99 measures to obtain the similarity information between each measure and other measures.
[0161] Subsequently, the server can perform Gaussian smoothing on the self-similarity matrix and then set a threshold to convert it into a connectivity matrix (0 / 1), thereby obtaining existing similar segments. Assume there is a self-similarity matrix, and it is converted into a connectivity matrix (0 / 1). Each number in the matrix is either 1 or 0. When the value in the first row and third column is 1, it indicates that the 1st segment and the 3rd segment are similar segments. When the value in the first row and second column is 0, it indicates that the 1st segment and the 2nd segment are not similar segments.
[0162] Optionally, the server calculates the Laplacian matrix from the self-similarity matrix and performs spectral clustering on the Laplacian matrix to obtain the categories of each segment. Thus, the segments corresponding to each category can be obtained. For example, the segments corresponding to the first category include segment 1, segment 3, segment 8, segment 11, and segment 22.
[0163] Optionally, spectral clustering is an algorithm evolved from graph theory and has been widely used in clustering later. Its main idea is to regard all data as points in space, and these points can be connected by edges. The edge weight value between two points with a large distance is low, while the edge weight value between two points with a small distance is high. By cutting the graph composed of all data points, the sum of the edge weights between different subgraphs after cutting is made as low as possible, and the sum of the edge weights within the subgraph is made as high as possible, thereby achieving the purpose of clustering.
[0164] In step S507, clustering is performed on multiple segments corresponding to the audio to be processed, and the multiple segments are divided into multiple second clustering segments corresponding to the audio to be processed.
[0165] Assume that the above segment sequence includes 100 segments, and these 100 segments can be numbered as segment 1, segment 2, segment 3... segment 100.
[0166] The server can cluster multiple segments using constrained agglomerative hierarchical clustering and divide the multiple segments into multiple second clustering segments. Among them, each second clustering segment contains at least one segment. When multiple segments are included in a second clustering segment, the multiple segments are adjacent segments according to the sorting of timestamps.
[0167] For example, assuming that it is divided into 5 second clustering segments, the first second clustering segment may include section 1, section 2, section 3... section 20; the second second clustering segment may include section 21, section 22, section 23... section 40; the third second clustering segment may include section 41, section 42, section 43... section 60; the fourth second clustering segment may include section 61, section 62, section 63... section 80; the fifth second clustering segment may include section 81, section 82, section 83... section 100. And it will not occur that the first second clustering segment includes section 1, section 2... section 18, section 20, section 22, while the second second clustering segment includes section 19, section 21, section 23... section 40.
[0168] In step S509, based on the similar sections and the category information of each section, the multiple second clustering segments corresponding to the audio to be processed are adjusted to obtain the updated multiple second clustering segments.
[0169] Optionally, on the premise that the server does not accurately split adjacent similar sections, based on the agglomerative hierarchical clustering result, the spectral clustering result is used to adjust the overall boundary. For example, when section 1, section 3, section 21, and section 23 are similar sections, the sections corresponding to the first category include section 1, section 3, section 8, section 11, and section 22, and a second clustering segment may include section 1, section 2, section 3... section 20. The server can re-determine the boundary of the second clustering segment based on the similar sections and the category information of each section, and obtain the included section 1, section 2, section 3... section 20, section 21, section 22, section 23. In this way, the similar sections and the sections of the same category can be placed in one clustering segment.
[0170] In step S511, the second target clustering segment is determined from the multiple second clustering segments.
[0171] In the embodiment of the present application, the server can determine the second target clustering segment from the updated multiple second clustering segments. Optionally, the server can determine the second target clustering segment from the updated multiple second clustering segments. Specifically, the server can determine the short-time energy information of each second clustering segment in the updated multiple second clustering segments, and determine the second target clustering segment from the updated multiple second clustering segments based on the short-time energy information of each second clustering segment.
[0172] Optionally, the server may determine the short-time energy information of each frame in each second clustering segment, and determine the short-time energy information of the second clustering segment based on the average value of the short-time energy information of each frame in the same second clustering segment. Subsequently, the server may determine the second clustering segment with the highest short-time energy information as the second target clustering segment.
[0173] In step S513, the starting time point of the second target clustering segment is determined as the starting time point of the chorus of the audio.
[0174] The server may determine the starting time point of the second target clustering segment as the starting time point of the chorus of the audio.
[0175] Figure 7 is a block diagram of an audio processing apparatus shown according to an exemplary embodiment. Referring to Figure 7 , the apparatus includes: a voice detection module 701, a beat detection module 702, a clustering module 703, a segment determination module 704, and a chorus start point determination module 705.
[0176] The voice detection module 701 is configured to perform voice detection on the audio to obtain a voice segment;
[0177] The beat detection module 702 is configured to perform beat detection on the voice segment to obtain a plurality of bars corresponding to the voice segment; the plurality of bars are sorted according to timestamps;
[0178] The clustering module 703 is configured to perform clustering on the plurality of bars, and divide the plurality of bars into a plurality of first clustering segments; each first clustering segment in the plurality of first clustering segments includes at least one bar;
[0179] The segment determination module 704 is configured to perform determining a first target clustering segment from the plurality of first clustering segments;
[0180] The chorus start point determination module 705 is configured to perform determining the starting time point of the first target clustering segment as the starting time point of the chorus of the audio.
[0181] In some possible embodiments, the beat detection module is configured to perform:
[0182] Perform beat detection on the voice segment to obtain the timestamp corresponding to each bar in the voice segment;
[0183] Segment the voice segment according to the timestamp corresponding to each bar to obtain a plurality of bars corresponding to the voice segment.
[0184] In some possible embodiments, the beat detection module is configured to perform:
[0185] Extract the Mel-frequency cepstral coefficients of the audio as the feature information of the audio;
[0186] Perform beat detection on the vocal segment based on the feature information of the audio to obtain the time stamps corresponding to each measure in the vocal segment.
[0187] In some possible embodiments, the vocal segment includes at least one vocal sub-segment, and each vocal sub-segment carries a vocal start time point and a vocal end time point; the apparatus further includes:
[0188] A time point determination module configured to determine the start time point and the end time point of each first clustering segment among a plurality of first clustering segments;
[0189] A segment update module configured to perform boundary adjustment on the plurality of first clustering segments based on the start time point and the end time point of each first clustering segment and the vocal start time point and the vocal end time point carried by each vocal sub-segment to obtain an updated plurality of first clustering segments.
[0190] In some possible embodiments, the segment determination module is configured to perform:
[0191] Determine a first target clustering segment from the updated plurality of first clustering segments.
[0192] In some possible embodiments, the segment determination module is configured to perform:
[0193] Determine the short-time energy information of each first clustering segment in the updated plurality of first clustering segments;
[0194] Determine a first target clustering segment from the updated plurality of first clustering segments based on the short-time energy information of each first clustering segment.
[0195] In some possible embodiments, the apparatus further includes:
[0196] A to-be-processed audio determination module configured to perform, when the duration of the vocal segment is less than a first preset duration, determine the to-be-processed audio based on the audio;
[0197] A beat detection module configured to perform beat detection on the to-be-processed audio to obtain a plurality of measures corresponding to the to-be-processed audio;
[0198] A measure information determination module configured to perform determining similar measures and the category information of each measure from the plurality of measures corresponding to the to-be-processed audio;
[0199] A clustering module configured to perform clustering on the plurality of measures corresponding to the to-be-processed audio and divide the plurality of measures into a plurality of second clustering segments corresponding to the to-be-processed audio;
[0200] The segment update module is configured to perform adjustments on multiple second clustering segments corresponding to the audio to be processed based on similar segments and the category information of each segment, so as to obtain multiple updated second clustering segments;
[0201] The segment determination module is configured to perform determining a second target clustering segment from multiple second clustering segments;
[0202] The chorus start point determination module is configured to perform determining the start time point of the second target clustering segment as the chorus start time point of the audio.
[0203] In some possible embodiments, the beat detection module is configured to perform:
[0204] Perform beat detection on the audio to be processed to obtain the time stamps corresponding to each segment in the audio to be processed;
[0205] Segment the audio to be processed according to the time stamps corresponding to each segment in the audio to be processed, so as to obtain multiple segments corresponding to the audio to be processed.
[0206] In some possible embodiments, the beat detection module is configured to perform:
[0207] Extract the Mel-frequency cepstral coefficients of the audio to be processed as the feature information of the audio to be processed;
[0208] Perform beat detection on the audio to be processed based on the feature information of the audio to be processed to obtain the time stamps corresponding to each segment in the audio to be processed.
[0209] In some possible embodiments, the segment information determination module is configured to perform:
[0210] Calculate the similarity information between any two segments among the multiple segments corresponding to the audio to be processed;
[0211] Determine similar segments based on the similarity information;
[0212] Perform spectral clustering processing on the similarity information to determine the category information of each segment.
[0213] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0214] Figure 8 It is a block diagram of a device 800 for audio processing shown according to an exemplary embodiment. For example, the device 800 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0215] Reference Figure 8 , the apparatus 800 may include one or more of the following components: a processing component 802, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 818.
[0216] The processing component 802 generally controls the overall operation of the apparatus 800, such as operations associated with display, telephone calls, data communications, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above-described methods. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0217] The memory 804 is configured to store various types of data to support the operation of the device 800. Examples of such data include instructions for any application or method operating on the apparatus 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0218] The power component 806 provides power to the various components of the apparatus 800. The power component 806 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the apparatus 800.
[0219] The multimedia component 808 includes a screen that provides an output interface between the device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of the touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operation mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0220] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive external audio signals when the device 800 is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 804 or sent via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.
[0221] The I / O interface 812 provides an interface between the processing component 802 and a peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include but are not limited to: a home button, a volume button, a power button, and a lock button.
[0222] The sensor component 814 includes one or more sensors for providing a status assessment of various aspects of the device 800. For example, the sensor component 814 can detect the on / off state of the device 800, the relative positioning of components, such as the display and the keypad of the device 800. The sensor component 814 can also detect a change in the position of the device 800 or a component of the device 800, the presence or absence of user contact with the device 800, the orientation or acceleration / deceleration of the device 800, and the temperature change of the device 800. The sensor component 814 can include a proximity sensor that is configured to detect the presence of nearby objects without any physical contact. The sensor component 814 can also include a light sensor, such as a CMOS or a CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 814 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0223] The communication component 816 is configured to facilitate communication, either wired or wirelessly, between the device 800 and other devices. The device 800 may access a wireless network based on a communication standard, such as WiFi, a carrier network (such as 2G, 3G, 4G, or 5G), or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0224] In an exemplary embodiment, the device 800 may be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above-described method.
[0225] In an exemplary embodiment, a storage medium including instructions, such as the memory 804 including instructions, is also provided, and the above instructions can be executed by the processor 820 of the device 800 to complete the above-described method. Optionally, the storage medium may be a non-transitory computer-readable storage medium. For example, the non-transitory computer-readable storage medium may be a ROM, a Random Access Memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
Claims
1. An audio processing method, characterized in that, Including: Performing voice detection on the audio to obtain voice segments; Performing beat detection on the voice segments to obtain multiple bars corresponding to the voice segments; Sorting the multiple bars according to timestamps; Clustering the multiple bars to divide the multiple bars into multiple first clustering segments; each first clustering segment in the multiple first clustering segments includes at least one bar; Determining a first target clustering segment from the multiple first clustering segments; Determining the start time point of the first target clustering segment as the start time point of the chorus of the audio; The voice segments include at least one voice sub-segment, and each voice sub-segment carries a voice start time point and a voice end time point; the method further includes: Determining the start time point and the end time point of each first clustering segment in the multiple first clustering segments; Performing boundary adjustment on the multiple first clustering segments based on the start time point and the end time point of each first clustering segment and the voice start time point and the voice end time point carried by each voice sub-segment to obtain updated multiple first clustering segments.
2. The audio processing method according to claim 1, wherein The performing beat detection on the voice segments to obtain multiple bars corresponding to the voice segments includes: Performing beat detection on the voice segments to obtain the timestamps corresponding to each bar in the voice segments; Segmenting the voice segments according to the timestamps corresponding to each bar to obtain multiple bars corresponding to the voice segments.
3. The audio processing method according to claim 2, wherein, The performing beat detection on the voice segments to obtain the timestamps corresponding to each bar in the voice segments includes: Extracting the Mel-frequency cepstral coefficients of the audio as the feature information of the audio; Performing beat detection on the voice segments based on the feature information of the audio to obtain the timestamps corresponding to each bar in the voice segments.
4. The audio processing method according to claim 1, characterized in that The determining a first target clustering segment from the multiple first clustering segments includes: Determining a first target clustering segment from the updated multiple first clustering segments.
5. The audio processing method according to claim 4, wherein The determining a first target clustering segment from the updated multiple first clustering segments includes: Determining the short-time energy information of each first clustering segment in the updated multiple first clustering segments; Determining a first target clustering segment from the updated multiple first clustering segments based on the short-time energy information of each first clustering segment.
6. The audio processing method according to claim 1, characterized in that, The method further includes: When the duration of the voice segments is less than a first preset duration, determining the audio to be processed based on the audio; Performing beat detection on the audio to be processed to obtain multiple bars corresponding to the audio to be processed; Determining similar bars and the category information of each bar from the multiple bars corresponding to the audio to be processed; Clustering the multiple bars corresponding to the audio to be processed to divide the multiple bars into multiple second clustering segments corresponding to the audio to be processed; Adjusting the multiple second clustering segments corresponding to the audio to be processed based on the similar bars and the category information of each bar to obtain updated multiple second clustering segments; Determining a second target clustering segment from the multiple second clustering segments; Determine the start time point of the second target clustering segment as the start time point of the chorus of the audio.
7. The audio processing method according to claim 6, wherein The beat detection of the audio to be processed to obtain multiple bars corresponding to the audio to be processed includes: Perform beat detection on the audio to be processed to obtain the time stamps corresponding to each bar in the audio to be processed; Segment the audio to be processed according to the time stamps corresponding to each bar in the audio to be processed to obtain multiple bars corresponding to the audio to be processed.
8. The audio processing method according to claim 7, wherein The beat detection of the audio to be processed to obtain the time stamps corresponding to each bar in the audio to be processed includes: Extract the Mel Frequency Cepstral Coefficients of the audio to be processed as the feature information of the audio to be processed; Perform beat detection on the audio to be processed based on the feature information of the audio to be processed to obtain the time stamps corresponding to each bar in the audio to be processed.
9. The audio processing method according to claim 6, wherein The determination of similar bars and the class information of each bar from the multiple bars corresponding to the audio to be processed includes: Calculate the similarity information between any two bars among the multiple bars corresponding to the audio to be processed; Determine the similar bars based on the similarity information; Perform spectral clustering on the similarity information to determine the class information of each bar.
10. An audio processing device, characterized in that, Including: A voice detection module configured to perform voice detection on the audio to obtain voice segments; A beat detection module configured to perform beat detection on the voice segments to obtain multiple bars corresponding to the voice segments; the multiple bars are sorted according to time stamps; A clustering module configured to perform clustering on the multiple bars to divide the multiple bars into multiple first clustering segments; each first clustering segment in the multiple first clustering segments includes at least one bar; A segment determination module configured to perform determining a first target clustering segment from the multiple first clustering segments; A chorus start point determination module configured to perform determining the start time point of the first target clustering segment as the start time point of the chorus of the audio; The voice segments include at least one voice sub-segment, and each voice sub-segment carries a voice start time point and a voice end time point; the apparatus further includes: A time point determination module configured to perform determining the start time point and the end time point of each first clustering segment in the multiple first clustering segments; A segment update module configured to perform boundary adjustment on the multiple first clustering segments based on the start time point and the end time point of each first clustering segment and the voice start time point and the voice end time point carried by each voice sub-segment to obtain the updated multiple first clustering segments.
11. The audio processing device according to claim 10, characterized in that The beat detection module is configured to perform: Perform beat detection on the voice segments to obtain the time stamps corresponding to each bar in the voice segments; Segment the voice segments according to the time stamps corresponding to each bar to obtain multiple bars corresponding to the voice segments.
12. The audio processing device according to claim 11, wherein The beat detection module is configured to perform: Extract the Mel Frequency Cepstral Coefficients of the audio as the feature information of the audio; Perform beat detection on the vocal segment based on the feature information of the audio to obtain the time stamps corresponding to each measure in the vocal segment.
13. The audio processing device according to claim 10, characterized in that, The segment determination module is configured to perform: Determine a first target clustering segment from the updated multiple first clustering segments.
14. The audio processing device according to claim 13, wherein The segment determination module is configured to perform: Determine the short-time energy information of each first clustering segment in the updated multiple first clustering segments; Based on the short-time energy information of each first clustering segment, determine a first target clustering segment from the updated multiple first clustering segments.
15. The audio processing device according to claim 10, characterized in that, The apparatus further includes: A to-be-processed audio determination module, configured to perform, when the duration of the vocal segment is less than a first preset duration, determine the to-be-processed audio based on the audio; The beat detection module is configured to perform beat detection on the to-be-processed audio to obtain multiple measures corresponding to the to-be-processed audio; A measure information determination module, configured to perform determining similar measures and the category information of each measure from the multiple measures corresponding to the to-be-processed audio; The clustering module is configured to perform clustering on the multiple measures corresponding to the to-be-processed audio, and divide the multiple measures into multiple second clustering segments corresponding to the to-be-processed audio; A segment update module, configured to perform adjusting the multiple second clustering segments corresponding to the to-be-processed audio based on the similar measures and the category information of each measure to obtain the updated multiple second clustering segments; The segment determination module is configured to perform determining a second target clustering segment from the multiple second clustering segments; The chorus start point determination module is configured to perform determining the start time point of the second target clustering segment as the chorus start time point of the audio.
16. The audio processing device according to claim 15, characterized in that, The beat detection module is configured to perform: Perform beat detection on the to-be-processed audio to obtain the time stamps corresponding to each measure in the to-be-processed audio; Segment the to-be-processed audio according to the time stamps corresponding to each measure in the to-be-processed audio to obtain multiple measures corresponding to the to-be-processed audio.
17. The audio processing device according to claim 16, wherein The beat detection module is configured to perform: Extract the Mel-frequency cepstral coefficients of the to-be-processed audio as the feature information of the to-be-processed audio; Perform beat detection on the to-be-processed audio based on the feature information of the to-be-processed audio to obtain the time stamps corresponding to each measure in the to-be-processed audio.
18. The audio processing device according to claim 15, characterized in that, The measure information determination module is configured to perform: Calculate the similarity information between any two measures in the multiple measures corresponding to the to-be-processed audio; Determine the similar measures based on the similarity information; Perform spectral clustering processing on the similarity information to determine the category information of each measure.
19. An electronic device, characterized in that, Includes: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the instructions to implement the audio processing method according to any one of claims 1 to 9.
20. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is enabled to execute the audio processing method according to any one of claims 1 to 9.
21. A computer program product, characterized in that, The computer program product includes a computer program which is stored in a readable storage medium. At least one processor of a computer device reads and executes the computer program, so that the computer device executes the audio processing method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Music section information determination method and device, storage medium and equipment
CN111128232A
Method and device for determining specific human voice segment in audio and electronic equipment
CN111243618A
Beat detection model training method, beat detection method and beat detection device
CN113223485A