Method, apparatus, device and storage medium for determining transition points
Through audio segmentation and importance parameter matching, different types of candidate points are selected as transition points, which solves the problem of singleness of video transition points at the pause point, realizes the diversity and richness of transition effects, and enhances the artistic expression of video production.
Patent Information
- Application Number
- CN202310185019.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-20
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2043-02-20
AI Technical Summary
In the existing video production technology of card point videos, the singularity of transition points leads to insufficient video transition effects and lack of diversity and richness.
By segmenting based on the chromaticity and rhythm characteristics of the audio frame, the importance parameters of the audio clip are determined, and the matching candidate points are selected from the beat points, remake points and mutation points as transition points, improving the diversity and richness of transition points.
It improves the transition effect of the stop video, enhances the diversity and richness of the video transition, and enhances the artistic expression of video production.
Smart Images

Figure CN116189708B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of video technology, and in particular to a method, apparatus, device, and storage medium for determining transition points. Background Art
[0002] The technology of making beat videos refers to a video technology that generates a video in which the picture matches the rhythm of the audio, so that the picture smoothly transitions at the rhythm points of the audio. In related technologies, when making a beat video, generally the beat points and accent points in the audio are used as transition points to make a beat video. The beat points and accent points in the audio are generally evenly distributed, and the transition points determined in this way make the transition effect of the video relatively single, thus reducing the transition effect of the beat video. Summary of the Invention
[0003] The present disclosure provides a method, apparatus, device, and storage medium for determining transition points. For audio segments with different spectral change situations in the audio, different types of candidate points can be selected as transition points, which improves the diversity and richness of the transition points in the audio. Furthermore, when making a beat video based on the transition points in the audio, the transition effect of the beat video can be improved. The technical solution of the present disclosure is as follows:
[0004] According to the first aspect of the embodiments of the present disclosure, a method for determining transition points is provided. The method includes:
[0005] Segment the audio based on the chromaticity features and rhythm features of multiple audio frames of the audio, obtaining multiple audio segments. The difference between the chromaticity features of the audio frames within the same audio segment is within a first preset range, and the difference between the rhythm features is within a second preset range;
[0006] Based on the spectral information of the multiple audio segments, determine the importance parameters of the multiple audio segments, where the importance parameters are used to reflect the spectral change situation of the audio segments;
[0007] For each audio segment, based on the importance parameter of the audio segment, determine a target candidate point type from three candidate point types: beat points, accent points, and mutation points, where the target candidate point type matches the importance parameter;
[0008] Determine the boundary time point of the target audio frame in the audio segment as the transition point of the audio segment, where the target audio frame includes the time points of the target candidate point type.
[0009] According to the second aspect of the embodiments of the present disclosure, a device for determining transition points is provided. The device includes:
[0010] A segmentation unit, configured to segment the audio based on the chromaticity features and rhythm features of each of multiple audio frames of the audio, to obtain multiple audio segments, wherein the difference between the chromaticity features of the audio frames within the same audio segment is within a first preset range, and the difference between the rhythm features is within a second preset range;
[0011] A parameter determination unit, configured to determine an importance parameter for each of the multiple audio segments based on the spectral information of each of the multiple audio segments, and the importance parameter is used to reflect the spectral change condition of the audio segment;
[0012] A candidate point type determination unit, configured to, for each of the audio segments, determine a target candidate point type from three candidate point types including a beat point, a downbeat point, and a mutation point based on the importance parameter of the audio segment, and the target candidate point type matches the importance parameter;
[0013] A transition point determination unit, configured to determine the boundary time point of a target audio frame in the audio segment as the transition point of the audio segment, and the target audio frame includes the time point of the target candidate point type.
[0014] In some embodiments, the candidate point type determination unit is configured to: for each of the audio segments, classify the audio segment into a target segment category among the multiple segment categories based on the importance parameter of the audio segment and the importance parameter ranges respectively corresponding to the multiple segment categories, and the importance parameter of the audio segment belongs to the importance parameter range corresponding to the target segment category; and use at least one type of candidate point corresponding to the target segment category among the three candidate point types as the target candidate point type.
[0015] In some embodiments, the multiple segment categories include a first segment category, a second segment category, and a third segment category, the importance parameter of the audio segment in the first segment category is greater than the importance parameter of the audio segment in the second segment category, and the importance parameter of the audio segment in the second segment category is greater than the importance parameter of the audio segment in the third segment category; the candidate point type determination unit is configured to: if the target segment category is the first segment category, use the mutation point among the three candidate point types as the target candidate point type; if the target segment category is the second segment category, use the beat point among the three candidate point types as the target candidate point type; if the target segment category is the third segment category, use the downbeat point among the three candidate point types as the target candidate point type.
[0016] In some embodiments, the segmentation unit is configured to: for each audio frame, perform feature splicing of chromaticity features and rhythm features on the audio frame to obtain the segmentation feature of the audio frame; perform change point detection on the segmentation features of each of the multiple audio frames to obtain at least one changed audio frame among the multiple audio frames, where the changed audio frame refers to an audio frame whose segmentation feature changes relative to the previous audio frame; and segment the audio with the at least one changed audio frame as the demarcation point to obtain the multiple audio segments.
[0017] In some embodiments, the spectral information includes spectral fluctuation intensity, spectral richness, and spectral centroid. The parameter determination unit is configured to: for each of the audio segments, perform weighted summation on the spectral fluctuation intensity, spectral richness, and spectral centroid of the audio segment to obtain the importance parameter of the audio segment.
[0018] In some embodiments, the apparatus further includes:
[0019] An audio frame type determination unit, configured to determine a first type of audio frame, a second type of audio frame, and a third type of audio frame from multiple audio frames of the audio. The first type of audio frame is an audio frame including a beat point, the second type of audio frame is an audio frame including a downbeat point, the third type of audio frame is an audio frame including a mutation point, and the signal energy of the third type of audio frame is greater than an energy threshold;
[0020] A target audio frame determination unit, configured to use the audio frames belonging to the target audio frame type in the audio segment as target audio frames, where the target audio frame type is the audio frame type that matches the target candidate point type among the first type of audio frame, the second type of audio frame, and the third type of audio frame.
[0021] In some embodiments, the audio frame type determination unit is configured to: resample the audio based on a first sampling rate to obtain a first audio signal, perform beat point detection and downbeat point detection on each of the multiple audio frames in the first audio signal to obtain the first type of audio frame and the second type of audio frame; resample the audio based on a second sampling rate to obtain a second audio signal, and perform mutation point detection on each of the multiple audio frames in the second audio signal to obtain the third type of audio frame, where the first sampling rate is less than the second sampling rate.
[0022] In some embodiments, the apparatus further includes: a rhythm determination unit, configured to for each audio frame, determine the rhythm feature of the audio frame based on the mutation points within a target period, where the target period is obtained by extending a preset duration forward and backward respectively with the audio frame as the center, and the rhythm feature is used to reflect the frequency of the mutation points.
[0023] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, the electronic device including:
[0024] one or more processors;
[0025] a memory for storing program codes executable by the processor;
[0026] The processor is configured to execute the program code to implement the above-mentioned method for determining the transition point.
[0027] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When the program code in the computer-readable storage medium is executed by a processor of an electronic device, the electronic device can execute the above-mentioned method for determining the transition point.
[0028] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, including a computer program, which implements the above-mentioned method for determining a transition point when executed by a processor.
[0029] An embodiment of the present disclosure provides a method for determining transition points. The method divides audio into multiple audio segments based on chromaticity features and rhythmic features, so that audio frames within the same audio segment have similar chromaticity features and rhythmic features, thereby achieving accurate segmentation of the audio; and determines an importance parameter of each audio segment for reflecting its spectral change, and then selects candidate points of matching types as transition points based on the importance parameter of each audio segment. In this way, for audio segments with different spectral changes in the audio, different types of candidate points can be selected as transition points, thereby improving the diversity and richness of the transition points in the audio, and thus improving the transition effect of the card point video when producing the card point video based on the transition points in the audio.
[0030] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.
[0032] Figure 1 It is a schematic diagram showing an implementation environment according to an exemplary embodiment.
[0033] Figure 2 The figure is a flowchart of a method for determining a transition point according to an exemplary embodiment.
[0034] Figure 3 It is a flowchart of another method for determining a transition point shown according to an exemplary embodiment.
[0035] Figure 4 It is a framework diagram of a method for determining a transition point shown according to an exemplary embodiment.
[0036] Figure 5 It is a block diagram of a device for determining a transition point shown according to an exemplary embodiment.
[0037] Figure 6 It is a block diagram of a terminal shown according to an exemplary embodiment.
[0038] Figure 7 It is a block diagram of a server shown according to an exemplary embodiment. Detailed implementation manners
[0039] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0040] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0041] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the present disclosure are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions. For example, the audio involved in the present disclosure is obtained under full authorization.
[0042] The method for determining a transition point provided in the embodiments of the present disclosure can be executed by an electronic device, and the electronic device is provided as at least one of a terminal and a server. Figure 1 It is a schematic diagram of an implementation environment provided in the embodiments of the present disclosure. See Figure 1, the implementation environment includes: a terminal 101 and a server 102. In an embodiment of the present disclosure, a target application is installed on the terminal 101, and the target application can be an application for making videos. The server 102 is the background server of the target application. In some embodiments, the terminal 101 is used to determine the transition points in the audio and make a beat video based on the transition points. In other embodiments, the server 102 is used to determine the transition points in the audio, send the transition points to the terminal 101, and the terminal 101 is used to make a beat video based on the transition points.
[0043] In some embodiments, when the terminal 101 selects a certain video and makes a beat video based on the video, it can trigger the process of determining the transition points in the audio, and then display the time points of at least one determined transition point based on the video. For example, the terminal 101 displays a progress bar of the video on the video production interface and displays the transition points on the progress bar. In response to a trigger operation on any transition point, the terminal 101 jumps the video interface to the video interface starting from the transition point.
[0044] The terminal 101 can be at least one of devices such as a smart phone, a smart watch, a desktop computer, a laptop computer, a virtual reality terminal, an augmented reality terminal, a wireless terminal, and a laptop portable computer. The terminal 101 has a communication function and can access a wired network or a wireless network. The terminal 101 can generally refer to one of multiple terminals. Those skilled in the art can know that the number of the above terminals can be more or less. The server 102 can be an independent physical server, or a server cluster or a distributed file system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. In some embodiments, the server 102 is directly or indirectly connected to the terminal 101 through a wired or wireless communication method, and the embodiments of the present disclosure do not limit this. Optionally, the number of the above servers 102 can be more or less, and the embodiments of the present disclosure do not limit this. Of course, the server 102 can also include other functional servers to provide more comprehensive and diverse services. Among them, the server 102 undertakes the main computing work, and the terminal 101 undertakes the secondary computing work; or, the server 102 undertakes the secondary computing work, and the terminal 101 undertakes the main computing work; or, the server 102 or the terminal 101 can separately undertake the computing work, and the embodiments of the present disclosure do not limit this.
[0045] Figure 2 is a flowchart of a method for determining transition points shown according to an exemplary embodiment, as Figure 2As shown, the method is executed by an electronic device, and the method includes the following steps:
[0046] In step S201, the electronic device segments the audio based on the chromaticity features and rhythm features of multiple audio frames of the audio, obtaining multiple audio segments. The difference between the chromaticity features of the audio frames within the same audio segment is within a first preset range, and the difference between the rhythm features is within a second preset range.
[0047] In the embodiments of the present disclosure, the multiple audio frames are obtained by frame-dividing the audio. Optionally, the electronic device frame-divides the audio based on a preset duration, obtaining multiple audio frames each having the preset duration.
[0048] In the embodiments of the present disclosure, both the chromaticity feature and the rhythm feature are feature vectors, respectively used to describe the features of the audio frame in audio chromaticity and in audio rhythm.
[0049] In step S202, the electronic device determines the importance parameter of each of the multiple audio segments based on the spectral information of each of the multiple audio segments, and the importance parameter is used to reflect the spectral change situation of the audio segment.
[0050] In the embodiments of the present disclosure, the spectral change situation of the audio segment includes at least one of soothing, ordinary, strong, etc., and is used to reflect whether the audio segment is a soothing audio segment, an ordinary audio segment, or a strong audio segment.
[0051] In step S203, for each audio segment, the electronic device determines a target candidate point type from three candidate point types of beat points, downbeat points, and mutation points based on the importance parameter of the audio segment, and the target candidate point type matches the importance parameter.
[0052] In the embodiments of the present disclosure, the three candidate point types can be respectively matched with different importance parameters. Optionally, the importance parameter matched by the mutation point is greater than the importance parameter matched by the beat point, and the importance parameter matched by the beat point is greater than the importance parameter matched by the downbeat point. And at least two candidate point types can also match the same importance parameter, which is not specifically limited here. For example, the beat point and the downbeat point match the same importance parameter.
[0053] In step S204, the electronic device determines the boundary time point of the target audio frame in the audio segment as the transition point of the audio segment, and the target audio frame includes the time point of the target candidate point type.
[0054] In the embodiments of the present disclosure, the boundary time point is the end point of the target audio frame or the time point at a preset duration from the end point in the target audio frame. For example, the duration of an audio frame is 10 ms, and the boundary time point is the 9th ms, which is 1 ms from the end point. The boundary time point of the target audio frame represents one of the beat points, accent points, and mutation points in the target audio frame. If the target candidate point type is a beat point, the boundary time point of the target audio frame represents a beat point; if the target candidate point type is an accent point, the boundary time point of the target audio frame represents an accent point; if the target candidate point type is a mutation point, the boundary time point of the target audio frame represents a mutation point.
[0055] In the embodiments of the present disclosure, the transition point can be used to produce a beat-matching video. The transition point is the time point for picture switching or video interface switching in the beat-matching video.
[0056] The embodiments of the present disclosure provide a method for determining a transition point. The method divides an audio into multiple audio segments based on chromaticity features and rhythm features, such that the audio frames within the same audio segment have similar chromaticity features and rhythm features, that is, accurate segmentation of the audio is achieved; and an importance parameter for reflecting the spectral change of each audio segment is determined. Furthermore, based on the importance parameter of each audio segment, a candidate point with a matching type is selected as the transition point. In this way, for audio segments with different spectral change situations in the audio, candidate points of different types can be selected as the transition point, which improves the diversity and richness of the transition points in the audio. Furthermore, when producing a beat-matching video based on the transition points in the audio, the transition effect of the beat-matching video can be improved.
[0057] In some embodiments, for each audio segment, based on the importance parameter of the audio segment, to determine the target candidate point type from three candidate point types of beat points, accent points, and mutation points, includes: for each audio segment, based on the importance parameter of the audio segment and the importance parameter ranges respectively corresponding to multiple segment categories, classify the audio segment into the target segment category among the multiple segment categories, where the importance parameter of the audio segment belongs to the importance parameter range corresponding to the target segment category; take at least one type of candidate point corresponding to the target segment category among the three candidate point types as the target candidate point type.
[0058] In the embodiments of the present disclosure, determining the segment category to which each audio segment belongs respectively in this way, and then determining the target candidate point type based on the segment category, improves the efficiency and accuracy of determining the target candidate point type.
[0059] In some embodiments, multiple segment categories include a first segment category, a second segment category, and a third segment category, the importance parameter of the audio segment of the first segment category is greater than the importance parameter of the audio segment of the second segment category, and the importance parameter of the audio segment of the second segment category is greater than the importance parameter of the audio segment of the third segment category; at least one type of candidate point corresponding to the target segment category among the three candidate point types is used as the target candidate point type, including at least one of the following: if the target segment category is the first segment category, the mutation point among the three candidate point types is used as the target candidate point type; if the target segment category is the second segment category, the beat point among the three candidate point types is used as the target candidate point type; if the target segment category is the third segment category, the rebeat point among the three candidate point types is used as the target candidate point type.
[0060] In the embodiment of the present disclosure, if the importance parameter of the audio clip of the first segment category is large, it means that the audio clip of this category is a relatively strong audio clip, and the distribution of mutation points in the audio is generally uneven, and can reflect the intensity of the audio, and then the unevenly distributed mutation points are used as transition points, and a card point video is produced based on the transition point, which can produce a video effect that is relatively "exciting" following the change of the mutation point. If the importance parameter of the audio clip of the second segment category is in the middle, it means that the audio clip of this category is an ordinary audio clip, and the beat points in the audio are evenly distributed, and the number of beat points is greater than the number of rebeat points, and then the beat points are used as transition points, and a card point video is produced based on the transition point, which can achieve a better transition effect. If the importance parameter of the audio clip of the third segment category is small, it means that the audio clip of this category is a relatively soothing audio clip, and the number of rebeat points in the audio is small, and then the rebeat points are used as transition points, and a card point video is produced based on the transition point, which can achieve a better transition effect.
[0061] In some embodiments, the audio is segmented based on the chromaticity features and rhythmic features of each of the multiple audio frames to obtain multiple audio clips, including: for each audio frame, the chromaticity features and rhythmic features of the audio frame are concatenated to obtain the segmentation features of the audio frame; change point detection is performed on the segmentation features of each of the multiple audio frames to obtain at least one changed audio frame among the multiple audio frames, where the changed audio frame refers to an audio frame whose segmentation features have changed relative to the previous audio frame; and the audio is segmented with at least one changed audio frame as a dividing point to obtain multiple audio clips.
[0062] In the embodiment of the present disclosure, the segmentation feature integrates the chromaticity feature and rhythm feature of the audio frame, and then based on the segmentation feature, it can detect the audio frame whose chromaticity feature and rhythm feature have changed, thereby ensuring the accuracy of the detected changed audio frame, and then segmenting is performed with the changed audio frame as the dividing point, thereby improving the rationality and accuracy of the segmentation.
[0063] In some embodiments, the spectral information includes spectral fluctuation intensity, spectral richness, and spectral centroid. Based on the spectral information of each of the multiple audio segments, importance parameters of each of the multiple audio segments are determined, including: for each audio segment, performing a weighted sum of the spectral fluctuation intensity, spectral richness, and spectral centroid of the audio segment to obtain the importance parameter of the audio segment.
[0064] In the embodiments of the present disclosure, fusion of multiple spectral information such as spectral fluctuation intensity, spectral richness, and spectral centroid is achieved, and thus, based on this importance parameter, the spectral change situation of the audio segment can be accurately reflected.
[0065] In some embodiments, the method further includes: determining a first type of audio frame, a second type of audio frame, and a third type of audio frame from multiple audio frames of the audio. The first type of audio frame is an audio frame including a beat point, the second type of audio frame is an audio frame including a downbeat point, and the third type of audio frame is an audio frame including a mutation point. The signal energy of the third type of audio frame is greater than an energy threshold; using the audio frames belonging to the target audio frame type in the audio segment as target audio frames, where the target audio frame type is the audio frame type that matches the target candidate point type among the first type of audio frame, the second type of audio frame, and the third type of audio frame.
[0066] In the embodiments of the present disclosure, first, three types of audio frames including beat points, downbeat points, and mutation points are determined, that is, detection of beat points, downbeat points, and mutation points is performed at the granularity of audio frames, which improves the detection efficiency; and it is convenient to subsequently determine the audio frame type that matches the target candidate point type, thereby improving the efficiency of determining the target audio frame type.
[0067] In some embodiments, determining a first type of audio frame, a second type of audio frame, and a third type of audio frame from multiple audio frames of the audio includes: resampling the audio based on a first sampling rate to obtain a first audio signal, performing beat point detection and downbeat point detection on each of the multiple audio frames in the first audio signal to obtain the first type of audio frame and the second type of audio frame; resampling the audio based on a second sampling rate to obtain a second audio signal, and performing mutation point detection on each of the multiple audio frames in the second audio signal to obtain the third type of audio frame, where the first sampling rate is less than the second sampling rate.
[0068] In the embodiments of the present disclosure, the requirements for the sampling rate of the audio signal for beat point and downbeat point detection are relatively low, that is, accurate beat points and downbeat points can be obtained based on a relatively low sampling rate. Therefore, detecting beat points and downbeat points based on an audio signal with a relatively low sampling rate can improve the detection efficiency and reduce resource consumption. The mutation point detection requires a relatively high sampling rate for the audio signal. By using an audio signal with a relatively high sampling rate, the mutation situation of the audio in the high-frequency range can be detected, thereby ensuring the effect of mutation point detection.
[0069] In some embodiments, the method further includes: for each audio frame, based on the mutation points within the target period, determining the rhythm feature of the audio frame. The target period is centered on the audio frame and extends forward and backward by a preset duration respectively. The rhythm feature is used to reflect the frequency of mutation points.
[0070] In the embodiments of the present disclosure, since the mutation points in the audio can effectively reflect the rhythm speed of the audio mutation, and then determining the rhythm feature based on the mutation points can improve the rationality and accuracy of the determined rhythm feature.
[0071] The above Figure 2 is the basic process for determining the transition point. Next, based on Figure 3 the process of determining the transition point will be further elaborated. Refer to Figure 3 Figure 3 FIG. [FIG. NUMBER] is a flowchart of a method for determining a transition point according to an exemplary embodiment. This method is executed by an electronic device, and the method includes the following steps:
[0072] In step S301, the electronic device determines a first type of audio frame, a second type of audio frame, and a third type of audio frame from multiple audio frames of the audio. The first type of audio frame is an audio frame including a beat point, the second type of audio frame is an audio frame including a downbeat point, the third type of audio frame is an audio frame including a mutation point, and the signal energy of the third type of audio frame is greater than an energy threshold.
[0073] In the embodiments of the present disclosure, the electronic device performs beat point detection, downbeat point detection, and mutation point detection on multiple audio frames respectively, and obtains at least one audio frame of the first type, at least one audio frame of the second type, and at least one audio frame of the third type.
[0074] In some embodiments, the electronic device performs beat point and downbeat point detection through a deep learning detection model. Through this deep learning detection model, it can be determined whether a certain audio frame is an audio frame including a beat point, and it can also be determined whether a certain audio frame is an audio frame including a downbeat point. Performing beat point and downbeat point detection based on this deep learning detection model can improve the detection efficiency and accuracy.
[0075] In some embodiments, the electronic device detects the mutation points by detecting the signal energy of multiple audio frames. The electronic device regards the audio frames with signal energy greater than the energy threshold as the third type of audio frames.
[0076] In the embodiments of the present disclosure, the electronic device uses an audio frame as a detection unit to determine the beat points, accent points, and mutation points. After the electronic device determines three types of audio frames, it uses the boundary time points of the first type of audio frames as the beat points, the boundary time points of the second type of audio frames as the accent points, and the boundary time points of the third type of audio frames as the mutation points, which improves the efficiency of determining the beat points, accent points, and mutation points.
[0077] In the embodiments of the present disclosure, the signal energy can be the decibel value in the time-domain spectrum or the frequency value in the frequency-domain spectrum. The signal energy of any audio frame is the comprehensive value of the signal energies of multiple sampling points in this audio frame, and this comprehensive value can be an average value or a sum value.
[0078] In some embodiments, the energy threshold is a fixed value. In other embodiments, the energy threshold can be adjusted in real time. For example, the energy threshold is the sum value between the signal energy of the previous audio frame of this audio frame and the first preset difference, that is, the audio frame with a large energy gap from the previous audio frame is regarded as the audio frame where the mutation point is located. Another example is that the energy threshold is the sum value between the signal energy of the background audio of this audio frame and the second preset difference, that is, the audio frame with a large energy gap from the background audio is regarded as the audio frame where the mutation point is located. Adjusting the energy threshold in real time like this can further improve the accuracy of the determined mutation points.
[0079] In some embodiments, the electronic device performs beat point, accent point, and mutation point detection based on the audio signals obtained by sampling the audio at different sampling rates. Correspondingly, the process of the above-mentioned electronic device determining the first type of audio frames, the second type of audio frames, and the third type of audio frames from multiple audio frames of the audio includes the following steps: The electronic device resamples the audio based on the first sampling rate to obtain the first audio signal, and performs beat point detection and accent point detection on multiple audio frames in the first audio signal respectively to obtain the first type of audio frames and the second type of audio frames. The electronic device resamples the audio based on the second sampling rate to obtain the second audio signal, and performs mutation point detection on multiple audio frames in the second audio signal to obtain the third type of audio frames, where the first sampling rate is less than the second sampling rate.
[0080] In the embodiments of the present disclosure, the first sampling rate and the second sampling rate can be set and changed as needed, and no specific limitation is made here. For example, the first sampling rate is 16 kHz and the second sampling rate is 44.1 kHz.
[0081] In the embodiments of the present disclosure, the requirements for the sampling rate of the audio signal for detecting beat points and accent points are relatively low, that is, accurate beat points and accent points can be obtained based on a relatively low sampling rate. Therefore, detecting beat points and accent points based on an audio signal with a relatively low sampling rate can improve the detection efficiency and reduce resource consumption. However, the requirements for the sampling rate of the audio signal for detecting mutation points are relatively high. By using an audio signal with a relatively high sampling rate, the mutation situation of the audio in the high-frequency range can be detected, thereby ensuring the detection effect of mutation points.
[0082] In step S302, the electronic device determines the chromaticity features and rhythm features of each of the multiple audio frames of the audio.
[0083] In some embodiments, the rhythm feature is used to reflect the frequency of mutation points. Correspondingly, the process by which the electronic device determines the rhythm features of each of the multiple audio frames includes the following steps: For each audio frame, the electronic device determines the rhythm feature of the audio frame based on the mutation points within the target period, and the target period is obtained by extending a preset duration forward and backward respectively centered on the audio frame. In this embodiment, since the mutation points in the audio can effectively reflect the rhythm speed of the audio mutation, and then determining the rhythm feature based on the mutation points can improve the rationality and accuracy of the determined rhythm feature.
[0084] In the embodiments of the present disclosure, the target period is centered on the audio frame. Further, the center point of the audio frame can be used as the center point of the target period, or the starting point of the audio frame can be used as the center point of the target period, or the ending point of the audio frame can be used as the center point of the target period, and no specific limitation is made here. In the embodiments of the present disclosure, the preset duration can be set and changed as needed, and no specific limitation is made here. Among them, the preset duration extended forward and the preset duration extended backward can be the same or different.
[0085] In the embodiments of the present disclosure, the electronic device extracts the rhythm feature of the audio frame through the rhythm detection module based on the mutation points within the target period, and obtains the rhythm feature of the audio frame. Optionally, the electronic device inputs all the time points within the target period into the rhythm detection module, and the time points within the target period carry tags indicating whether they are mutation points. The rhythm detection module is used to output the rhythm feature of the audio frame based on the time points within the target period, which can improve the efficiency of determining the rhythm feature of the audio frame.
[0086] In some other embodiments, for each audio frame, the electronic device determines the rhythm feature of the audio frame based on the beat points within the target period, and the rhythm feature is used to reflect the frequency of the beat points. In some other embodiments, for each audio frame, the electronic device determines the rhythm feature of the audio frame based on the downbeat points within the target period, and the rhythm feature is used to reflect the frequency of the downbeat points. Since the frequencies of both the beat points and the downbeat points can reflect the speed of the rhythm, determining the rhythm feature based on the beat points or the downbeat points can improve the rationality and accuracy of the determined rhythm feature.
[0087] In the embodiments of the present disclosure, the chromaticity feature is a feature vector, and the chromaticity feature of each audio frame indicates the signal energy of each of the multiple pitch classes present in the audio frame. In some embodiments, the process by which the electronic device determines the chromaticity features of the multiple audio frames includes the following steps: The electronic device separately extracts the chromaticity features of the multiple audio frames in the second audio signal to obtain the chromaticity features of the multiple audio frames. It should be noted that a relatively high sampling rate of the audio signal is required for extracting the chromaticity feature, and in the embodiments of the present disclosure, with an audio signal having a relatively high sampling rate, the energy of the pitch classes in the high-frequency range of the audio can be extracted, thereby ensuring the effect of chromaticity feature extraction.
[0088] It should be noted that step S302 may be executed before step S303, or may be executed in advance, such as before triggering the process of determining the transition point. The embodiments of the present disclosure do not specifically limit the execution timing of step S302. By determining the audio features and chromaticity features of the audio frames in advance, when segmenting the audio, the previously determined chromaticity features and rhythm features can be directly applied, improving the efficiency of audio segmentation and thereby improving the efficiency of determining the transition point.
[0089] In step S303, the electronic device segments the audio based on the chromaticity features and rhythm features of the multiple audio frames of the audio to obtain multiple audio segments. The difference between the chromaticity features of the audio frames within the same audio segment is within the first preset range, and the difference between the rhythm features is within the second preset range.
[0090] In the embodiments of the present disclosure, if the difference between the chromaticity features of the audio frames within the same audio segment is within the preset range and the difference between the rhythm features is within the preset range, it indicates that the audio frames within the same audio segment have similar chromaticity features and audio features. The first preset range and the second preset range may be the same or different, and the first preset range and the second preset range can be set and changed as needed, and no specific limitation is made here.
[0091] In some embodiments, the process of segmenting the audio by the electronic device based on the chromaticity features and rhythm features of multiple audio frames of the audio includes the following steps: For each audio frame, the electronic device performs feature splicing on the chromaticity feature and rhythm feature of the audio frame to obtain the segmentation feature of the audio frame; the electronic device performs change point detection on the segmentation features of multiple audio frames to obtain at least one changed audio frame among the multiple audio frames, where the changed audio frame refers to an audio frame whose segmentation feature changes relative to the previous audio frame, and the electronic device segments the audio with the at least one changed audio frame as the demarcation point to obtain multiple audio segments.
[0092] In the embodiments of the present disclosure, a changed audio frame refers to an audio frame whose segmentation feature changes relative to the previous audio frame. Further, the changed audio frame refers to an audio frame whose gap between the segmentation feature and the segmentation feature of the previous audio frame is greater than the gap threshold.
[0093] In the embodiments of the present disclosure, both the chromaticity feature and the rhythm feature are feature vectors; correspondingly, splicing the chromaticity feature and the rhythm feature of the audio frame means splicing the vectors of the chromaticity feature and the rhythm feature. The segmentation features of multiple audio frames can be stored in a feature time series, and the multiple audio frames are arranged in chronological order, and the feature time series can be expressed as [y1, y2,..., yT], where y1, y2, and yT respectively represent the first audio frame, the second audio frame, and the Tth audio frame, and T is an integer greater than 1.
[0094] In some embodiments, the electronic device performs change point detection on the segmentation features of multiple audio frames through a change point detection algorithm to obtain at least one changed audio frame. For example, the change point detection algorithm is a kernel change point detection algorithm (Kernel change point detection). Through this kernel change point detection algorithm, K changed segmentation features can be found from the multiple segmentation features in the feature time series, and then K changed audio frames can be found, where K is an integer greater than 0. Furthermore, the audio can be divided into K + 1 audio segments. K can be determined according to the duration of the audio. For example, an audio segment is divided every 10s on average, or the most suitable number of segments for segmentation can be automatically determined according to the change point detection algorithm, or it can be manually set as needed, and no specific limitation is made here. In the embodiments of the present disclosure, the changed audio frame is an audio frame whose segmentation feature changes, so the changed audio frame is divided into the subsequent audio segment to improve the rationality and accuracy of segment division.
[0095] In an embodiment of the present disclosure, the segmented feature combines the chromaticity feature and the rhythm feature of the audio frame. Further, based on the segmented feature, the audio frame in which the chromaticity feature and the rhythm feature change can be detected, ensuring the accuracy of the detected changed audio frame. Then, segmentation is performed with the changed audio frame as the demarcation point, improving the rationality and accuracy of the segmentation.
[0096] In step S304, the electronic device determines the importance parameter of each of the multiple audio segments based on the spectral information of each of the multiple audio segments, and the importance parameter is used to reflect the spectral change condition of the audio segment.
[0097] In an embodiment of the present disclosure, the spectral information includes at least one of spectral fluctuation intensity, spectral richness, and spectral centroid. The spectral fluctuation intensity is used to reflect the fluctuation degree of the signal in the spectrum, the spectral richness is used to reflect the richness of various types of audio such as high frequency, low frequency, human voice, and musical instrument sound, and the spectral centroid is used to reflect the brightness of the sound in the audio.
[0098] In some embodiments, if the spectral information includes spectral fluctuation intensity, spectral richness, and spectral centroid, the process by which the above-mentioned electronic device determines the importance parameter of each of the multiple audio segments based on the spectral information of each of the multiple audio segments includes the following steps: For each audio segment, the electronic device performs a weighted sum of the spectral fluctuation intensity, spectral richness, and spectral centroid of the audio segment to obtain the importance parameter of the audio segment. In this embodiment, the fusion of multiple spectral information such as spectral fluctuation intensity, spectral richness, and spectral centroid is achieved, and further, based on the importance parameter, the spectral change condition of the audio segment can be accurately reflected.
[0099] In some other embodiments, if the spectral information includes two of spectral fluctuation intensity, spectral richness, and spectral centroid, a weighted sum of these two of the audio segment is performed to obtain the importance parameter of the audio segment. If the spectral information includes one of spectral fluctuation intensity, spectral richness, and spectral centroid, this item is directly used as the importance parameter of the audio segment.
[0100] For example, the electronic device performs a weighted sum of the spectral fluctuation intensity, spectral richness, and spectral centroid of the audio segment through the following formula (1) to obtain the importance parameter of the audio segment.
[0101] Ipm(k) = a1 * SF(k) + a2 * SR(k) + a3 * SC(k) (1)
[0102] Among them, Ipm(k) represents the importance parameter of the k-th audio frame, a1 represents the weighted parameter of the spectral fluctuation intensity, SF(k) represents the spectral fluctuation intensity of the k-th audio frame, a2 represents the weighted parameter of the spectral richness, SR(k) represents the spectral richness of the k-th audio frame, a3 represents the weighted parameter of the spectral centroid, and SC(k) represents the spectral centroid of the k-th audio frame.
[0103] In step S305, for each audio segment, the electronic device classifies the audio segment into a target segment category among multiple segment categories based on the importance parameter of the audio segment and the importance parameter ranges respectively corresponding to the multiple segment categories, where the importance parameter of the audio segment belongs to the importance parameter range corresponding to the target segment category.
[0104] In some embodiments, the electronic device can pre-establish a correspondence relationship, which stores multiple segment categories and the importance parameter ranges respectively corresponding to the multiple segment categories, and then directly finds the target segment category to which the audio segment belongs from the first correspondence relationship based on the importance parameter of the audio segment.
[0105] In other embodiments, the electronic device can adjust the importance parameter ranges respectively corresponding to the multiple segment categories in real time based on the importance parameters of the multiple audio segments in the audio. For example, if the maximum importance parameter among the importance parameters of the multiple audio segments is less than the first parameter threshold, it indicates that the importance parameters of the multiple audio segments are generally small, then the importance parameter ranges respectively corresponding to the multiple segment categories are adjusted to be smaller, that is, both the upper limit value and the lower limit value of the importance parameter range are adjusted to be smaller, and then the audio segments are classified based on the adjusted importance parameter ranges. Another example is that if the minimum importance parameter among the importance parameters of the multiple audio segments is greater than the second parameter threshold, it indicates that the importance parameters of the multiple audio segments are generally large, then the importance parameter ranges respectively corresponding to the multiple segment categories are adjusted to be larger, that is, both the upper limit value and the lower limit value of the importance parameter range are adjusted to be larger, and then the audio segments are classified based on the adjusted importance parameter ranges. Another example is that if the maximum importance parameter among the importance parameters of the multiple audio segments is less than the first parameter threshold and the minimum importance parameter is greater than the second parameter threshold, it indicates that the importance parameters of the multiple audio segments as a whole correspond to a relatively narrow importance parameter range, then the ranges of the importance parameter ranges respectively corresponding to the multiple segment categories are narrowed, that is, each segment category corresponds to a smaller importance parameter range, thereby improving the accuracy of audio segment classification. Adjusting the importance parameter ranges respectively corresponding to the multiple segment categories in real time like this can improve the rationality and accuracy of audio segment classification.
[0106] In step S306 , the electronic device selects at least one type of candidate point corresponding to the target segment category among the three candidate point types as a target candidate point type.
[0107] In the embodiment of the present disclosure, the candidate point type corresponding to any segment category can be set and changed as needed, and is not specifically limited here.
[0108] In some embodiments, the multiple segment categories include a first segment category, a second segment category, and a third segment category, and the importance parameters of the segments of the first segment category are greater than the importance parameters of the segments of the second segment category, and the importance parameters of the segments of the second segment category are greater than the importance parameters of the segments of the third segment category. That is, the lower limit of the importance parameter range corresponding to the first segment category is greater than the upper limit of the importance parameter range corresponding to the second segment category, and the lower limit of the importance parameter range corresponding to the second segment category is greater than the upper limit of the importance parameter range corresponding to the third segment category.
[0109] Correspondingly, the process of the above-mentioned electronic device taking at least one type of candidate point corresponding to the target segment category among the three candidate point types as the target candidate point type includes at least one of the following items: if the target segment category is the first segment category, the electronic device takes the mutation point among the three candidate point types as the target candidate point type; if the target segment category is the second segment category, the electronic device takes the beat point among the three candidate point types as the target candidate point type; if the target segment category is the third segment category, the electronic device takes the rebeat point among the three candidate point types as the target candidate point type.
[0110] In the embodiment of the present disclosure, if the importance parameter of the audio clip of the first segment category is large, it means that the audio clip of this category is a relatively strong audio clip, and the mutation points are generally unevenly distributed in the audio and can reflect the intensity of the audio. Then, the unevenly distributed mutation points are used as transition points, and a card point video is produced based on the transition points, which can produce a video effect that is relatively "exciting" following the changes in the mutation points. If the importance parameter of the audio clip of the second segment category is in the middle, it means that the audio clip of this category is an ordinary audio clip, and the beat points in the audio are generally evenly distributed, and the number of beat points is greater than the number of rebeat points. Then, the beat points are used as transition points, and a card point video is produced based on the transition points, which can achieve a stable transition effect. If the importance parameter of the audio clip of the third segment category is small, it means that the audio clip of this category is a relatively soothing audio clip, and the number of rebeat points in the audio is small. Then, the rebeat points are used as transition points, and a card point video is produced based on the transition points, which can achieve a relatively soothing transition effect.
[0111] In some embodiments, if the target segment category is the second segment category or the third segment category, the electronic device also uses the boundary time point of the audio frames of the third type in the audio segment, that is, the mutation point, as the display point of the video special effect. Further, the electronic device determines the display point of the video special effect based on the signal energy strength of the audio frames of the third type. For example, the boundary time point of the audio frames with signal energy greater than the target energy threshold in the audio frames of the third type is used as the display point of the video special effect, and the video special effect includes but is not limited to the screen flashing effect or the jitter effect. In the embodiments of the present disclosure, by designing different video special effects for different audio segments, the diversity of the beat videos can be improved.
[0112] In the embodiments of the present disclosure, based on the above steps S305-S306, the process of determining the target candidate point type from the three candidate point types of the beat point, the strong beat point, and the mutation point for each audio segment based on the importance parameter of the audio segment is realized. In this way, the segment category to which each audio segment belongs is determined, and then the target candidate point type is determined based on this segment category, which improves the efficiency and accuracy of determining the target candidate point type.
[0113] It should be noted that the above steps S305-S306 are only an optional implementation manner for determining the target candidate point type. The embodiments of the present disclosure can also determine the target candidate point type through other optional implementation manners, which are not specifically limited herein. For example, for the audio segments with importance parameters greater than the parameter threshold, the mutation point is used as the target candidate point type, and for the audio segments with importance parameters not greater than the parameter threshold, both the beat point and the strong beat point are used as the target candidate point types.
[0114] In step S307, the electronic device uses the audio frames belonging to the target audio frame type in the audio segment as the target audio frames, and the target audio frame type is the audio frame type that matches the target candidate point type among the audio frames of the first type, the audio frames of the second type, and the audio frames of the third type.
[0115] In the embodiments of the present disclosure, if the target candidate point type is the beat point, the target audio frame type is the first type, and the target audio frames are the audio frames of the first type; if the target candidate point type is the strong beat point, the target audio frame type is the second type, and the target audio frames are the audio frames of the second type; if the target candidate point type is the mutation point, the target audio frame type is the third type, and the target audio frames are the audio frames of the third type.
[0116] In the embodiments of the present disclosure, the serial number of step S301 is only for convenience of description. Step S301 may be executed before step S302, or may be executed after step S302 and before any step of step S307, or may be executed in advance, such as before the process of triggering the determination of the transition point. The embodiments of the present disclosure do not specifically limit the execution timing of step S301. By determining these three types of audio frames in the audio in advance, when determining the target audio frame, the three types of audio frames determined in advance can be directly applied to improve the efficiency of determining the target audio frame, and further improve the efficiency of determining the transition point.
[0117] In the embodiments of the present disclosure, after determining multiple types of audio frames through step S301, the type of the target audio frame is determined from the multiple types of audio frames in step S307. First, three types of audio frames including beat points, downbeat points, and mutation points are determined, that is, the detection of beat points, downbeat points, and mutation points is performed from the granularity of audio frames, which can improve the detection efficiency and facilitate subsequent determination of the audio frame type matching the target candidate point type, thereby improving the efficiency of determining the type of the target audio frame.
[0118] In step S308, the electronic device determines the boundary time point of the target audio frame in the audio segment as the transition point of the audio segment, and the target audio frame includes the time point of the target candidate point type.
[0119] In the embodiments of the present disclosure, the time point of the target candidate point type included in the target audio frame is the boundary time point of the target audio frame. If the type of the target audio frame is the first type, the time point of the target candidate point type is the boundary time point representing the beat point; if the type of the target audio frame is the second type, the time point of the target candidate point type is the boundary time point representing the downbeat point; if the type of the target audio frame is the third type, the time point of the target candidate point type is the boundary time point representing the mutation point.
[0120] See Figure 4 , Figure 4A framework diagram of a method for determining transition points provided by an embodiment of the present disclosure. The framework includes a resampling module (Resampling), a beat point and downbeat point detection module (Beats Detection), an onset point detection module (Onsets Detection), a segmentation module (Segmentation), a segment importance determination module (SegmentImportance), and a fusion module (Reformat). Among them, the electronic device inputs the audio into the resampling module, and the resampling module resamples the audio to generate a first audio signal and a second audio signal with sampling rates of 16 kHz and 44.1 kHz respectively. The beat point and downbeat point detection module detects the beat points and downbeat points of multiple audio frames in the first audio signal respectively, and obtains the first type of audio frames and the second type of audio frames. The boundary time points of the first type of audio frames are used as beat points, and the boundary time points of the second type of audio frames are used as downbeat points. The onset point detection module detects the onset points of multiple audio frames in the second audio signal, and obtains the third type of audio frames. The boundary time points of the third type of audio frames are used as onset points. Through the segmentation module, the tempo of multiple audio frames is generated based on the onset points in the audio, and at the same time, the chroma features of multiple audio frames in the second audio signal are extracted by the segmentation module to obtain the chroma features of each of the multiple audio frames. Based on the chroma features and tempo of each of the multiple audio frames, the audio is segmented to obtain multiple audio segments. Through the segment importance determination module, for each audio segment, based on information such as the spectral flux, spectral richness, and spectral centroid / brightness in the audio frame, the importance parameter of the audio segment is determined. Based on this importance parameter, the multiple audio segments can be divided into types such as gentle, ordinary, and intense. Through the fusion module, information such as the beat points (beats), downbeat points (downbeats), onset points (onsets), and audio segments (segments) of the audio is fused to generate multiple audio segments of the audio and the information of each audio segment. For each audio segment, the information of the audio segment includes the union of all change point times such as the importance parameter of the audio segment, the beat points, downbeat points, and onset points within the audio segment, and includes the mutation type (strong or weak) and beat point type (non-beat point, beat point, downbeat point) corresponding to each time point.
[0121] Embodiments of the present disclosure provide a method for determining transition points. This method provides multiple types of candidate points that can serve as transition points, including uniform beat points, accent points, and non-uniform mutation points. Additionally, video special effects can be created based on the energy intensity of the mutation points. Furthermore, the method segments the audio by combining the chromaticity feature and the rhythm feature for reflecting the tempo, obtaining multiple audio segments, and determines the importance of each audio segment. By combining the above information, embodiments of the present disclosure can generate a variety of diverse beat-matching videos.
[0122] Embodiments of the present disclosure provide an overall solution for a music understanding algorithm that combines multiple types of beat-matching and audio segmentation, an audio segmentation method based on the rhythm feature and chromaticity feature of the audio, and a beat-matching video design based on music understanding information. By combining the music understanding algorithm of multiple types of audio beat-matching and audio segmentation, rich and variable audio beat-matching information is provided, which can significantly enhance the diversity of beat-matching video production.
[0123] Embodiments of the present disclosure provide a method for determining transition points. This method divides the audio into multiple audio segments based on the chromaticity feature and rhythm feature, such that the audio frames within the same audio segment have similar chromaticity features and rhythm features, thereby achieving accurate segmentation of the audio. And it determines the importance parameter for each audio segment to reflect its spectral change situation. Then, based on the importance parameter of each audio segment, candidate points with a matching type are selected as transition points. In this way, for audio segments with different spectral change situations in the audio, different types of candidate points can be selected as transition points, improving the diversity and richness of the transition points in the audio. Furthermore, when making a beat-matching video based on the transition points in the audio, the transition effect of the beat-matching video can be improved.
[0124] Figure 5 is a block diagram of a device for determining transition points shown according to an exemplary embodiment. Referring to Figure 5 , the device includes:
[0125] A segmentation unit 501, configured to segment the audio based on the chromaticity feature and rhythm feature of each of multiple audio frames of the audio, obtaining multiple audio segments. The difference between the chromaticity features of the audio frames within the same audio segment is within a first preset range, and the difference between the rhythm features is within a second preset range;
[0126] A parameter determination unit 502, configured to determine the importance parameter of each of the multiple audio segments based on the spectral information of each of the multiple audio segments. The importance parameter is used to reflect the spectral change situation of the audio segment;
[0127] A candidate point type determination unit 503, configured to, for each audio segment, determine a target candidate point type from three candidate point types, namely, beat points, accent points, and mutation points, based on the importance parameter of the audio segment, where the target candidate point type matches the importance parameter;
[0128] A transition point determination unit 504, configured to determine the boundary time point of a target audio frame in the audio segment as the transition point of the audio segment, where the target audio frame includes the time point of the target candidate point type.
[0129] In some embodiments, the candidate point type determination unit 503 is configured to: for each audio segment, classify the audio segment into a target segment category among multiple segment categories based on the importance parameter of the audio segment and the importance parameter ranges respectively corresponding to the multiple segment categories, where the importance parameter of the audio segment belongs to the importance parameter range corresponding to the target segment category; and use at least one type of candidate point corresponding to the target segment category among the three candidate point types as the target candidate point type.
[0130] In some embodiments, the multiple segment categories include a first segment category, a second segment category, and a third segment category, the importance parameter of the audio segment in the first segment category is greater than the importance parameter of the audio segment in the second segment category, and the importance parameter of the audio segment in the second segment category is greater than the importance parameter of the audio segment in the third segment category; the candidate point type determination unit 503 is configured to: if the target segment category is the first segment category, use the mutation point among the three candidate point types as the target candidate point type; if the target segment category is the second segment category, use the beat point among the three candidate point types as the target candidate point type; if the target segment category is the third segment category, use the accent point among the three candidate point types as the target candidate point type.
[0131] In some embodiments, a segmentation unit 501 is configured to: for each audio frame, perform feature splicing on the chromaticity feature and the rhythm feature of the audio frame to obtain the segmentation feature of the audio frame; perform change point detection on the segmentation features of the multiple audio frames to obtain at least one changed audio frame among the multiple audio frames, where the changed audio frame refers to an audio frame whose segmentation feature changes relative to the previous audio frame; and segment the audio with at least one changed audio frame as the demarcation point to obtain multiple audio segments.
[0132] In some embodiments, the spectral information includes spectral fluctuation intensity, spectral richness, and spectral centroid, and a parameter determination unit 502 is configured to: for each audio segment, perform weighted summation on the spectral fluctuation intensity, spectral richness, and spectral centroid of the audio segment to obtain the importance parameter of the audio segment.
[0133] In some embodiments, the apparatus further includes:
[0134] an audio frame type determining unit configured to determine, from a plurality of audio frames of the audio, a first type of audio frame, a second type of audio frame, and a third type of audio frame, wherein the first type of audio frame is an audio frame including a beat point, the second type of audio frame is an audio frame including a rebeat point, and the third type of audio frame is an audio frame including a mutation point, wherein signal energy of the third type of audio frame is greater than an energy threshold;
[0135] The target audio frame determination unit is configured to take an audio frame in the audio segment that belongs to a target audio frame type as a target audio frame, and the target audio frame type is an audio frame type that matches the target candidate point type among the first type of audio frame, the second type of audio frame, and the third type of audio frame.
[0136] In some embodiments, the audio frame type determination unit is configured to: resample the audio based on a first sampling rate to obtain a first audio signal, perform beat point detection and rebeat point detection on multiple audio frames in the first audio signal, respectively, to obtain a first type of audio frame and a second type of audio frame; resample the audio based on a second sampling rate to obtain a second audio signal, perform mutation point detection on multiple audio frames in the second audio signal, to obtain a third type of audio frame, and the first sampling rate is less than the second sampling rate.
[0137] In some embodiments, the device also includes: a rhythm determination unit, which is configured to determine the rhythm characteristics of the audio frame for each audio frame based on the mutation points within a target time period. The target time period is centered on the audio frame and is extended forward and backward by preset time lengths. The rhythm characteristics are used to reflect the frequency of the mutation points.
[0138] The disclosed embodiments provide a device for determining transition points, which divides audio into multiple audio segments based on chromaticity features and rhythmic features, so that audio frames within the same audio segment have similar chromaticity features and rhythmic features, thereby achieving accurate segmentation of the audio; and determines an importance parameter of each audio segment for reflecting its spectral change, and then selects candidate points of matching types as transition points based on the importance parameter of each audio segment. In this way, for audio segments with different spectral changes in the audio, different types of candidate points can be selected as transition points, thereby improving the diversity and richness of the transition points in the audio, and thus improving the transition effect of the card point video when producing the card point video based on the transition points in the audio.
[0139] Regarding the apparatus in the above embodiment, the specific manner in which each unit performs operations has been described in detail in the embodiment of the method, and will not be elaborated on here.
[0140] Figure 6The structural block diagram of the terminal 600 provided by an exemplary embodiment of the present disclosure is shown. The terminal 600 may be: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a notebook computer or a desktop computer. The terminal 600 may also be referred to by other names such as user equipment, portable terminal, laptop terminal, desktop terminal, etc.
[0141] Generally, the terminal 600 includes: a processor 601 and a memory 602.
[0142] The processor 601 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 601 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), PLA (Programmable Logic Array). The processor 601 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 601 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 601 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0143] The memory 602 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 602 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices, flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 602 is used to store at least one program code, and the at least one program code is used to be executed by the processor 601 to implement the method for determining the transition point provided in the method embodiments of the present disclosure.
[0144] In some embodiments, the terminal 600 may further optionally include: a peripheral device interface 603 and at least one peripheral device. The processor 601, the memory 602, and the peripheral device interface 603 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 603 through a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of: a radio frequency circuit 604, a display screen 605, a camera module 606, an audio circuit 607, and a power supply 608.
[0145] The peripheral device interface 603 can be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 601 and the memory 602. In some embodiments, the processor 601, the memory 602, and the peripheral device interface 603 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 601, the memory 602, and the peripheral device interface 603 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.
[0146] The radio frequency circuit 604 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 604 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 604 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 604 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 604 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: a metropolitan area network, generations of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 604 may further include a circuit related to NFC (Near Field Communication), and this disclosure does not limit this.
[0147] The display screen 605 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 605 is a touch display screen, the display screen 605 also has the ability to collect touch signals on or above the surface of the display screen 605. The touch signals can be input as control signals to the processor 601 for processing. At this time, the display screen 605 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, there can be one display screen 605, which is provided on the front panel of the terminal 600; in other embodiments, there can be at least two display screens 605, which are respectively provided on different surfaces of the terminal 600 or are in a foldable design; in still other embodiments, the display screen 605 can be a flexible display screen, which is provided on the curved surface or the folding surface of the terminal 600. Even further, the display screen 605 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 605 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0148] The camera module 606 is used to capture images or videos. Optionally, the camera module 606 includes a front camera and a rear camera. Generally, the front camera is provided on the front panel of the terminal, and the rear camera is provided on the back of the terminal. In some embodiments, there are at least two rear cameras, which are respectively any one of the main camera, the depth-of-field camera, the wide-angle camera, and the telephoto camera, so as to implement functions such as the combination of the main camera and the depth-of-field camera to achieve the background blurring function, the combination of the main camera and the wide-angle camera to achieve panoramic shooting and VR (Virtual Reality) shooting functions or other combined shooting functions. In some embodiments, the camera module 606 can also include a flash. The flash can be a single-color temperature flash or a two-color temperature flash. A two-color temperature flash refers to the combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.
[0149] The audio circuit 607 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 601 for processing, or input to the radio frequency circuit 604 to enable voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal 600. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 601 or the radio frequency circuit 604 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 607 may further include a headset jack.
[0150] The power supply 608 is used to supply power to each component in the terminal 600. The power supply 608 may be alternating current, direct current, a disposable battery or a rechargeable battery. When the power supply 608 includes a rechargeable battery, the rechargeable battery may support wired charging or wireless charging. The rechargeable battery may also be used to support fast charging technology.
[0151] Those skilled in the art can understand that Figure 6 the structure shown in
[0152] Figure 7 is a schematic structural diagram of a server provided according to an embodiment of the present disclosure. The server 700 may vary greatly due to different configurations or performances, and may include one or more processors (Central Processing Units, CPUs) 701 and one or more memories 702. Among them, the memory 702 is used to store executable program codes, and the processor 701 is configured to execute the above-mentioned executable program codes to implement the method for determining the transition point provided in each of the above method embodiments. Of course, the server may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input / output. The server may also include other components for implementing the functions of the device, which will not be elaborated here.
[0153] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory including instructions. The above instructions can be executed by the processor of the terminal to complete the method for determining the transition point. Optionally, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0154] In an exemplary embodiment, a computer program product is further provided, including a computer program which, when executed by a processor, implements the above-described method for determining transition points. In some embodiments, the computer program product involved in the embodiments of the present disclosure may be deployed to be executed on an electronic device, or on multiple electronic devices located at one location, or alternatively, on multiple electronic devices distributed at multiple locations and interconnected through a communication network. The multiple electronic devices distributed at multiple locations and interconnected through a communication network may form a blockchain system.
[0155] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include well-known knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.
[0156] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A method for determining a transition point, characterized in that: The method comprises: Segmenting the audio based on chroma features and rhythm features of a plurality of audio frames to obtain a plurality of audio segments, wherein a difference between chroma features of audio frames in the same audio segment is within a first preset range, and a difference between rhythm features of audio frames in the same audio segment is within a second preset range; determining, based on the spectral information of each of the plurality of audio clips, an importance parameter for each of the plurality of audio clips, wherein the importance parameter is used to reflect a spectral change of the audio clip, where the spectral change of the audio clip includes at least one of soothing, normal, and intense; For each of the audio segments, based on an importance parameter of the audio segment and importance parameter ranges corresponding to the multiple segment categories, the audio segment is classified into a target segment category among the multiple segment categories, wherein the importance parameter of the audio segment belongs to the importance parameter range corresponding to the target segment category; and at least one type of candidate point corresponding to the target segment category among the three candidate point types of beat points, re-beat points, and break points is selected as a target candidate point type, wherein the target candidate point type is matched with the importance parameter; A boundary time point of a target audio frame in the audio segment is determined as a transition point of the audio segment, wherein the target audio frame includes a time point of the target candidate point type.
2. The method for determining a transition point according to claim 1, wherein: The multiple segment categories include a first segment category, a second segment category, and a third segment category, the importance parameter of the audio segment of the first segment category is greater than the importance parameter of the audio segment of the second segment category, and the importance parameter of the audio segment of the second segment category is greater than the importance parameter of the audio segment of the third segment category; the selecting at least one type of candidate point corresponding to the target segment category among the three candidate point types of beat points, rebeat points, and mutation points as the target candidate point type includes at least one of the following: If the target segment category is the first segment category, taking the mutation point among the three candidate point types as the target candidate point type; If the target segment category is the second segment category, taking the beat point among the three candidate point types as the target candidate point type; If the target segment category is the third segment category, the retake point among the three candidate point types is used as the target candidate point type.
3. The method for determining a transition point according to claim 1, wherein: The audio is segmented based on the chrominance features and rhythm features of the plurality of audio frames to obtain a plurality of audio segments, including: For each audio frame, performing feature concatenation of the chrominance feature and the rhythm feature of the audio frame to obtain a segmented feature of the audio frame; performing change point detection on the segmentation features of each of the plurality of audio frames to obtain at least one changed audio frame among the plurality of audio frames, wherein the changed audio frame refers to an audio frame whose segmentation features have changed relative to a previous audio frame; The audio is segmented using the at least one changed audio frame as a demarcation point to obtain the multiple audio segments.
4. The method for determining a transition point according to claim 1, wherein: The spectrum information includes spectrum fluctuation intensity, spectrum richness, and spectrum centroid. Determining the importance parameter of each of the multiple audio segments based on the spectrum information of each of the multiple audio segments includes: For each of the audio segments, a weighted sum is performed on the spectrum fluctuation intensity, spectrum richness, and spectrum centroid of the audio segment to obtain an importance parameter of the audio segment.
5. The method for determining a transition point according to claim 1, wherein: The method further comprises: determining, from the plurality of audio frames of the audio, a first type of audio frame, a second type of audio frame, and a third type of audio frame, wherein the first type of audio frame is an audio frame including a beat point, the second type of audio frame is an audio frame including a rebeat point, the third type of audio frame is an audio frame including a break point, and signal energy of the third type of audio frame is greater than an energy threshold; An audio frame in the audio segment that belongs to a target audio frame type is used as a target audio frame, where the target audio frame type is an audio frame type that matches the target candidate point type among the audio frames of the first type, the audio frames of the second type, and the audio frames of the third type.
6. The method for determining a transition point according to claim 5, wherein: The determining, from the plurality of audio frames of the audio, an audio frame of the first type, an audio frame of the second type, and an audio frame of the third type comprises: resampling the audio based on a first sampling rate to obtain a first audio signal, and performing beat point detection and re-beat point detection on multiple audio frames in the first audio signal to obtain audio frames of the first type and audio frames of the second type; The audio is resampled based on a second sampling rate to obtain a second audio signal, and mutation point detection is performed on multiple audio frames in the second audio signal to obtain the third type of audio frames, where the first sampling rate is lower than the second sampling rate.
7. The method for determining a transition point according to claim 1, wherein: The method further comprises: For each audio frame, a rhythmic feature of the audio frame is determined based on a mutation point within a target time period. The target time period is centered on the audio frame and extended forward and backward by a preset duration. The rhythmic feature is used to reflect the frequency of the mutation point.
8. A device for determining a transition point, characterized in that: The device comprises: a segmentation unit configured to segment the audio based on the chroma features and rhythm features of the plurality of audio frames of the audio to obtain a plurality of audio segments, wherein the difference between the chroma features of the audio frames in the same audio segment is within a first preset range, and the difference between the rhythm features is within a second preset range; a parameter determination unit configured to determine, based on the spectral information of each of the plurality of audio segments, an importance parameter for each of the plurality of audio segments, wherein the importance parameter is used to reflect a spectral change of the audio segment, where the spectral change of the audio segment includes at least one of soothing, normal, and intense; The candidate point type determination unit is configured to, for each of the audio segments, classify the audio segment into a target segment category among the multiple segment categories based on an importance parameter of the audio segment and importance parameter ranges corresponding to the multiple segment categories, wherein the importance parameter of the audio segment falls within the importance parameter range corresponding to the target segment category; select at least one type of candidate point corresponding to the target segment category among the three candidate point types of beat points, re-beat points, and sudden change points as a target candidate point type, wherein the target candidate point type matches the importance parameter; The transition point determination unit is configured to determine a boundary time point of a target audio frame in the audio segment as a transition point of the audio segment, wherein the target audio frame includes a time point of the target candidate point type.
9. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the method for determining the transition point according to any one of claims 1 to 7.
10. A computer-readable storage medium, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, enables the electronic device to execute the method for determining a transition point according to any one of claims 1 to 7.
Citation Information
Patent Citations
Music transition time point detection method, equipment and medium
CN113436641A
Marking method and device of card point label, equipment and medium
CN115240618A