Method and apparatus for correcting the rhythm of audio
By collecting the audio's fundamental frequency, pitch and speech recognition sequence, calculating the duration of the bars and variable speed processing, the problem that users cannot accurately follow the rhythm in karaoke software is solved, and automatic adjustment of the audio rhythm and personalized singing experience are achieved.
Patent Information
- Application Number
- CN202110977913.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-18
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-08-18
AI Technical Summary
Users cannot accurately sing with the music rhythm when using karaoke software, especially novices lack professional music knowledge when dividing the sections independently. The traditional method of correcting the audio rhythm is not applicable to young people's personalized singing and original works.
By collecting the fundamental frequency, pitch and speech recognition sequence of the audio, the theoretical and actual duration of each bar are calculated, the speed change coefficient is derived and the audio speed change processing is performed, and the audio rhythm is automatically adjusted.
It realizes accurate correction of the audio rhythm of the user's humming, and is suitable for original music singing with accompaniment or without accompaniment, improving the user's personalized singing experience.
Smart Images

Figure CN115708153B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of audio signal processing and algorithmic composition, and more particularly, to a method and apparatus for correcting the rhythm of audio. Background Art
[0002] With the development of the music industry and the Internet industry, music creation (original or secondary creation), online karaoke (i.e., Karaoke, song singing based on an accompaniment system), voice humming and other functions have become a popular field direction in the mobile Internet industry. Users' personalized needs for audio signals are becoming increasingly strong.
[0003] However, in the process of using karaoke software, users often cannot sing accurately along with the rhythm of the music, and there are problems of singing slightly earlier or slightly later. Moreover, for novice users' music creation, there are bottlenecks in music professional knowledge in the issue of independently dividing measures. Traditional methods and apparatuses for correcting the rhythm of audio need to provide a lyrics template or a singing audio template for matching in order to provide the function of rhythm correction, which is not applicable to the personalized singing of young people and original works. Therefore, there is a need for a method and system suitable for correcting pitch deviations during the singing of original music with or without accompaniment. Summary of the Invention
[0004] To solve the above problems, embodiments of the present application provide a method for correcting the rhythm of audio, which includes the steps of: determining the actual duration of each measure based on one or more of the speech recognition sequence, the original pitch sequence, and the theoretical duration of the audio, and re-dividing each measure according to the actual duration; deriving a variable speed coefficient sequence for each measure based on the theoretical duration and the actual duration of each measure; and performing variable speed processing on each re-divided measure based on the variable speed coefficient sequence of each measure to restore it to the theoretical duration of the measure.
[0005] In some embodiments, it further includes the steps of obtaining the fundamental frequency sequence of the audio, where the fundamental frequency sequence includes a plurality of time points and the fundamental frequency value of each time point; obtaining the speech recognition sequence of the audio; obtaining the original pitch sequence of the audio based on the fundamental frequency sequence; and calculating the theoretical duration of each measure of the audio based on the tempo and / or time signature.
[0006] In some embodiments, the obtaining the fundamental frequency sequence of the audio includes: obtaining a continuous fundamental frequency sequence given in the form of an array, where the time point corresponding to the fundamental frequency can be judged according to the array; and obtaining a fundamental frequency sequence given by the time points of fundamental frequency changes, which includes the fundamental frequency and the time points of fundamental frequency changes.
[0007] In some embodiments, obtaining the speech recognition sequence of the audio includes: obtaining a continuous text or coding sequence in the form of an array, where the coding sequence is used to distinguish the speech features and duration of different texts, and the time points corresponding to the texts or codings can be determined according to the array of the continuous text or coding sequence; and the text or coding sequence includes the text and the time points of text changes or the coding and the time points of coding changes.
[0008] In some embodiments, the time signature and tempo are set by the user, or if the user does not set the tempo and / or time signature, default time signature and tempo values are used.
[0009] In some embodiments, calculating the duration of each measure based on the tempo and time signature includes: calculating the theoretical duration of each measure based on the time signature and tempo; or directly giving the theoretical duration of each measure; or according to the total duration of the audio, giving the total number of measure divisions of the audio, and determining the theoretical duration of each measure based on the total duration of the audio and the total number of measure divisions of the audio.
[0010] In some embodiments, the speech recognition sequence includes each recognized word and the pronunciation duration of each word; the original pitch sequence of the audio includes the original pitch of the audio and the duration of each original pitch; determining the actual duration of each measure of the audio based on the speech recognition sequence, original pitch sequence, and the theoretical duration of each measure of the audio includes: if no pitch information and no speech sequence information are detected at the end point of the theoretical duration, using the theoretical duration as the actual duration of the measure; determining and re-dividing the actual duration of the measure of the audio based on the two reference bases of the speech recognition sequence and the original pitch sequence of the audio; or determining and re-dividing the actual duration of the measure of the audio based on only one of the speech recognition sequence or the original pitch sequence of the audio; or finding the time point closest to the end point of the theoretical duration of the measure among the time points where the speech recognition sequence changes and the time points where the pitch recognition result changes, and determining the closest time point as the actual end time point of the measure; or assigning different weights to the speech recognition sequence and the pitch recognition result, finding the time point closest to the end point of the theoretical duration of the measure among the time points where the speech recognition sequence changes and the time points where the pitch recognition result changes, and determining the time point with the highest weight as the actual end time point of the measure through weight calculation.
[0011] In some embodiments, deriving the tempo change sequence for each measure based on the theoretical duration and the actual duration of each measure includes: for a tempo change coefficient sequence given in the form of positive and negative offsets, adjusting the tempo of the audio for each measure with the positive and negative offsets; or, for a tempo change coefficient sequence given as a tempo change ratio, using the tempo change ratio as the basis for tempo adjustment for each measure.
[0012] In some embodiments, tempo processing is performed on the audio for each measure having an actual duration according to each measure tempo change coefficient sequence.
[0013] The present application also discloses an apparatus for correcting the rhythm of audio, which includes at least one processor and at least one memory including computer program code, wherein the at least one memory and the computer program code are configured to, together with the at least one processor, cause the apparatus to at least execute the method for correcting the rhythm of audio in any one of the above.
[0014] The present application determines the actual duration of each measure by collecting the fundamental frequency and pitch contained in the audio itself and deriving the theoretical duration of each measure, and accordingly re - divides the duration of each measure, then calculates the tempo change coefficient and restores the audio according to the tempo change coefficient, so as to achieve rhythm adjustment of the audio. Description of the Drawings
[0015] Figure 1 is a schematic diagram of a method for correcting the rhythm of audio according to an embodiment of the present application; Detailed Description of the Embodiments
[0016] The following will describe in detail the specific embodiments of the present application with reference to the accompanying drawings.
[0017] In order to clearly express the scope of the present application and avoid ambiguity, the following are the definitions of common terms in the present application:
[0018] Time signature: The time signature is a common symbol in musical scores, marked in the form of a fraction. Generally, there is a time signature at the beginning of the musical score, and if the rhythm changes in the middle, the changed time signature will be marked. The time signature is like a fraction, such as 2 / 4, 3 / 4, etc. The denominator represents the value of the beat, that is, which note value is used as one beat. For example, 2 / 4 means using a quarter note as one beat, and there are two beats in each measure. The numerator represents how many beats there are in each measure. As mentioned before, in 2 / 4 time, a quarter note is one beat and there are two beats in a measure; in 3 / 4 time, a quarter note is one beat and there are three beats in each measure... The reading method is to read the denominator first and then the numerator. For example, 2 / 4 is called two - four time, 3 / 4 is called three - four time, and 6 / 8 is called six - eight time.
[0019] Tempo: The speed of the beats, which is usually in the unit of beats per minute (BPM) in this application. It refers to the number of sound beats emitted within a one-minute time period, and the unit of this number is BPM.
[0020] Measure: In the progress of music, its strong beats and weak beats always regularly appear in a cycle. The part from one strong beat to the next strong beat is called a measure. In sheet music, measures are separated by short vertical lines (measure lines). In this application, a measure specifically refers to the duration of a complete time signature response.
[0021] Pitch: The standard height of a sound. The internationally common standard height (the first international height) is the a sound with a mechanical wave of 440 Hz and a wavelength of 78 cm, that is, the a in the first leger line above middle C is the "standard pitch". The corresponding pitches of the equal temperament are extended from this.
[0022] Fundamental frequency: The lowest and usually the strongest frequency in a complex wave, which is the lowest natural frequency of the vibration system. It determines the pitch at the current time in audio.
[0023] It is easily understandable that, as generally described and depicted in the accompanying drawings of this article, the components of certain exemplary embodiments can be arranged and designed in various different configurations. Therefore, the following detailed description of some example embodiments of systems, methods, devices, and computer program products related to the interactive multimedia structure is not intended to limit the scope of certain embodiments, but is representative of the selected example embodiments.
[0024] The features, structures, or characteristics of the example embodiments described throughout the specification can be combined in any suitable manner in one or more example embodiments. For example, throughout the specification, the use of phrases such as "certain embodiments", "some embodiments", or other similar languages refers to the fact that the specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment. Therefore, throughout the specification, the appearance of phrases such as "in certain embodiments", "in some embodiments", "in other embodiments", or other similar languages does not necessarily all refer to the same set of embodiments, and the described features, structures can be combined in any suitable manner in one or more example embodiments. Additionally, the phrase "a set" refers to a set including one or more of the recited members. Therefore, the phrases "a set", "one or more", and "at least one" or equivalent terms can be used interchangeably. Additionally, unless otherwise explicitly stated, "or" is intended to mean "and / or".
[0025] Additionally, if needed, the different functions or operations discussed below can be performed in different orders and / or simultaneously with each other. Furthermore, if desired, one or more of the described functions or operations can be optional or can be combined. Thus, the following description should be considered only as an illustration of the principles and teachings of certain exemplary embodiments and not as a limitation thereof.
[0026] In scenarios where the user is humming an original song, adapting an existing song, etc., which contain human voices, the present application provides a method and apparatus for correcting the rhythm of the human voice in the song.
[0027] Figure 1 It is a flowchart of a method for correcting the rhythm of an audio according to an embodiment of the present application. This embodiment relates to a method for correcting the rhythm of an audio, especially a method for correcting the rhythm of a human voice audio, which includes collecting the singing voice of the user humming a song at a sampling rate to obtain an audio file in, for example, WAV format, step S110, and simultaneously or sequentially collecting the tempo and time signature defined by the user, step S120; after obtaining the audio file, audio processing can be performed on the audio file to obtain the fundamental frequency sequence of the audio file, step S210, and the pitch sequence, step S220; simultaneously or sequentially performing speech recognition on the audio file to obtain the speech recognition sequence of the audio file, step S300, and simultaneously or sequentially, the tempo and time signature defined by the user collected can be obtained, and the theoretical duration of each measure can be calculated according to the tempo and time signature, step S310; thereafter, based on the speech recognition sequence, the original pitch sequence of the audio, and the theoretical duration of each measure, the actual duration of each measure is calculated, and the audio is re-divided into multiple measures according to the actual duration, step S400; the tempo change coefficient is calculated for the duration of each measure divided from the audio according to the time offset amount to obtain a tempo change coefficient sequence for each measure, step S500, and the audio of each measure is subjected to tempo change processing according to the tempo change coefficient sequence for each measure to obtain the audio with the corrected rhythm, step S600.
[0028] Among them, the obtained fundamental frequency sequence of the audio includes multiple time points and the fundamental frequency values at each of the time points; among them, the obtained speech recognition sequence of the audio includes each recognized character and the pronunciation duration of each character; among them, the original pitch sequence of the audio can be obtained based on the fundamental frequency sequence, and the original pitch sequence of the audio includes the original pitch of the audio and the duration of each pitch; among them, the actual duration of each measure can be calculated based on the speech recognition sequence, the original pitch sequence, and the theoretical duration of each measure of the audio, and the measure of the audio can be re-divided according to the actual duration; among them, the variable speed coefficient of each measure is calculated according to the offset of the actual duration of each measure of the audio relative to the theoretical duration of each measure calculated based on the tempo value and the time signature value, and a variable speed coefficient sequence of each measure is obtained; subsequently, variable speed processing is performed on each measure of the audio according to the variable speed coefficient sequence of each measure.
[0029] In the method of the present application, a function can be provided to enable the user to set the tempo value and the time signature value, for example, select from existing options. When the user does not make a setting, default values of the tempo value and the time signature value can be provided. Or this function can be not provided and the tempo value and the time signature value of each audio can be directly set to default values.
[0030] If no fundamental frequency information and no speech recognition information are detected at the time point where the theoretical end of the measure is, it is determined that the theoretical duration of the measure is its actual duration;
[0031] Either the speech recognition result or the original pitch sequence of the audio can be used alone as a judgment basis to determine the actual duration of the measure of the audio. For example, if the characters in the speech recognition result change or the original pitch changes, such as detecting the end or start of the duration of a character, or for example, detecting the end of one original pitch or the start of another original pitch, it is determined that the theoretical duration of this measure is the actual duration of this measure.
[0032] The actual duration of the measure of the audio can be determined based on both the speech recognition result and the original pitch sequence of the audio. For example, by finding the time point where the speech recognition result changes, such as the change of characters, and the time point where the pitch recognition result changes, such as the change of pitch, the time point closest to the end of the theoretical duration of the measure among the two can be determined as the actual end time point of the measure; or
[0033] Assign different weights to the speech recognition result and the pitch recognition result, find the time point closest to the end of the theoretical duration of the measure among the time points where the speech recognition result changes and the time points where the pitch recognition result changes, and through the calculation of the weights, determine the actual end time point of the measure as the time point with the highest weight.
[0034] In addition, derive the sequence of tempo change coefficients for each measure based on the theoretical duration and the actual duration of each measure. The derivation of this sequence of tempo change coefficients can include deriving the sequence of tempo change coefficients based on the offset, and / or deriving the sequence of tempo change coefficients based on the tempo change ratio. For the sequence of tempo change coefficients given in the form of positive and negative offsets, perform measure tempo adjustment on the audio with the positive and negative offsets; or, for the sequence of tempo change coefficients given as the tempo change ratio, use the tempo change ratio as the basis for measure tempo adjustment.
[0035] This application determines the actual duration of the measure by collecting the fundamental frequency and pitch contained in the audio itself and deriving the theoretical duration of each measure, re-divides the duration of the measure accordingly, then calculates the tempo change coefficient and restores the audio according to the tempo change coefficient, so as to realize the rhythm adjustment of the audio.
[0036] The applicant designed a software that provides a new way of playing the KTV mode. Its operation mode is as follows: the user selects the bpm and time signature, and then starts humming and recording. When the recording ends, the system enters the following rhythm automatic correction steps:
[0037] Step 1: Obtain the audio file recorded by the user;
[0038] Step 2: Use the PYIN algorithm to obtain the fundamental frequency sequence of the audio file, and use the librosa module to convert the audio fundamental frequency into a pitch sequence;
[0039] Step 3: Call the third-party API interface for speech recognition to obtain the corresponding relationship between the speech recognition result and the time;
[0040] Step 4: Assume that the bpm value selected by the user is 120 and the time signature is 4 / 4, then calculate that the theoretical measure duration is 2s for each measure, that is, the duration of each theoretical measure is 2s;
[0041] Step 5. Based on the speech recognition result of the audio, the original pitch sequence, and the theoretical duration of each measure, re-partition the audio, where the re-partitioning rule is as follows: 1) If, when the theoretical duration of a measure ends, the text recognized in the speech recognition result exactly changes, for example, the duration of a word ends or starts, or the original pitch changes, such as by 0.5 s, then determine the end point of the theoretical duration as the end point of the actual duration, and re-partition it into a new measure with this time point as the boundary, and jump to the next measure partitioning; 2) If, when the theoretical duration of a measure ends, the speech recognition result does not change and the original pitch does not change, then find the time point where the speech recognition result changes and the time point where the pitch recognition result changes within this measure, and select the time point closest to the end time point of the theoretical duration of this measure as the actual end time point of this measure, and re-partition this measure with this actual end time point as the boundary, and jump to the determination of the actual duration of the next measure and the re-partitioning of the measure.
[0042] Calculate the time offset of each measure after re-partitioning; the calculation method of the time offset is: offset = the duration of the currently adjusted measure in the audio partitioning - the theoretical duration of each measure calculated based on the tempo; if the cumulative pitch conversion point closest to 2 s is calculated as 1.9 s and the speech change point is 2.2 s, then select the pitch conversion point of 1.9 s as the partitioning standard for this measure, and calculate the offset as -0.1 s; if the cumulative pitch conversion point closest to 2 s is calculated as 2.2 s and the speech change point is 2.1 s, then select the speech conversion point of 2.1 s as the re-partitioning time point for this measure, and calculate the offset as +0.1 s; after calculating all measures, form the duration sequence of the re-partitioned measures and the offset duration sequence, and by analogy, obtain the tempo change coefficient sequence including all measures in the whole audio or a part of it.
[0043] Step 6. Use the WSOLA algorithm to perform tempo change processing on each measure of the audio according to the tempo change coefficient sequence of each measure. For example, if the re-partitioned durations of an audio including five measures are [1.9 s, 2.3 s, 1.8 s, 2.1 s, 2.0 s], then use the WSOLA algorithm to perform tempo change adjustment on these measures so that each measure is 2 s.
[0044] It should be understood that the above example is only a possible embodiment of the method for implementing the present application. Some steps can be adjusted as needed.
[0045] For example, the ratio of the theoretical duration of a measure to the actual duration of the re-partitioned measure can also be used as the tempo change coefficient. And perform tempo change processing on the audio based on this tempo change coefficient.
[0046] For another example, the above rhythm adjustment can be performed on the collected audio as a whole after the user's humming recording is completed, or the audio can be directly collected in real time during the user's humming and adjusted according to the above steps.
[0047] For another example, a continuous fundamental frequency sequence given in the form of an array can be obtained, and the time points corresponding to the fundamental frequencies can be determined according to the array; the fundamental frequency sequence given by the time points of the fundamental frequency change includes the fundamental frequency and the time points of the fundamental frequency change.
[0048] For another example, a continuous text or coding sequence can be given in the form of an array, where the function of the coding sequence is to distinguish the speech features and duration of different texts. According to the array, the time points corresponding to the texts or codings can be determined; the text or coding sequence given by the time points of the text or coding change includes the text and the time points of the text change or the coding and the time points of the coding change.
[0049] In some example embodiments, the functions of any method, process, signaling diagram, algorithm, or flowchart described herein can be implemented by software and / or computer program code or code portions stored in a memory or other computer-readable or tangible medium and executed by a processor.
[0050] In some example embodiments, a device may be included or associated with at least one software application, module, unit, or entity configured to perform arithmetic operations, or as its program or portion (including added or updated software routines), executed by at least one operating processor. A program, also referred to as a program product or a computer program, includes software routines, applets, and macros, can be stored in any device-readable data storage medium, and can include program instructions for performing specific tasks.
[0051] A sequence is a unit of data structure, which may include strings, lists, tuples, etc.
[0052] A computer program product may include one or more computer-executable components configured to perform some example embodiments when the program runs. The one or more computer-executable components may be at least one software code or code portion. Changes and configurations for implementing the functions of the example embodiments can be executed as routines, which can be implemented as added or updated software routines. In one example, the software routines can be downloaded to the device.
[0053] As an example, software or computer program code or a part of the code can be in source code form, object code form or some intermediate form, and can be stored in some carrier, distribution medium or computer-readable medium, which can be any entity or device capable of carrying the program. For example, such a carrier can include a recording medium, computer memory, read-only memory, optoelectronic and / or electrical carrier signals, telecommunication signals and / or software distribution packages. Depending on the required processing power, the computer program can be executed in a single electronic digital computer or can be distributed among multiple computers. The computer-readable medium or computer-readable storage medium can be a non-transitory medium.
[0054] In other example embodiments, the functionality can be performed by circuitry, such as by using an application specific integrated circuit (ASIC), programmable gate array (PGA), field programmable gate array (FPGA) or any other combination of hardware and software. In yet another example embodiment, the functionality can be implemented as a signal, such as a non-tangible means carried by an electromagnetic signal that can be downloaded from the Internet or other network.
[0055] According to an example embodiment, a device such as a node, device or responsive component can be configured as a circuit, computer or microprocessor (such as a single-chip computer element) or a chip set, which can at least include a memory for providing storage capacity for arithmetic operations and / or an arithmetic processor for performing arithmetic operations.
[0056] The example embodiments described herein apply equally to both singular and plural implementations, regardless of whether the language used to describe certain embodiments is in the singular or plural form. For example, an embodiment describing the operation of a single computing device equally applies to an embodiment including multiple instances of the computing device, and vice versa.
[0057] Those of ordinary skill in the art will readily understand that the example embodiments described above can be implemented with operations in a different order and / or with hardware elements in a configuration different from the disclosed configuration. Thus, although some embodiments have been described based on these example embodiments, it will be apparent to those skilled in the art that certain modifications, variations and alternative constructs will be apparent while still within the spirit and scope of the example embodiments.
Claims
1. A method for correcting the rhythm of audio, characterized in that: Including steps Determine the actual duration of each measure based on the speech recognition sequence and / or the original pitch sequence of the audio, and re-divide each measure according to the actual duration; wherein, the speech recognition sequence includes each recognized word and the pronunciation duration of each word; the original pitch sequence includes the original pitch of the audio and the duration of each original pitch; Derive the variable speed coefficient sequence of each measure according to the theoretical duration and the actual duration of each measure; Perform variable speed processing on each re-divided measure based on the variable speed coefficient sequence of each measure to restore it to the theoretical duration of the measure.
2. The method for correcting the rhythm of an audio according to claim 1, characterized in that: It also includes determining the actual duration of each measure in combination with the theoretical duration.
3. The method for correcting the rhythm of an audio according to claim 2, characterized in that: It also includes obtaining the fundamental frequency sequence of the audio, the fundamental frequency sequence including a plurality of time points and the fundamental frequency value of each time point; obtaining the speech recognition sequence of the audio; obtaining the original pitch sequence of the audio based on the fundamental frequency sequence; And calculating the theoretical duration of each measure of the audio based on the tempo and the time signature value.
4. The method for correcting the rhythm of the audio according to claim 3, characterized in that: The obtaining the fundamental frequency sequence of the audio includes: Obtaining a continuous fundamental frequency sequence given in the form of an array, wherein the time point corresponding to the fundamental frequency can be judged according to the array; and Obtaining a fundamental frequency sequence given by the time points of the fundamental frequency change, which includes the fundamental frequency and the time points of the fundamental frequency change.
5. The method for correcting the rhythm of an audio according to claim 3, characterized in that: The obtaining the speech recognition sequence of the audio includes: obtaining a continuous text or coding sequence in the form of an array, the coding sequence being used to distinguish the speech features and the duration of different texts, and the time point corresponding to the text or coding can be judged according to the array of the continuous text or coding sequence; and the text or coding sequence includes the text and the time points of the text change or the coding and the time points of the coding change.
6. The method for correcting the rhythm of an audio according to claim 3, characterized in that: The time signature and the tempo value are set by the user, or if the user does not set the tempo and / or the time signature, the default time signature and / or tempo value is used.
7. The method for correcting the rhythm of an audio according to claim 3, characterized in that: The calculating the duration of each measure based on the tempo and / or the time signature value includes: calculating the theoretical duration of each measure based on the time signature and the tempo; or directly giving the theoretical duration of each measure; or according to the total duration of the audio, giving the total number of measure divisions of the audio, and determining the theoretical duration of each measure based on the total duration of the audio and the total number of measure divisions of the audio.
8. The method for correcting the rhythm of an audio according to claim 2, characterized in that: The determining the actual duration of each measure of the audio based on the speech recognition sequence, the original pitch sequence, and the theoretical duration of each measure includes: If no pitch information and no speech sequence information are detected at the end point of the theoretical duration, then use the theoretical duration as the actual duration of the measure; and Determine the actual duration of the measure of the audio based on the two reference bases of the speech recognition sequence and the original pitch sequence of the audio and re-divide the measure; or Determine the actual duration of the measure of the audio based on a single reference basis of the speech recognition sequence or the original pitch sequence of the audio and re-divide the measure; or Find the time point closest to the end point of the theoretical duration of the measure among the time points where the speech recognition sequence changes and the time points where the pitch recognition result changes, and determine this closest time point as the actual end time point of the measure; or Assign different weights to the speech recognition sequence and the pitch recognition result, find the time point closest to the end point of the theoretical duration of the measure among the time points where the speech recognition sequence changes and the time points where the pitch recognition result changes, and through weight calculation, determine the time point with the highest weight as the actual end time point of the measure.
9. The method for correcting the rhythm of an audio according to claim 2, characterized in that: Derive the speed change sequence of each measure based on the theoretical duration and the actual duration of each measure, including: for a speed change coefficient sequence given in the form of positive and negative offsets, perform measure speed change adjustment on the audio with the positive and negative offsets; or, for a speed change coefficient sequence given as a speed change ratio, use the speed change ratio as the basis for measure speed change adjustment.
10. The method for correcting the rhythm of an audio according to claim 2, characterized in that: Perform speed change processing on each measure audio with an actual duration according to each measure speed change coefficient sequence.
11. Apparatus for correcting the rhythm of audio, characterized in that: Comprising at least one processor and at least one memory including computer program code, wherein the at least one memory and the computer program code are configured to, together with the at least one processor, cause the device to at least execute the method for correcting the rhythm of audio according to any one of claims 1 to 10.
Citation Information
Patent Citations
Device for adjusting rhythm of music based on sensor in which music changes by following physiological changes of a human body
TW202141998A