Composition method and device and computer readable storage medium

By performing vectorized representation and annealing algorithm processing on music samples, combined with the screening and merging of audio databases, music composition automation is realized to ensure that the created music remains consistent with the sample, and the problem of poor music creation effect in the existing technology is solved.

CN120014998AInactive Publication Date: 2025-05-16NANYANG NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411891546.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing auxiliary music creation technology cannot achieve automation of composition, resulting in inconsistent music styles and poor auditory effects.

Method used

By vectorizing the first sample to be created, the first approximate sequence is obtained by using annealing algorithm, and the second approximate sequence and the third approximate sequence are screened from the preset audio database based on the pitch time distribution characteristics and the pitch intensity distribution characteristics, and finally merge them using the annealing algorithm to obtain the target sequence.

Benefits of technology

The target sequence obtained by the creation has similar factor characteristics to the sample, so that the created music is consistent with the style of the sample, solving the problem that existing auxiliary music creation cannot achieve the automation of composition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014998A_ABST
    Figure CN120014998A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of auxiliary music composition, in particular to a composition method and device and a computer readable storage medium, and the method comprises the steps: carrying out the vectorization expression of a first sample to be subjected to composition processing, and obtaining a sample sequence; processing the sample sequence by using an annealing algorithm to obtain a first approximate sequence; based on the pitch duration distribution characteristics of the first approximate sequence and the pitch and intensity distribution characteristics of the first approximate sequence, screening is carried out from a preset audio database, a second approximate sequence and a third approximate sequence are obtained, and the audio database comprises a plurality of audio clips and phoneme characteristics corresponding to the audio clips; the phoneme features comprise pitch duration distribution features and pitch and pitch intensity distribution features; the second approximate sequence and the third approximate sequence are combined through an annealing algorithm, a target sequence is obtained, and the pitch approximation degree of the target sequence is larger than a preset first threshold value in the preset time length. The problem that existing auxiliary music creation cannot refer to automatic composition is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of music-assisted composition, and in particular, to a composition method, device and computer-readable storage medium. Background Art

[0002] Assisted music creation refers to a method that uses computer technology and artificial intelligence technology to help music creators improve their creative efficiency and enrich their creative inspiration and means. Currently, most assisted music creations only realize the automation of input and output, and do not require users to use paper, pen and musical instruments to create, which provides users with corresponding convenience. However, users still rely on music theory knowledge to create effective music works. Therefore, it is urgent to invent a solution that can automate composition. Summary of the invention

[0003] The present application provides a composition method, device and computer-readable storage medium, which can solve the problem of poor auditory effects caused by inconsistent music styles in existing composition methods. The technical solution is as follows:

[0004] In a first aspect, a composition method is provided, the method comprising:

[0005] Vectorize the first sample to be created and processed to obtain a sample sequence;

[0006] The sample sequence is processed using the annealing algorithm to obtain the first approximate sequence;

[0007] Based on the pitch duration distribution feature of the first approximate sequence and the pitch intensity distribution feature of the first approximate sequence, a second approximate sequence and a third approximate sequence are respectively screened from a preset audio database, where the audio database includes a plurality of audio clips and their corresponding phoneme features, and the phoneme features include pitch duration distribution features and pitch intensity distribution features;

[0008] The second approximate sequence and the third approximate sequence are merged by using an annealing algorithm to obtain a target sequence, wherein the target sequence includes pitch similarity greater than a preset first threshold over a preset time length.

[0009] In a second aspect, a composition device is provided, the device comprising:

[0010] A sample sequence determination module, used for vectorizing the first sample to be created and processed to obtain a sample sequence;

[0011] A sample sequence processing module, used to process the sample sequence using an annealing algorithm to obtain a first approximate sequence;

[0012] An approximate sequence screening module is used to screen from a preset audio database based on the pitch-duration distribution feature of the first approximate sequence and the pitch-intensity distribution feature of the first approximate sequence to obtain a second approximate sequence and a third approximate sequence, wherein the audio database includes a plurality of audio clips and their corresponding phoneme features, and the phoneme features include a pitch-duration distribution feature and a pitch-intensity distribution feature;

[0013] The target sequence determination module is used to combine the second approximate sequence and the third approximate sequence using an annealing algorithm to obtain a target sequence, wherein the target sequence includes a pitch similarity greater than a preset first threshold over a preset time length.

[0014] In a third aspect of an embodiment of the present application, an electronic device is disclosed. The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.

[0015] In a fourth aspect of the embodiments of the present application, a computer-readable storage medium is disclosed, which stores a computer program. When the computer program is executed by a processor, the steps of the above method are implemented.

[0016] In the embodiment of the present application, a first sample to be created is vectorized to obtain a sample sequence, and then an annealing algorithm is used to process the sample sequence to obtain a first approximate sequence. Based on the pitch-duration distribution characteristics of the first approximate sequence and the pitch-intensity distribution characteristics of the first approximate sequence, the pitch-duration distribution characteristics of the first approximate sequence are respectively screened from a preset audio database to obtain a second approximate sequence and a third approximate sequence. The audio database includes a number of audio clips and their corresponding phoneme features, and the phoneme features include pitch-duration distribution characteristics and pitch-intensity distribution characteristics. The second approximate sequence and the third approximate sequence are then merged using an annealing algorithm to obtain a target sequence, wherein the target sequence includes a pitch similarity greater than a preset first threshold over a preset time length. This method of vectorizing the sample and then screening out the approximate sequence in the audio database and merging them to obtain the target sequence makes the created target sequence have similar factor characteristics to the sample, so that the created music is consistent with the style of the sample, which solves the problem that the existing auxiliary music creation cannot be automated in composition. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in describing the embodiments of the present application.

[0018] Figure 1 A schematic diagram of the structure of the composition method provided in the embodiment of the present application;

[0019] Figure 2A schematic diagram of the structure of a composing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0020] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be interpreted as limiting the present application.

[0021] The present application embodiment provides a composition method, such as Figure 1 As shown, the method includes: step S101 to step S104.

[0022] Step S101: vectorize the first sample to be created and processed to obtain a sample sequence.

[0023] Specifically, the first sample can be a clip recorded when the creator sings a song, or a clip of an instrument played by the creator. When applied, these clips can be imported and used as the first sample directly, or these clips can be cropped and the cropped result can be used as the first sample. For example, assuming that the imported audio clip of the creator is 2 minutes and 30 seconds long, the clip between the 59th second and the 120th second can be used as the first sample through cropping.

[0024] Step S102: Process the sample sequence using an annealing algorithm to obtain a first approximate sequence.

[0025] Step S103: Based on the pitch-duration distribution characteristics of the first approximate sequence and the pitch-intensity distribution characteristics of the first approximate sequence, a second approximate sequence and a third approximate sequence are respectively screened from a preset audio database to obtain the second approximate sequence and the third approximate sequence, wherein the audio database includes a plurality of audio clips and their corresponding phoneme characteristics, and the phoneme characteristics include pitch-duration distribution characteristics and pitch-intensity distribution characteristics.

[0026] In the embodiment of the present application, the pitch-duration distribution feature is used to characterize the temporal variation trend of the pitch of the sound.

[0027] Step S104: merging the second approximate sequence and the third approximate sequence to obtain a target sequence, wherein the target sequence includes a pitch similarity greater than a preset first threshold over a preset time length.

[0028] Specifically, a preset algorithm (such as an annealing algorithm) may be used to merge the second approximate sequence and the third approximate sequence to obtain a target sequence.

[0029] In the embodiment of the present application, a first sample to be created is vectorized to obtain a sample sequence, and then an annealing algorithm is used to process the sample sequence to obtain a first approximate sequence. Based on the pitch-duration distribution characteristics of the first approximate sequence and the pitch-intensity distribution characteristics of the first approximate sequence, the pitch-duration distribution characteristics of the first approximate sequence are respectively screened from a preset audio database to obtain a second approximate sequence and a third approximate sequence. The audio database includes a number of audio clips and their corresponding phoneme features, and the phoneme features include pitch-duration distribution characteristics and pitch-intensity distribution characteristics. The second approximate sequence and the third approximate sequence are then merged using an annealing algorithm to obtain a target sequence, wherein the target sequence includes a pitch similarity greater than a preset first threshold over a preset time length. This method of vectorizing the sample and then screening out the approximate sequence in the audio database and merging them to obtain the target sequence makes the created target sequence have similar factor characteristics to the sample, so that the created music is consistent with the style of the sample, which solves the problem that the existing auxiliary music creation cannot be automated in composition.

[0030] In some embodiments, the method for obtaining the second approximate sequence includes: determining a pitch and duration association matrix according to the pitch and duration characteristics of the first approximate sequence over a preset first time length, and determining the upper and lower limits of the pitch-duration distribution corresponding to each pitch of the first approximate sequence, using the upper and lower limits of the pitch-duration distribution corresponding to each pitch of the first approximate sequence as a first filter, screening in an audio database to obtain a number of first audio clips, and merging the first audio clips after vectorizing them to obtain a second approximate sequence; the method for obtaining the third approximate sequence includes: determining a pitch and sound intensity association matrix according to the pitch and duration characteristics of the first approximate sequence over a preset second time length, and determining the upper and lower limits of the sound intensity distribution corresponding to each pitch of the first approximate sequence, using the upper and lower limits of the sound intensity distribution corresponding to each pitch of the first approximate sequence as a second filter, screening in an audio database to obtain a number of second audio clips, and merging the second audio clips after vectorizing them to obtain a third approximate sequence. Specifically, pitch is also called tone, which mainly depends on the frequency of vibration of the sound-emitting body. Generally, the higher the vibration frequency, the higher the pitch; the lower the vibration frequency, the lower the pitch. When applied, the pitch feature can be obtained by calculating the autocorrelation function or the cross-correlation function. When applied, the first approximate sequence can be Fourier transformed into a spectrogram, and the pitch and duration correlation matrix can be determined according to the change of the amplitude of the spectrum over time. Specifically, sound intensity refers to the strength characteristics of the audio, which is usually manifested as the size of the sound wave amplitude. When applied, the amplitude or energy of the first approximate sequence can be used to characterize the sound intensity. When applied, the first approximate sequence can also be feature extracted using a preset deep learning model (such as a convolutional neural network CNN, a recurrent neural network RNN, a long short-term memory network LSTM, etc.), to obtain the pitch and duration features on the first time length, and the pitch and duration features on the second time length. Feature extraction can be performed using a trained model, which can effectively retain the characteristics of the sample on which the composition is based, and provide accurate matching features for screening approximate sequences.

[0031] In some embodiments, step S104 further includes:

[0032] After randomly selecting audio segments from the second approximate sequence and the third approximate sequence, an intermediate sequence of a target length is generated;

[0033] The intermediate sequence is adjusted using an annealing algorithm based on the initialization system temperature, and the objective function of the adjusted intermediate sequence is calculated, and the adjustment of the intermediate sequence is determined according to the objective function;

[0034] Repeat the adjustment steps of the intermediate sequence until the system equilibrium is reached.

[0035] Specifically, the second approximate sequence and the third approximate sequence are generally composed of several audio clips, which can be screened by a preset random algorithm, so as to obtain an intermediate sequence after the screened audio clips are vectorized and represented, and then the intermediate sequence is adjusted by using an annealing algorithm. The objective function of the adjusted intermediate sequence is calculated and used to adjust the intermediate sequence until the system is balanced, thereby improving the similarity between the intermediate sequence used for creation and the sample, and avoiding excessive differences between the target sequence and the sample due to large differences in the selection of the intermediate sequence and the sample, thereby affecting the composition effect.

[0036] In some embodiments, the objective function is obtained based on the main pitch change characteristics obtained in the third time window and the main pitch change characteristics obtained in the fourth time window; wherein the main pitch is the pitch with the largest proportion in the first approximate sequence, and the change characteristics of the main pitch are the duration change characteristics or the intensity change characteristics of the main pitch relative to the previous time window. Specifically, the change trends of the amplitudes corresponding to the pitches in the first approximate sequence can be statistically analyzed, and the corresponding window approximate sequence can be intercepted according to the preset time window size, and the proportion of the highest pitch in the window approximate sequence can be calculated. The objective function can be quickly determined by the window size, which reduces the adjustment calculation of the objective function and reduces the computational overhead of the target sequence.

[0037] In some embodiments, the value of the objective function is obtained in the following manner: based on the third time window, the change characteristics corresponding to the main pitch in the intermediate sequence are obtained to obtain the first characteristic change sequence; based on the fourth time window, the change characteristics corresponding to the main pitch in the intermediate sequence are obtained to obtain the second characteristic change sequence; after aligning the first characteristic change sequence and the second characteristic change sequence, the correlation sequence of the first characteristic change sequence and the second characteristic change sequence is determined; the mean of the correlation sequence is obtained as the value of the objective function. Specifically, the third time window and the fourth time window can be the same or different. The first characteristic change sequence can be determined by performing Fourier transformation on the sequence in the third time window to obtain the corresponding spectrum and then filtering the change trend of the main pitch therefrom; the sequence in the fourth time window is processed accordingly in the same way to obtain the second characteristic change sequence. The value of the objective function is determined by extracting the main pitch characteristics in the sequence in the window, thereby improving the determination speed of the objective function.

[0038] In some embodiments, the upper and lower limits of the sound intensity distribution are determined as follows: determine the total number of sound intensity intervals associated with each pitch according to the sound intensity association matrix, sort the pitches according to the total number of associated sound intensity intervals, obtain the sound intensity corresponding to the pitch with the largest total number of associated sound intensity intervals as the upper limit of the sound intensity distribution, and use 10-20% of the upper limit of the sound intensity distribution as the lower limit of the sound intensity distribution; the upper and lower limits of the sound duration distribution are determined as follows: determine the total number of sound duration intervals associated with each pitch according to the pitch and duration association matrix, sort the pitches according to the total number of associated sound duration intervals, obtain the sound duration corresponding to the pitch with the largest total number of associated sound duration intervals as the upper limit of the sound duration distribution, and use 15-30% of the upper limit of the sound duration distribution as the lower limit of the sound duration distribution. Specifically, the maximum value of the sound intensity can be used as the upper limit, and the minimum value can be used as the lower limit.

[0039] In some embodiments, the process of generating the first approximate sequence includes: changing the pitch in the sample based on the pitch in the first sample, and the pitch change in the first sample has a spatial distribution consistent with the pitch difference of the first sample, and after the pitch is changed, the length and intensity of the corresponding pitch are changed, and the amplitude of the change in length or intensity is -20% to 20%, and an updated sample sequence is obtained; determining the characteristic value of the updated sample sequence according to the pitch and intensity of the updated sample sequence; accepting a new solution when the characteristic value of the updated sample sequence is less than zero, or probabilistically accepting a new solution based on the Metropolis rule, or regenerating the updated sample sequence until the new solution is accepted or the number of iterations is reached. By setting the acceptance condition of the new solution to control the iterative process, the strong global convergence of the simulated annealing algorithm is retained, and the purpose of improving the solution efficiency is achieved.

[0040] In some embodiments, the feature value of the updated sample sequence is obtained based on the third similarity sequence, the third similarity sequence includes the cosine similarity of the first sample and the segment of the first approximate sequence obtained by sliding according to the second time window and the first step length, and the feature value of the updated sample sequence is the mean of the cosine similarity in the third similarity sequence. Specifically, the second time window is controlled to slide according to the first step length to extract a segment in the first approximate sequence, and the cosine similarity of the segment and the first sample in the pitch distribution is calculated, so as to perform mean calculation to obtain the feature value of the updated sample sequence. The stability of the updated sample sequence is ensured by calculating the mean of cosine similarity, so as to avoid the problem of excessive difference between the sample sequence and the first approximate sequence due to the feature value being too large or too small.

[0041] Another embodiment of the present application provides a composition device, such as Figure 2 As shown, the device 20 includes: a sample sequence determination module 201 , a sample sequence processing module 202 , an approximate sequence screening module 203 and a target sequence determination module 204 .

[0042] A sample sequence determination module 201, configured to vectorize a first sample to be created and processed to obtain a sample sequence;

[0043] The sample sequence processing module 202 is used to process the sample sequence using an annealing algorithm to obtain a first approximate sequence;

[0044] The approximate sequence screening module 203 is used to screen the second approximate sequence and the third approximate sequence from a preset audio database based on the pitch-duration distribution feature of the first approximate sequence and the pitch-intensity distribution feature of the first approximate sequence, respectively, wherein the audio database includes a plurality of audio clips and their corresponding phoneme features, and the phoneme features include the pitch-duration distribution feature and the pitch-intensity distribution feature;

[0045] The target sequence determination module 204 is used to combine the second approximate sequence and the third approximate sequence using an annealing algorithm to obtain a target sequence, wherein the target sequence includes pitch similarity greater than a preset first threshold over a preset time length.

[0046] In the embodiment of the present application, a first sample to be created is vectorized to obtain a sample sequence, and then an annealing algorithm is used to process the sample sequence to obtain a first approximate sequence. Based on the pitch-duration distribution characteristics of the first approximate sequence and the pitch-intensity distribution characteristics of the first approximate sequence, the pitch-duration distribution characteristics of the first approximate sequence are respectively screened from a preset audio database to obtain a second approximate sequence and a third approximate sequence. The audio database includes a number of audio clips and their corresponding phoneme features, and the phoneme features include pitch-duration distribution characteristics and pitch-intensity distribution characteristics. The second approximate sequence and the third approximate sequence are then merged using an annealing algorithm to obtain a target sequence, wherein the target sequence includes a pitch similarity greater than a preset first threshold over a preset time length. This method of vectorizing the sample and then screening out the approximate sequence in the audio database and merging them to obtain the target sequence makes the created target sequence have similar factor characteristics to the sample, so that the created music is consistent with the style of the sample, which solves the problem that the existing auxiliary music creation cannot be automated in composition.

[0047] Furthermore, the sample sequence processing module includes:

[0048] A second approximate sequence determination submodule is used to determine a pitch and duration association matrix according to pitch and duration features of the first approximate sequence over a preset first time length, and determine upper and lower limits of the pitch-duration distribution corresponding to each pitch of the first approximate sequence, use the upper and lower limits of the pitch-duration distribution corresponding to each pitch of the first approximate sequence as a first filter, perform screening in an audio database to obtain a plurality of first audio clips, and merge the plurality of first audio clips after vectorizing them to obtain a second approximate sequence;

[0049] The third approximate sequence determination module is used to determine the pitch and intensity correlation matrix according to the pitch and duration characteristics of the first approximate sequence over a preset second time length, and determine the upper and lower limits of the intensity distribution corresponding to each pitch of the first approximate sequence, and use the upper and lower limits of the intensity distribution corresponding to each pitch of the first approximate sequence as a second filter to filter in the audio database to obtain a plurality of second audio clips, and merge the plurality of second audio clips after vectorizing them to obtain a third approximate sequence.

[0050] Furthermore, the target sequence determination module includes:

[0051] An intermediate sequence determination submodule, used to generate an intermediate sequence of a target length after vectorizing the audio segments randomly selected from the second approximate sequence and the third approximate sequence;

[0052] The intermediate sequence adjustment submodule is used to adjust the intermediate sequence using the annealing algorithm based on the initialization system temperature, calculate the objective function of the adjusted intermediate sequence, and select and determine the adjustment of the intermediate sequence according to the objective function, and repeat the adjustment steps of the intermediate sequence until the system equilibrium is reached.

[0053] Furthermore, the objective function in the intermediate sequence adjustment submodule is obtained based on the main pitch change characteristics obtained in the third time window and the main pitch change characteristics obtained in the fourth time window; wherein the main pitch is the pitch with the largest proportion in the first approximate sequence, and the change characteristics of the main pitch are the duration change characteristics or the intensity change characteristics of the main pitch relative to the previous time window.

[0054] Furthermore, the value of the objective function in the intermediate sequence adjustment submodule is obtained in the following manner: based on the third time window, the change characteristics corresponding to the main pitch in the intermediate sequence are obtained to obtain a first characteristic change sequence; based on the fourth time window, the change characteristics corresponding to the main pitch in the intermediate sequence are obtained to obtain a second characteristic change sequence; after aligning the first characteristic change sequence and the second characteristic change sequence, a correlation sequence between the first characteristic change sequence and the second characteristic change sequence is determined; and the mean of the correlation sequence is obtained as the value of the objective function.

[0055] Furthermore, the upper and lower limits of the sound intensity distribution in the sample sequence processing module are determined as follows: the total number of sound intensity intervals associated with each pitch is determined according to the sound intensity association matrix, the pitches are sorted according to the total number of associated sound intensity intervals, and the sound intensity corresponding to the pitch with the largest total number of associated sound intensity intervals is obtained as the upper limit of the sound intensity distribution, and 10-20% of the upper limit of the sound intensity distribution is used as the lower limit of the sound intensity distribution; the upper and lower limits of the sound length distribution are determined as follows: the total number of sound length intervals associated with each pitch is determined according to the pitch and duration association matrix, the pitches are sorted according to the total number of associated sound length intervals, and the sound length corresponding to the pitch with the largest total number of associated sound length intervals is obtained as the upper limit of the sound length distribution, and 15-30% of the upper limit of the sound length distribution is used as the lower limit of the sound length distribution.

[0056] Furthermore, the sample sequence processing module includes:

[0057] A sample sequence updating submodule is used to change the pitch in the sample based on the pitch in the first sample, and the pitch change in the first sample has a spatial distribution consistent with the pitch difference of the first sample. After the pitch is changed, the length and intensity of the corresponding pitch are changed, and the amplitude of the change in the length or intensity is -20% to 20%, to obtain an updated sample sequence;

[0058] A sample feature value determination module, used to determine the feature value of the updated sample sequence according to the pitch and sound intensity of the updated sample sequence;

[0059] The sample eigenvalue processing submodule is used to accept a new solution when the eigenvalue of the updated sample sequence is less than zero, or to accept a new solution probabilistically based on the Metropolis rule, or to regenerate the updated sample sequence until the new solution is accepted or the number of iterations is reached.

[0060] Furthermore, the characteristic value of the updated sample sequence in the sample characteristic value determination module is obtained based on the third approximation sequence; the third approximation sequence includes a fragment of the first approximation sequence obtained by sliding according to the second time window and the first step length and the cosine approximation of the first sample on the pitch distribution; the characteristic value of the updated sample sequence is the mean of the cosine approximation within the third approximation sequence.

[0061] The composition device of this embodiment can execute the composition method provided in the embodiment of the present application. The implementation principle is similar and will not be repeated here.

[0062] Another embodiment of the present application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above method when executing the computer program.

[0063] Specifically, the processor may be a CPU, a general-purpose processor, a DSP, an ASIC, an FPGA or other programmable logic device, a transistor logic device, a hardware component or any combination thereof. It may implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of this application. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0064] Specifically, the processor is connected to the memory via a bus, and the bus may include a path for transmitting information. The bus may be a PCI bus or an EISA bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc.

[0065] The memory can be a ROM or other type of static storage device that can store static information and instructions, a RAM or other type of dynamic storage device that can store information and instructions, or an EEPROM, CD-ROM or other optical disk storage, optical disk storage (including compressed optical disk, laser disk, optical disk, digital versatile disk, Blu-ray disk, etc.), magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to this.

[0066] Optionally, the memory is used to store the code of the computer program for executing the solution of the present application, and the execution is controlled by the processor. The processor is used to execute the application program code stored in the memory to implement the actions of the device provided in the above embodiment.

[0067] Another embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the above method when executed by a processor.

[0068] The device embodiments described above are only illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0069] It will be appreciated by those skilled in the art that all or some of the steps and systems in the disclosed method above may be implemented as software, firmware, hardware and appropriate combinations thereof. Some physical components or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor or a microprocessor, or may be implemented as hardware, or may be implemented as an integrated circuit, such as an application specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or a non-transitory medium) and a communication medium (or a temporary medium). As known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, disk storage or other magnetic storage devices, or any other medium that may be used to store desired information and may be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically embodies computer readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.

[0070] The above is a specific description of the preferred implementation of the present application, but the present application is not limited to the above-mentioned implementation mode. Technical personnel familiar with the field can also make various equivalent deformations or substitutions without violating the spirit of the present application. These equivalent deformations or substitutions are all included in the scope defined by the claims of the present application.

Claims

1. A composition method, characterized in that: include: Vectorize the first sample to be created and processed to obtain a sample sequence; Processing the sample sequence using an annealing algorithm to obtain a first approximate sequence; Based on the pitch-duration distribution feature of the first approximate sequence and the pitch-intensity distribution feature of the first approximate sequence, a second approximate sequence and a third approximate sequence are respectively screened from a preset audio database, wherein the audio database includes a plurality of audio clips and their corresponding phoneme features, and the phoneme features include a pitch-duration distribution feature and a pitch-intensity distribution feature; The second approximate sequence and the third approximate sequence are merged by using an annealing algorithm to obtain a target sequence, wherein the target sequence includes a pitch similarity greater than a preset first threshold over a preset time length.

2. The method according to claim 1, characterized in that The method for obtaining the second approximate sequence includes: Determine a pitch and duration association matrix according to the pitch and duration features of the first approximate sequence over a preset first time length, and determine upper and lower limits of the pitch-duration distribution corresponding to each pitch of the first approximate sequence, use the upper and lower limits of the pitch-duration distribution corresponding to each pitch of the first approximate sequence as a first filter, perform screening in the audio database to obtain a plurality of first audio clips, and merge the plurality of first audio clips after vectorizing them to obtain the second approximate sequence; The method for obtaining the third approximate sequence includes: According to the pitch and duration characteristics of the first approximate sequence over a preset second time length, a pitch and intensity correlation matrix is ​​determined, and the upper and lower limits of the intensity distribution corresponding to each pitch of the first approximate sequence are determined. The upper and lower limits of the intensity distribution corresponding to each pitch of the first approximate sequence are used as a second filter, and screening is performed in the audio database to obtain a plurality of second audio clips. The plurality of second audio clips are vectorized and then merged to obtain a third approximate sequence.

3. The method according to claim 1, characterized in that The step of combining the second approximate sequence and the third approximate sequence by using an annealing algorithm to obtain a target sequence includes: Vectorizing the audio segments randomly selected from the second approximate sequence and the third approximate sequence to generate an intermediate sequence of a target length; Based on the initialization system temperature, the intermediate sequence is adjusted using the annealing algorithm, and the objective function of the adjusted intermediate sequence is calculated. The adjustment of the intermediate sequence is determined according to the objective function, and the adjustment steps of the intermediate sequence are repeated until the system equilibrium is reached.

4. The method according to claim 3, characterized in that The objective function is obtained based on the main pitch change characteristics obtained in the third time window and the main pitch change characteristics obtained in the fourth time window; wherein the main pitch is the pitch with the largest proportion in the first approximate sequence, and the change characteristics of the main pitch are the duration change characteristics or the sound intensity change characteristics of the main pitch relative to the previous time window.

5. The method according to claim 4, characterized in that The value of the objective function is obtained as follows: Acquire the change feature corresponding to the main pitch in the intermediate sequence based on the third time window to obtain a first feature change sequence; Acquire the change feature corresponding to the main pitch in the intermediate sequence based on the fourth time window to obtain a second feature change sequence; After aligning the first feature change sequence and the second feature change sequence, determining a correlation sequence between the first feature change sequence and the second feature change sequence; Get the mean of the correlation series as the value of the objective function.

6. The method according to claim 2, characterized in that The upper and lower limits of the sound intensity distribution are determined in the following manner: the total number of sound intensity intervals associated with each pitch is determined according to the sound intensity association matrix, the pitches are sorted according to the total number of associated sound intensity intervals, the sound intensity corresponding to the pitch with the largest total number of associated sound intensity intervals is obtained as the upper limit of the sound intensity distribution, and 10-20% of the upper limit of the sound intensity distribution is used as the lower limit of the sound intensity distribution; The upper and lower limits of the sound length distribution are determined as follows: the total number of sound length intervals associated with each pitch is determined based on the pitch and duration association matrix, the pitches are sorted according to the total number of associated sound length intervals, and the sound length corresponding to the pitch with the largest total number of associated sound length intervals is obtained as the upper limit of the sound length distribution, and 15-30% of the upper limit of the sound length distribution is used as the lower limit of the sound length distribution.

7. The method according to claim 1, characterized in that The process of generating the first approximate sequence includes: The pitch in the sample is changed based on the pitch in the first sample, and the pitch change in the first sample has a spatial distribution consistent with the pitch difference of the first sample. After the pitch is changed, the length and intensity of the corresponding pitch are changed, and the amplitude of the change in the length or intensity is -20% to 20%, to obtain an updated sample sequence; Determine the feature value of the updated sample sequence according to the pitch and tone intensity of the updated sample sequence; When the eigenvalue of the updated sample sequence is less than zero, the new solution is accepted, or the new solution is probabilistically accepted based on the Metropolis rule, or the updated sample sequence is regenerated until the new solution is accepted or the number of iterations is reached.

8. The method according to claim 7, characterized in that The characteristic value of the updated sample sequence is obtained based on the third similarity sequence; The third approximation sequence includes segments of the first approximation sequence obtained by sliding according to the second time window and the first step length and cosine approximations of the first sample on the pitch distribution; The characteristic value of the updated sample sequence is the mean value of the cosine approximation in the third approximation sequence.

9. A composition device, characterized in that: include: A sample sequence determination module, used for vectorizing the first sample to be created and processed to obtain a sample sequence; A sample sequence processing module, used to process the sample sequence using an annealing algorithm to obtain a first approximate sequence; an approximate sequence screening module, configured to screen from a preset audio database based on the pitch-duration distribution feature of the first approximate sequence and the pitch-intensity distribution feature of the first approximate sequence, respectively, to obtain a second approximate sequence and a third approximate sequence, wherein the audio database includes a plurality of audio clips and their corresponding phoneme features, and the phoneme features include a pitch-duration distribution feature and a pitch-intensity distribution feature; The target sequence determination module is used to combine the second approximate sequence and the third approximate sequence using an annealing algorithm to obtain a target sequence, wherein the target sequence includes a pitch similarity greater than a preset first threshold over a preset time length.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer program is called by a processor to execute the method according to any one of claims 1 to 8.