Training method, device and computer program product for background music generator
By training the background music generator through the adversarial generative network, suitable background music is automatically generated for long audio, solving the time-consuming problem of manual music composition and achieving efficient and realistic music effects.
Patent Information
- Application Number
- CN202211025605.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-25
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2042-08-25
Smart Images

Figure CN115376475B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of audio processing, and in particular to a training method, computer device, and computer program product for a background music generator. Background Art
[0002] Long audio clips refer to audio formats such as storytelling, crosstalk, and radio programs. Some of these videos primarily feature speaking, supplemented by singing, and have become a crucial addition to popular pastimes. While some long audio clips are primarily speaking, background music is essential. Appropriate background music can help users understand the content, connect with the speaker's artistic conception, and provide greater scope for imagination. Therefore, creating music for long audio clips has always been a crucial element of creative production.
[0003] However, the current long audio soundtrack requires manual participation. Manually selecting background music for long audio soundtracks is time-consuming and the efficiency of long audio soundtracks is low. Summary of the Invention
[0004] Based on this, it is necessary to provide a training method, computer equipment and computer program product for a background music generator to address the above technical problems.
[0005] The present application provides a training method for a background music generator, the method comprising:
[0006] Acquire at least one background music set, wherein each background music set has a corresponding long audio category;
[0007] Adapting at least one long audio for each piece of background music in the background music set according to the long audio category corresponding to the background music set;
[0008] Each piece of background music in the at least one background music set and the long audio adapted for the background music are used as training samples; the training samples are used to train a generative adversarial network, and the generator in the trained generative adversarial network is used as a background music generator.
[0009] The present application provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the following steps:
[0010] Acquire at least one background music set, wherein each background music set has a corresponding long audio category;
[0011] Adapting at least one long audio for each piece of background music in the background music set according to the long audio category corresponding to the background music set;
[0012] Each piece of background music in the at least one background music set and the long audio adapted for the background music are used as training samples; the training samples are used to train a generative adversarial network, and the generator in the trained generative adversarial network is used as a background music generator.
[0013] The present application provides a computer program product having a computer program stored thereon, wherein the computer program is executed by a processor to perform the following steps:
[0014] Acquire at least one background music set, wherein each background music set has a corresponding long audio category;
[0015] Adapting at least one long audio for each piece of background music in the background music set according to the long audio category corresponding to the background music set;
[0016] Each piece of background music in the at least one background music set and the long audio adapted for the background music are used as training samples; the training samples are used to train a generative adversarial network, and the generator in the trained generative adversarial network is used as a background music generator.
[0017] In the above-mentioned training method, computer device and computer program product of the background music generator, the generator in the trained adversarial generative network is used as a background music generator, and the background music generator is used to compose music for the long audio to be composed. Without human intervention, suitable background music is automatically generated for the long audio to be composed, thereby realizing automatic composition of long audio and improving the efficiency of long audio composition. Moreover, the training of the background music generator is based on the adversarial generative network, so the background music generated by the background music generator can be more realistic and closer to real background music. In addition, according to the long audio category corresponding to the background music set, at least one long audio is adapted for each background music in the background music set, and each background music in the at least one background music set and the long audio adapted to the background music are used as training samples. Through such training samples, the adversarial generative network is trained, so that the background music generator can generate suitable background music for the long audio to be composed while taking into account the category of the long audio to be composed. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 1 is a flow chart of a training method for a background music generator according to an embodiment;
[0019] Figure 2 Schematic diagram of the composition of a training set in one embodiment;
[0020] Figure 3 A structural diagram of a generative adversarial network in one embodiment;
[0021] Figure 4 A schematic diagram of iteratively processing an audio spectrum in one embodiment;
[0022] Figure 5 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0023] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0024] References to "embodiments" in this application mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described in this application may be combined with other embodiments.
[0025] This application provides a training method for a background music generator, such as Figure 1 As shown, the method can be executed by a computer device, and specifically includes the following steps:
[0026] Step S101: Acquire at least one background music set.
[0027] Among them, each background music set has its own corresponding long audio category. The long audio category can be obtained by emotional classification. In this case, the long audio category can include sad, healing, radio, romantic, exciting, cheerful, etc. In addition, the long audio category can also be obtained by music genre classification. In this case, the long audio category can include: rock, heavy metal, folk songs, jazz, etc. In this embodiment, after obtaining a plurality of background music, these background music are divided according to the long audio category to obtain at least one background music set. For the background music set formed, the background music in the same background music set has the same long audio category. Among them, the number of background music in each background music set can be equal.
[0028] Step S102: Adapt at least one long audio to each piece of background music in the background music set according to the long audio category corresponding to the background music set.
[0029] After obtaining at least one background music set through division, k corresponding long audio files are determined for the background music in each background music set, where k is an integer greater than or equal to 1. For example, if a background music set includes background music a and b, and the long audio file category corresponding to the background music set is sad, then at least one long audio file in the sad category is determined to be compatible with background music a, and at least one long audio file in the sad category is determined to be compatible with background music b. In this manner, k corresponding long audio files can be determined for each background music piece in each background music set.
[0030] In some scenarios, if the same long audio is adapted to background music corresponding to different long audio categories, for example, the same long audio is adapted to background music a of the sad category and background music m of the cheerful category, then the generator of the adversarial generative network will find it difficult to learn: to generate suitable background music for the long audio while taking into account the long audio category, and the learning difficulty increases.
[0031] Based on this, in order to reduce the learning difficulty of the generator, in some embodiments of the present application, the same long audio is only adapted to the background music corresponding to the same long audio category.
[0032] Step S103: Use each piece of background music in at least one background music set and the long audio adapted to the background music as training samples; use the training samples to train a generative adversarial network, and use the generator in the trained generative adversarial network as a background music generator.
[0033] The adversarial generative network may include a generator and a discriminator. In this embodiment, the competitive relationship between the generator and the discriminator of the adversarial generative network is utilized to make the background music generated by the generator more realistic and closer to real background music.
[0034] Specifically, after adapting at least one long audio to each piece of background music in the background music set, for each piece of background music and its adapted long audio, the background music can be used as the discrimination criterion of the discriminator of the generative adversarial network, and the corresponding adapted long audio can be used as the data input to the generator of the generative adversarial network. Based on the background music as the discrimination criterion and the long audio as the input generator data, at least one training sample is obtained.
[0035] In the process of training the adversarial generative network using at least one training sample, the long audio in the training sample is input into the generator so that the generator generates background music based on the input long audio; the background music generated by the generator is input into the discriminator so that the discriminator can distinguish whether the input background music is the background music generated by the generator or the background music in the training sample.
[0036] The adversarial generative network is trained with the goal that the background music generated by the generator cannot be distinguished by the discriminator. If the background music input to the discriminator is generated by the generator, but the discriminator considers it to be a training sample, then the background music generated by the generator is considered to be indistinguishable by the discriminator. After the training is completed, the generator in the trained adversarial generative network is used as the background music generator. Then, the long audio to be matched with music can be input into the background music generator. The background music generator generates background music based on the long audio, and uses the background music to match the long audio.
[0037] In the above-mentioned training method of the background music generator, the generator in the trained adversarial generative network is used as the background music generator, and the background music generator is used to compose music for the long audio to be composed. Without human intervention, suitable background music is automatically generated for the long audio to be composed, thereby realizing automatic composition of long audio and improving the efficiency of long audio composition. Moreover, the training of the background music generator is based on the adversarial generative network, so the background music generated by the background music generator can be more realistic and closer to real background music. In addition, according to the long audio category corresponding to the background music set, at least one long audio is adapted for each background music in the background music set, and each background music in at least one background music set and the long audio adapted to the background music are used as training samples. Through such training samples, the adversarial generative network is trained, so that the background music generator can generate suitable background music for the long audio to be composed while taking into account the category of the long audio to be composed.
[0038] In order to form a background music collection corresponding to a long audio category, each background music needs to be divided according to the long audio category, and background music corresponding to the same long audio category is divided together to form a background music collection corresponding to the long audio category.
[0039] When determining which long audio category each piece of background music corresponds to, the matching degree between the same piece of background music and the long audios belonging to different long audio categories can be compared, and the long audio category to which the long audio with the highest matching degree belongs can be used as the long audio category corresponding to the background music; it should be noted that the long audio used to determine which long audio category the background music corresponds to can be the same as or different from the long audio mentioned in step S102.
[0040] Therefore, when a computer device obtains at least one background music set, it can specifically perform the following steps: obtain multiple background music and multiple long audios; different long audios belong to different long audio categories; among the multiple long audios, determine the most matching long audio for each background music; classify the background music belonging to the same long audio category of the most matching long audio into the same set, and obtain at least one background music set.
[0041] Furthermore, in order to determine the most matching long audio for each background music, the distance between feature vectors can be used for judgment. According to the size of the distance between the feature vectors, the most matching long audio for the background music can be determined, and then a more accurate long audio category can be determined for the background music, thereby improving the classification accuracy.
[0042] Specifically, the computer device can extract the feature vectors of each long audio and the feature vectors of each background music; obtain the distance between the feature vector of each background music and the feature vector of each long audio, and take the long audio with the smallest distance as the long audio that best matches the background music.
[0043] In this embodiment, the feature vectors of the long audio and the feature vectors of the background music can be extracted through the embedding layer. In this case, the extracted feature vectors can be called embedding vectors. For example, multiple long audios can be obtained, and these long audios can correspond to the sad category, the healing category, the radio category, the romantic category, the exciting category, and the cheerful category, respectively. The embedding layer is used to extract the embedding vector of each long audio; then, the embedding layer is used to extract the embedding vector of the background music, and then the distance between the embedding vector of the background music and the embedding vector of each long audio is calculated; among the multiple distances obtained, the long audio corresponding to the smallest distance is determined, and the long audio is used as the long audio that best matches the background music.
[0044] As described in step S102, the computer device can adapt at least one long audio to each background music; in order to determine a more suitable long audio for each background music, the distance between the feature vectors can be used for judgment, and based on the size of the distance between the feature vectors, a more suitable long audio can be determined for the background music, thereby improving the adaptability between the background music and the long audio.
[0045] Specifically, the computer device can obtain the feature vector of each background music in the background music set, as well as the feature vectors of multiple long audios under the long audio category corresponding to the background music set; obtain the distance between the feature vector of each background music and the feature vector of each long audio, and take the long audio corresponding to the distance less than the threshold as the long audio adapted to the background music.
[0046] Among them, the feature vector of the long audio and the feature vector of the background music can be extracted through the embedding layer. In this case, the extracted feature vector can be called an embedding vector.
[0047] Take background music a as an example to adapt to long audio:
[0048] Reference Figure 2The long audio category corresponding to the background music set to which background music a belongs is the sad category. Therefore, the computer device can obtain multiple long audios under the sad category, obtain the embedding vector of background music a and the embedding vectors of each of these long audios, calculate the distance between the embedding vector of background music a and the embedding vector of each long audio, and determine the long audio with a distance less than a threshold as the long audio adapted for background music a. The number of long audios with a distance less than the threshold can be one or more, and accordingly, the number of long audios adapted for background music a can be one or more.
[0049] If the number of long audios adapted to each background music is k, then each training sample includes: a background music and its k adapted long audios; if the adversarial generative network only includes one generator, then the k long audios of the training sample need to be input into the generator in sequence, and the generator generates the corresponding background music in sequence. It is impossible to generate the background music corresponding to the k long audios in parallel, and the training efficiency is low.
[0050] Based on this, in some embodiments provided by this application, the adversarial generative network may include k generators, the parameters of these generators are shared, and the k long audios of the training samples can be input into different generators, and then the background music corresponding to the k long audios can be generated in parallel, thereby improving training efficiency.
[0051] When the adversarial generation network includes k generators, its structure diagram is as follows Figure 3 As shown. Figure 3 , G represents the generator. The adversarial generative network includes k generators, and each generator is distinguished by a subscript, for example, G k-1 represents the k-1th generator. z is the long audio in the training sample, Pz is the feature vector of the long audio; G k-1 (z;θ g k-1 ) in θ g k-1 Represents the parameters of the k-1th generator. When k generators share parameters, G k-1 (z;θ g k-1 ) in θ g k-1 With G k (z;θ g k ) in θ g k Same; x represents the background music of the training sample, and Pd is the feature vector of the background music.
[0052] Refer again Figure 3, D represents the discriminator. The feature vectors of the background music generated by the k generators and the feature vectors of the background music of the training samples are input into the discriminator together; if the background music of the training samples is called "true" background music, and the background music generated by the generator is called "false" background music, then after the discriminator obtains the feature vector of the "true" background music and the feature vectors of each "false" background music, it will calculate the distance between the feature vector of the "true" background music and the feature vector of each "false" background music, and obtain d1, d2, ..., d k ; Among them, the distance between the feature vector of the “real” background music and the feature vector of the “fake” background music can represent the similarity between the “real” background music and the “fake” background music. In addition, the discriminator also calculates the distance between the feature vector of the “real” background music and itself, and obtains d k+1 .
[0053] Based on this, in some embodiments, the background music generated by the generator cannot be distinguished by the discriminator as a training target, and the adversarial generative network is trained, specifically including: obtaining the similarity between the background music input to the discriminator and the background music included in the training sample, and determining the value of the loss function based on the similarity; adjusting the parameters of the background music generator with the goal of minimizing the value of the loss function to train the adversarial generative network.
[0054] The training goal of the adversarial generative network is to make the background music generated by the generator indistinguishable from the discriminator. Specifically, the distance between the feature vector of the "fake" background music and the feature vector of the "real" background music is as close as possible, and the d1, d2, ..., d generated by the discriminator are k As much as possible with d k+1 consistent.
[0055] The specific training process is as follows: after each generator generates "fake" background music, first fix each generator, and then train the discriminator. The corresponding loss function is as follows:
[0056]
[0057] The larger the first term, the more accurately the discriminator can identify the real sample as the mathematical expectation of the real sample. In this embodiment, there is only one real background music, and the first term is 1; the larger the second term, the more accurately the discriminator can distinguish the fake sample from the real sample. D(x) can be expressed as the distance between different embedding vectors calculated by the discriminator, that is, Figure 3 d1, d2, ..., d shown k , d k+1 , from the loss function, d1,d2,…,d k and d k+1 The smaller , the smaller D(G(z)).
[0058] In each epoch, after the discriminator training is completed, the discriminator can be fixed and then the generator can be trained. The corresponding loss function is as follows:
[0059]
[0060] Since we want the generator to be stronger, we can set D(G(z)) to be larger. By using a larger D(G(z)), we can adjust the parameters of the generator to a greater extent, so that the "fake" background music generated by the generator is closer to the "real" background music. Similarly, in the above loss function, the larger the first term, the more accurately the discriminator can identify the real sample as the mathematical expectation of the real sample. In this embodiment, there is only one real background music, and the first term is 1; the larger the second term, the more accurately the discriminator can distinguish between fake samples and real samples. Among them, D(x) can be expressed as the distance between different embedding vectors calculated by the discriminator, that is, Figure 3 d1, d2, ..., d shown k , d k+1 , from the loss function, d1,d2,…,d k and d k+1 The smaller , the smaller D(G(z)).
[0061] When the value of the loss function converges, the generator at this time can be used as a background music generator and the corresponding parameters can be saved.
[0062] In one embodiment, after the computer device obtains the background music generator, it can perform the following steps: input the frequency spectrum of the long audio to be matched with music into the background music generator to obtain the audio frequency spectrum output by the background music generator; convert the audio frequency spectrum into background music, and use the background music to match the long audio to be matched with music.
[0063] Furthermore, converting the audio spectrum into background music may include: performing an inverse Fourier transform on the audio spectrum to obtain an initial time domain waveform; if the initial time domain waveform is unstable, performing multiple iterations with the initial time domain waveform as the input of the first iteration, and outputting a time domain waveform for each iteration; wherein, in each iteration, extracting the phase of the time domain waveform input for that iteration and performing a Fourier transform to obtain a spectrum, combining the spectrum and the phase of the time domain waveform input for that iteration and performing an inverse Fourier transform to obtain a time domain waveform as the output of that iteration; if the time domain waveform output by the iteration is unstable, using the time domain waveform as the input of the next iteration; if the time domain waveform output by the iteration is stable, stopping the iteration, and obtaining background music based on the stable time domain waveform.
[0064] In this embodiment, in the prediction stage, the frequency spectrum of the long audio to be matched with music is input into the background music generator, and the corresponding audio frequency spectrum is obtained at the output end of the background music generator.
[0065] In order to convert the audio spectrum into background music, the Griffin-Lim algorithm or other algorithms can be used to perform Fourier transform. Figure 4 As shown, the computer device can perform inverse Fourier transform on the audio spectrum, and the resulting time domain waveform is called the initial time domain waveform Y w (mT,ω)|, extract the phase θ of the initial time domain waveform i , and then Fourier transform to get the spectrum Combined with this spectrum and the phase θ of the initial time domain waveform i , perform inverse Fourier transform and get the time domain waveform x i+1 (n), if the time domain waveform x i+1 (n) is stable, then the iteration is stopped and the background music is obtained based on the stable time domain waveform; if the time domain waveform x i+1 (n) is unstable, then the time domain waveform x i+1 (n) is used as the input for the next iteration, and the time domain waveform x i+1 Phase θ of (n) i+1 Perform Fourier transform to obtain the spectrum Combined with this spectrum And the time domain waveform x i+1 Phase θ of (n) i+1 , perform inverse Fourier transform to obtain the next time domain waveform, and iterate until the obtained time domain waveform is stable.
[0066] In the above embodiment, more stable background music can be obtained by iteratively processing the audio spectrum generated by the background music generator through Fourier transform.
[0067] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0068] In one embodiment, a computer device is provided, whose internal structure diagram can be as follows: Figure 5As shown. The computer device includes a processor, a memory and a network interface connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store long audio soundtrack data. The network interface of the computer device is used to communicate with an external terminal via a network connection. The computer device also includes an input and output interface, which is a connection circuit for exchanging information between the processor and an external device. They are connected to the processor via a bus, referred to as an I / O interface. When the computer program is executed by the processor, a training method for a background music generator is implemented.
[0069] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0070] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.
[0071] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0072] In one embodiment, a computer program product is provided, on which a computer program is stored. The computer program is used by a processor to execute the steps in the above-mentioned various method embodiments.
[0073] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the above-mentioned computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0074] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0075] The above embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.
Claims
1. A training method for a background music generator, characterized in that: The method comprises: Acquire at least one background music set, wherein each background music set has a corresponding long audio category; Adapting at least one long audio for each piece of background music in the background music set according to the long audio category corresponding to the background music set; Each piece of background music in the at least one background music set and the long audio adapted to the background music are used as training samples; the training sample includes a piece of background music and multiple long audios matched to the background music; the training sample is used to train the adversarial generative network, and the generator in the trained adversarial generative network is used as the background music generator; including: inputting each long audio contained in any of the training samples into each generator in the adversarial generative network respectively to obtain the background music generated by each generator; inputting the background music generated by each generator into the discriminator of the adversarial generative network, so that the discriminator can distinguish whether the input background music is the background music included in the training sample or the background music generated by each generator; and training the adversarial generative network with the background music generated by each generator not being distinguishable by the discriminator as the training target.
2. The method according to claim 1, characterized in that Get at least one background music collection, including: Get multiple background music and multiple long audio files; different long audio files belong to different long audio categories; Determining the best-matching long audio for each piece of background music among the multiple long audios; The background music of the best-matched long audios belonging to the same long audio category are grouped into the same set to obtain at least one background music set.
3. The method according to claim 2, characterized in that Determining the best matching long audio for each piece of background music among the multiple long audios includes: Extracting the feature vectors of each of the long audios and the feature vectors of each piece of background music; The distances between the feature vectors of each background music and the feature vectors of each long audio are obtained, and the long audio with the smallest distance is taken as the long audio that best matches the background music.
4. The method according to claim 1, wherein Adapting at least one long audio file for each piece of background music in the background music collection includes: Obtaining a feature vector of each piece of background music in the background music set, and a feature vector of each of a plurality of long audios in a long audio category corresponding to the background music set; Obtain the distance between the feature vector of each background music and the feature vector of each long audio, and take the long audio corresponding to the distance less than the threshold as the long audio adapted to the background music.
5. The method according to claim 1, wherein The adversarial generative network includes a generator and a discriminator; and using the training sample to train the adversarial generative network includes: Inputting the long audio included in the training sample into the generator to obtain background music generated by the generator; Inputting the background music generated by the generator into the discriminator, so that the discriminator distinguishes whether the input background music is the background music included in the training sample or the background music generated by the generator; The adversarial generative network is trained by taking the background music generated by the generator as a training target that cannot be distinguished by the discriminator.
6. The method according to claim 5, characterized in that The background music generated by the generator cannot be distinguished by the discriminator as a training target, and the adversarial generative network is trained, including: Obtaining a similarity between the background music input to the discriminator and the background music included in the training sample, and determining a value of a loss function according to the similarity; The parameters of the background music generator are adjusted with the goal of minimizing the value of the loss function to train the adversarial generation network.
7. The method according to claim 1, characterized in that After using the generator in the trained adversarial generative network as a background music generator, the method further includes: Inputting the frequency spectrum of the long audio to be matched with music into the background music generator to obtain the audio frequency spectrum output by the background music generator; The audio spectrum is converted into background music, and the background music is used to compose music for the long audio to be composed.
8. The method according to claim 7, characterized in that Converting the audio spectrum into background music, comprising: Obtaining an audio amplitude spectrum corresponding to the audio spectrum and an initial time domain waveform phase; Obtaining a first frequency spectrum based on the audio amplitude spectrum and the phase of the initial time-domain waveform, and performing an inverse Fourier transform on the first frequency spectrum to obtain a time-domain waveform corresponding to the first frequency spectrum; If the time domain waveform corresponding to the first spectrum is unstable, performing a Fourier transform on the time domain waveform corresponding to the first spectrum to obtain a second spectrum, using the time domain waveform phase corresponding to the second spectrum as a new initial time domain waveform phase, and returning to the step of obtaining the first spectrum based on the audio amplitude spectrum and the initial time domain waveform phase until the time domain waveform corresponding to the first spectrum is stable; If the time domain waveform corresponding to the first spectrum is stable, background music is obtained based on the time domain waveform corresponding to the first spectrum.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 8 is implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Traditional opera synthesis method and apparatus, and computer readable storage medium
CN108766409A
Video music dubbing method and device, electronic equipment and storage medium
CN113572981A
Audio Generation Methods and Systems
US20230018661A1
Method, system, and medium for affective music recommendation and composition
US20230113072A1