Environmental sound generation method and apparatus, computer device, and storage medium

By using an environmental sound generation model based on generative adversarial networks and a target vocoder, the problem of insufficient realism of environmental sounds in existing technologies is solved, and personalized and highly realistic environmental sound generation is achieved.

CN115273807BActive Publication Date: 2025-12-30PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210903105.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-29
Publication Date
2025-12-30
Estimated Expiration
2042-07-29

AI Technical Summary

Technical Problem

Existing ambient sound generation models, without a large amount of training data, produce ambient sounds with low realism, failing to meet users' personalized business needs.

Method used

An ambient sound generation model based on generative adversarial networks is used to process the ambient sound tag vector, obtain the target spectrogram, and convert it into the target ambient sound through a target vocoder.

Benefits of technology

It improves the realism of generated environmental sounds, meeting users' personalized environmental sound generation needs under different conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115273807B_ABST
    Figure CN115273807B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of artificial intelligence, and discloses an environmental sound generation method and device, computer equipment and a storage medium.The method obtains environmental sound label data according to user requirements, vectorizes environmental audio data in the environmental sound label data and time identifiers corresponding to the environmental audio data, obtains environmental sound label vectors corresponding to the environmental sound label data, and uses the environmental sound label vectors for subsequent environmental sound generation.An environmental sound generation model based on a generative adversarial network is used to process the environmental sound label vectors to obtain target spectrograms corresponding to generated environmental sounds, and a target sound decoder is used to process the target spectrograms to obtain target environmental sounds corresponding to each time identifier in the environmental sound label data, thereby meeting different user requirements and improving the authenticity of the generated environmental sounds.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to an environmental sound generation method, apparatus, computer device, and storage medium. Background Technology

[0002] With the development of the entertainment industry, games and movies have become major forms of public entertainment, bringing joy to people's lives. These industries have evolved over time, requiring a significant amount of ambient sound to provide an immersive experience.

[0003] In existing technologies, the ambient sounds used in games or movies are usually pre-stored ambient sounds from different scenes and played in a loop. However, due to the diversity of ambient sounds used in games or movies, the workload of acquiring ambient sounds is extremely large. Furthermore, when using ambient sound generation models to generate ambient sounds, without a large amount of training data, the generated ambient sounds are relatively monotonous, unable to meet the personalized business needs of users, and lack realism. Summary of the Invention

[0004] This invention provides an environmental sound generation method, apparatus, computer device, and storage medium, which solves the problem that the environmental sounds generated by existing environmental sound generation models have low realism.

[0005] This invention provides a method for generating ambient sound, comprising:

[0006] Acquire ambient sound tag data, wherein the ambient sound tag data includes ambient audio data and a time identifier corresponding to the ambient audio data;

[0007] The environmental audio data and the time stamp are vectorized to obtain the environmental sound tag vector;

[0008] An ambient sound generation model based on generative adversarial networks is used to process the ambient sound tag vector to obtain the target spectrogram;

[0009] A target vocoder is used to transform the target spectrogram to obtain the target environmental sound corresponding to each time marker.

[0010] This invention also provides an environmental sound generation device, comprising:

[0011] An ambient sound tag data acquisition module is used to acquire ambient sound tag data, wherein the ambient sound tag data includes ambient audio data and a time identifier corresponding to the ambient audio data;

[0012] An ambient sound tag vector acquisition module is used to perform vectorization processing on the ambient audio data and the time identifier to obtain the ambient sound tag vector.

[0013] The target spectrogram acquisition module is used to process the ambient sound tag vector using an ambient sound generation model based on a generative adversarial network to obtain the target spectrogram.

[0014] The target environment sound acquisition module is used to convert the target spectrogram using a target vocoder to acquire the target environment sound corresponding to each time marker.

[0015] This invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described method for generating ambient sound.

[0016] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for generating ambient sound.

[0017] The aforementioned environmental sound generation method, apparatus, computer equipment, and storage medium acquire environmental sound tag data according to user needs, vectorize the environmental audio data and the corresponding time markers in the environmental sound tag data to obtain the environmental sound tag vectors corresponding to the environmental sound tag data, which are then used for subsequent environmental sound generation. An environmental sound generation model based on a generative adversarial network is used to process the environmental sound tag vectors to obtain the target spectrogram corresponding to the generated environmental sound. The target spectrogram is then transformed using a target vocoder to obtain the target environmental sound corresponding to each time marker in the environmental sound tag data, thereby meeting different user needs and improving the realism of the generated environmental sound. Attached Figure Description

[0018] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a schematic diagram of an application environment for an environmental sound generation method according to an embodiment of the present invention;

[0020] Figure 2 This is a flowchart of an environmental sound generation method according to an embodiment of the present invention;

[0021] Figure 3 This is another flowchart of an environmental sound generation method in one embodiment of the present invention;

[0022] Figure 4 This is another flowchart of an environmental sound generation method in one embodiment of the present invention;

[0023] Figure 5 This is another flowchart of an environmental sound generation method in one embodiment of the present invention;

[0024] Figure 6 This is a schematic diagram of an environmental sound generation device according to an embodiment of the present invention;

[0025] Figure 7 This is a schematic diagram of a computer device according to an embodiment of the present invention. Detailed Implementation

[0026] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0027] The environmental sound generation method provided in this embodiment of the invention can be applied to, for example... Figure 1 The application environment shown. Figure 1 As shown, the client (computer device) communicates with the server via a network. The client, also known as the user terminal, refers to the program that provides local services to the client, corresponding to the server. Client (computer device) includes, but is not limited to, various personal computers, laptops, smartphones, tablets, cameras, and portable wearable devices. The server can be a standalone server or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0028] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0029] The environmental sound generation method provided in this embodiment of the invention can be applied to, for example... Figure 1 The application environment is shown. Specifically, this environmental sound generation method is applied in an environmental sound generation system, which includes, for example, […]. Figure 1 The client and server shown communicate with each other over a network to generate and process environmental sound tag data corresponding to the target user, so as to obtain environmental sounds corresponding to more realistic environmental sound tag data.

[0030] In one embodiment, such as Figure 2 As shown, an environmental sound generation method is provided, which can be applied to... Figure 1 Taking the server in the example, the following steps are included:

[0031] S201: Obtain ambient sound tag data, which includes ambient audio data and the time stamp corresponding to the ambient audio data;

[0032] S202: Vectorize the ambient audio data and time stamp to obtain the ambient sound tag vector;

[0033] S203: An ambient sound generation model based on generative adversarial networks is used to process the ambient sound tag vectors and obtain the target spectrogram;

[0034] S204: Using a target vocoder, the target spectrogram is transformed to obtain the target environmental sound corresponding to each time identifier.

[0035] Ambient sound refers to the sounds associated with scenes in games or movies. For example, in a game, when a player enters a scene, the audio engine automatically plays the ambient sound effects until they leave. These ambient sound effects are played back according to certain rules to simulate a more "natural" atmosphere. Common ambient sounds include natural sounds such as wind and rain, and human-made sounds such as the clicking of keyboards in a public office area.

[0036] As an example, in step S201, the server acquires the ambient sound tag data set by the user according to their needs. This ambient sound tag data includes ambient audio data and a corresponding time stamp. The ambient audio data is ambient audio data with an ambient sound type tag, such as ambient audio data tagged with "thunder." The time stamp is the time stamp of the frame corresponding to the ambient audio data. In this example, the user can set the corresponding ambient sound type tag on the corresponding time stamp according to business needs, thereby controlling the ambient audio type corresponding to each frame of the ambient audio data to obtain the required ambient audio data.

[0037] As an example, in step S202, after obtaining the ambient sound tag data, the server vectorizes the ambient audio data and corresponding time markers in the ambient sound tag data to obtain the ambient sound tag vector for the ambient sound generation model. In this example, if time markers 1 to n frames represent ambient audio data A, and time markers n+1 to m frames represent ambient audio data B, a vector z = {z1,...,z...} is constructed. n |z n+1 ,...,z m}, and for any z i have Among them l i Let be the label vector at time i.

[0038] The methods for converting environmental audio data and corresponding time markers into vectors include, but are not limited to, using a pre-set vector lookup table. Vector conversion can also be performed using a word2vec word segmenter or a one-hot encoder. No further restrictions are imposed.

[0039] As an example, in step S203, after confirming the ambient sound label vector, the server uses an ambient sound generation model based on generative adversarial networks to process the ambient sound label vector to obtain the target spectrogram corresponding to the generated target ambient sound. By generating ambient sound through a trained ambient sound generation model, the realism of the generated target ambient sound can be improved.

[0040] Generative Adversarial Networks (GANs) are a type of unsupervised learning neural network. A GAN consists of two models: a generative model and a discriminative model. The generative model generates instances that appear natural and realistic, similar to the original data. The discriminative model determines whether a given instance appears natural and realistic or is artificially created. Real instances come from the dataset, while artificial instances come from the generative model. This example uses a GAN to build an ambient sound generation model for generating ambient sounds.

[0041] The spectrogram is a frequency distribution graph that can be converted to and from the speech signal. This complex signal mainly consists of signals containing different frequency components. The horizontal axis of the spectrogram represents frequency, and the vertical axis represents amplitude. The frequency is obtained through Fourier transform.

[0042] As an example, in step S204, after obtaining the generated target spectrogram, the server uses a target vocoder to convert the target spectrogram into a signal, obtaining the target ambient sound corresponding to the target spectrogram. The time marker in the target ambient sound corresponds to the target ambient sound, ensuring that the generated target ambient sound meets the user's business requirements. In this example, a FreGAN-based target vocoder is used to convert the target spectrogram into a sound waveform file, i.e., the target ambient sound.

[0043] In this example, by acquiring the ambient sound tag data required by the user, the ambient audio data and the corresponding time markers in the ambient sound tag data are vectorized to obtain the ambient sound tag vectors corresponding to the ambient sound tag data, which are then used for subsequent ambient sound generation. An ambient sound generation model based on generative adversarial networks is used to process the ambient sound tag vectors to obtain the target spectrogram corresponding to the generated ambient sound. The target spectrogram is then transformed by a target vocoder to obtain the target ambient sound corresponding to each time marker in the ambient sound tag data, thereby meeting different user needs and improving the realism of the generated ambient sound.

[0044] In one embodiment, such as Figure 3 As shown, step S202: Vectorize the environmental audio data and time stamps to obtain the environmental sound tag vector, including:

[0045] S301: Perform vectorization processing on the environmental audio data to obtain the audio type vector;

[0046] S302: Vectorize the time stamp to obtain the time stamp vector;

[0047] S303: Concatenate the audio type vector and the time identifier vector to obtain the ambient sound tag vector.

[0048] As an example, in step S301, after obtaining environmental audio data with environmental audio data and time identifiers, the server first performs vectorization processing on the environmental audio data to obtain an audio type vector. In this example, the environmental audio data is vectorized according to the pre-set environmental audio category label vector to obtain an audio type vector that can be used for model recognition.

[0049] As an example, in step S302, after obtaining the audio type vector, the server performs vectorization processing on the time identifier to obtain the time identifier vector.

[0050] As an example, in step S303, after confirming the audio type vector and the time identifier vector, the server concatenates the audio type vector and the time identifier vector to obtain an ambient sound tag vector containing both the time identifier vector and the audio type vector. In this example, the audio type vector and the time identifier vector can also be associated through mapping to obtain an ambient sound tag vector containing both the time identifier vector and the audio type vector.

[0051] In this example, the acquired ambient sound tag data with ambient audio data and time identifiers is converted into corresponding vectors and concatenated to obtain an ambient sound tag vector with a time identifier vector and an audio type vector, which is used for model recognition. At the same time, the ambient sound tag vector is guaranteed to have multiple features.

[0052] In one embodiment, such as Figure 4 As shown, step S203: For the ambient sound generation model built using a generative adversarial network, the ambient sound tag vector is processed to obtain the target spectrogram, including:

[0053] S401: The ambient sound generation model based on generative adversarial networks includes N generation modules connected in sequence and N head modules connected to each generation module;

[0054] S402: An ambient sound generation model based on generative adversarial networks is used to process the ambient sound label vectors and obtain the target spectrogram.

[0055] As an example, in step S401, the ambient sound generation model used by the server includes N generation modules connected in sequence and N head modules connected to each generation module to generate the target ambient sound. In this example, the generation modules include grouped convolutions and Bi-GRU recurrent neural network side branches to ensure the efficiency and smoothness of the generated target spectrogram. The head modules include residual side branches to ensure the dimensionality of the target spectrogram and the smooth transmission of dimensional information of the generated spectrogram.

[0056] Setting N corresponding generation modules and header modules means that the ambient sound tag vector is amplified N times, and the number of data amplification times and amplification coefficients are not limited.

[0057] As an example, in step S402, after obtaining the generated target spectrogram, the server uses a target vocoder to convert the target spectrogram into a signal and obtain the target ambient sound corresponding to the target spectrogram. The time marker in the target ambient sound corresponds to the target ambient sound to ensure that the generated target ambient sound meets the user's business requirements.

[0058] In this example, N sequentially connected generation modules and N head modules connected to each generation module are used to generate target environmental sounds. At the same time, an environmental sound generation model based on generative adversarial networks is used to process the environmental sound tag vectors to obtain the target spectrogram corresponding to the generated environmental sounds, thereby optimizing the generated target spectrogram for greater efficiency and smoothness.

[0059] In one embodiment, step S402: using an ambient sound generation model based on a generative adversarial network to process the ambient sound tag vector and obtain the target spectrogram, including:

[0060] S4021: Using the current generation module, process the current input information of the current frame to obtain the first generated spectrum corresponding to the current frame, process the first generated spectrum to obtain the first upsampled spectrum corresponding to the current frame; the current input information is the ambient sound tag vector or the first upsampled spectrum output by the previous generation module;

[0061] S4022: Using the current header module, process the first generated spectrum corresponding to the current frame to obtain the second generated spectrum corresponding to the current frame, perform upsampling processing on the second generated spectrum to obtain the second upsampled spectrum corresponding to the current frame;

[0062] S4023: Obtain the target spectrum based on N second upsampled spectra.

[0063] Upsampling, also known as signal interpolation, involves inserting n zeros between two points in the original sequence, which is equivalent to spectral compression in the frequency domain. Upsampling, combined with a filter, can increase the sampling frequency by a certain multiple, which means that the corresponding dimensionality is amplified.

[0064] As an example, in step S4021, the server uses the current generation module in the ambient sound generation model to process the ambient sound tag vector for the first time. The generation module generates the corresponding first generated spectrum and performs upsampling processing on the first generated spectrum to obtain the first upsampled spectrum corresponding to the current frame.

[0065] In this example, the upsampling module uses an interpolation function to augment the dimension of the first generated spectrum, obtaining the augmented first upsampled spectrum. Let's assume the initial generated spectrum is f. x Each frame is amplified by a coefficient of 2, and through N=4 corresponding generation modules, corresponding to 4 upsampling operations, f is obtained. x =16×f z This means that the spectral dimension per frame has increased by 16 times.

[0066] In this example, the second current generation module processes the first upsampled spectrum output by the previous generation module, which means it performs a second-dimensional amplification on the first generated spectrum from the previous generation, and the number of amplifications is N.

[0067] As an example, in step S4022, the server uses the current header module in the ambient sound generation model to process the first generated spectrum corresponding to the current frame, obtain the second generated spectrum corresponding to the current frame, and then performs upsampling processing on the second generated spectrum to obtain the second upsampled spectrum corresponding to the current frame. In this example, the upsampling method is the same as that of the generation module, both of which perform corresponding dimensionality augmentation.

[0068] As an example, in step S4023, after acquiring the second upsampled spectrum, the server obtains the final second upsampled spectrum based on N second upsampled spectra, which is used as the target spectrogram after dimensional amplification for subsequent target environmental sound conversion.

[0069] In this example, the current generation module in the ambient sound generation model first processes the ambient sound tag vector to obtain the corresponding first generated spectrum. After upsampling, the first upsampled spectrum with increased dimensions is obtained, which is then used by the second generation module to process the first upsampled spectrum. At the same time, the head module processes the first generated spectrum and upsamples it to obtain the second upsampled spectrum with increased dimensions. After N iterations as required, the final target spectrogram is obtained, thereby improving the efficiency and smoothness of the generated target spectrogram, as well as the dimensionality of the target spectrogram.

[0070] In another embodiment, such as Figure 5 As shown, before step S201: obtaining ambient sound tag data, the following steps are also included:

[0071] S501: Acquire training ambient sound data;

[0072] S502: Convert and process the training ambient sound data to obtain the training spectrogram;

[0073] S503: The generator in the generative adversarial network is used to generate and process the training ambient sound data to obtain the generated spectrogram;

[0074] S504: Use the discriminator in a generative adversarial network to discriminate between the generated spectrogram and the training spectrogram, and obtain the model loss value;

[0075] S505: Based on the model loss value, construct an ambient sound generation model using a generative adversarial network.

[0076] In generative adversarial networks, the discriminator and the generator are trained separately. The discriminator is trained by fixing the parameters of the generator, inputting real cases into the discriminator and outputting a result labeled as 1, and inputting the output of the generator into the discriminator and obtaining an output result labeled as 0. The discriminator is trained until it converges.

[0077] The generator is trained by fixing the discriminator's parameters, inputting random noise into the generated spectrogram, and then inputting the result into the discriminator with a label of 1, training the generator to convergence. This process is repeated cyclically to obtain a more reliable model-based generative adversarial network.

[0078] As an example, in step S501, the server obtains training ambient sound data for training the ambient sound generation model.

[0079] As an example, in step S502, the server converts the acquired training ambient sound data into a corresponding training spectrum for use. The discriminator uses the generator as a real case to judge the generated spectrogram generated by the generator, thereby training the generative adversarial network.

[0080] As an example, in step S503, the server uses a generator in a generative adversarial network to process the training ambient sound data and obtain the generated spectrogram generated by the generator, which is then used by the discriminator to judge it, thereby improving the authenticity of the generated spectrogram generated by the generator.

[0081] In this example, a loop structure can be used to determine the quality of the generated spectrogram. By introducing a restoration function, the generated spectrogram can be restored to the generated ambient sound, that is, the generated ambient sound can be used to generate a spectrogram, thereby calculating the loss value and improving the generator parameters.

[0082] As an example, in step S504, the server uses a discriminator in a generative adversarial network to verify the training spectrogram, fix the discriminator parameters, and then performs discriminative processing on the generated spectrogram to obtain the model loss value for updating the network parameters.

[0083] In this example, the discriminator structure first extracts the spectrogram through three 2D convolutional layers, then combines the 2D feature dimensions into 1D through the Flatten layer for temporal information inference, then uses four 1D convolutional layers with Bi-GRU side branches for comprehensive inference of short-range temporal information, and finally uses a three-layer self-attention mechanism for comprehensive inference of long-range temporal information.

[0084] As an example, in step S505, the server updates the generator's parameters based on the model loss value, thereby completing the generator's convergence and constructing the environmental sound generation model built by the generative adversarial network.

[0085] In this example, the discriminator and generator are trained alternately to train each other and improve the reliability of the ambient sound generation model.

[0086] In this example, the generator is trained by using a discriminator with fixed parameters in a generative adversarial network, thereby improving the realism of the spectrograms generated by the generator and thus improving the reliability of the ambient sound generation model.

[0087] In another embodiment, before step S201: acquiring ambient sound tag data, the method further includes:

[0088] S501A: Obtain the raw vocoder;

[0089] S502A: Based on the training ambient sound data, fine-tune the original vocoder to obtain the target vocoder.

[0090] As an example, in step S501A, the server obtains the original vocoder for fine-tuning. In this example, a FreGAN-based original vocoder is used. The original vocoder can be trained using a human voice dataset to obtain the corresponding general background model.

[0091] Among them, the Universal Background Model (UBM) is a Gaussian mixture model that represents the distribution of speech features of a large number of non-speaker-specific speech. The training of the Universal Background Model usually uses a large amount of speech data that is independent of specific speakers and channels.

[0092] As an example, in step S502A, the server fine-tunes the original vocoder by acquiring the training environment sound used for training, thereby obtaining the generation of the environment sound used in this application and improving the accuracy of the sound conversion of the spectrum generated by the model.

[0093] In this example, a FreGAN-based vocoder is trained using a human voice dataset to obtain a general background model. Then, based on the generation of environmental sounds in this application, appropriate parameters are fine-tuned to improve the transcoding accuracy.

[0094] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0095] In one embodiment, an ambient sound generation device is provided, which corresponds one-to-one with the ambient sound generation methods described in the above embodiments. For example... Figure 6 As shown, the environmental sound generation device includes an environmental sound tag data acquisition module 601, an environmental sound tag vector acquisition module 602, a target spectrogram acquisition module 603, and a target environmental sound acquisition module 604. Detailed descriptions of each functional module are as follows:

[0096] The ambient sound tag data acquisition module 601 is used to acquire ambient sound tag data, which includes ambient audio data and the time identifier corresponding to the ambient audio data.

[0097] The ambient sound tag vector acquisition module 602 is used to perform vectorization processing on ambient audio data and time stamps to obtain ambient sound tag vectors.

[0098] The target spectrogram acquisition module 603 is used to process the ambient sound tag vector and acquire the target spectrogram by employing an ambient sound generation model based on a generative adversarial network.

[0099] The target environment sound acquisition module 604 is used to convert the target spectrogram using a target vocoder to acquire the target environment sound corresponding to each time identifier.

[0100] In one embodiment, the ambient sound tag vector acquisition module 602 includes:

[0101] The audio type vector acquisition unit is used to vectorize environmental audio data and obtain audio type vectors.

[0102] The time identifier vector acquisition unit is used to vectorize the time identifier and obtain the time identifier vector.

[0103] The ambient sound tag vector acquisition unit is used to concatenate the audio type vector and the time identifier vector to obtain the ambient sound tag vector.

[0104] In one embodiment, the target spectrum acquisition module 603 includes:

[0105] The first upsampled spectrum acquisition unit is used to process the current input information of the current frame using the current generation module, obtain the first generated spectrum corresponding to the current frame, process the first generated spectrum, and obtain the first upsampled spectrum corresponding to the current frame; the current input information is the ambient sound tag vector or the first upsampled spectrum output by the previous generation module;

[0106] The second upsampled spectrum acquisition unit is used to process the first generated spectrum corresponding to the current frame using the current header module, obtain the second generated spectrum corresponding to the current frame, and perform upsampling processing on the second generated spectrum to obtain the second upsampled spectrum corresponding to the current frame.

[0107] The target spectrogram acquisition unit is used to acquire the target spectrogram based on N second upsampled spectra.

[0108] In another embodiment, the ambient sound generating device further includes:

[0109] The training environment sound data acquisition module is used to acquire training environment sound data;

[0110] The training spectrogram acquisition module is used to transform and process the training ambient sound data to obtain the training spectrogram.

[0111] The spectrogram acquisition module is used to process the training environment sound data using a generator in a generative adversarial network to obtain a generated spectrogram.

[0112] The model loss value acquisition module is used to discriminate between the generated spectrogram and the training spectrogram using a discriminator in a generative adversarial network, and to obtain the model loss value.

[0113] The ambient sound generation model acquisition module is used to construct an ambient sound generation model based on the model loss value using a generative adversarial network.

[0114] In yet another embodiment, the ambient sound generating apparatus further includes:

[0115] The raw vocoder acquisition module is used to acquire the raw vocoder.

[0116] The target vocoder acquisition module is used to fine-tune the original vocoder based on the training ambient sound data to acquire the target vocoder.

[0117] Specific limitations regarding the environmental sound generation device can be found in the limitations of the environmental sound generation method described above, and will not be repeated here. Each module in the aforementioned environmental sound generation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0118] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 7As shown. The computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system, computer programs, and the database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database is used to store data employed or generated during the execution of the environmental sound generation method. The network interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements an environmental sound generation method.

[0119] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the environmental sound generation method described in the above embodiment, for example... Figure 2 As shown in S201-S204, or Figures 3 to 5 As shown, to avoid repetition, it will not be described again here. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit in this embodiment of the ambient sound generation device, for example... Figure 6 The functions of the ambient sound tag data acquisition module 601, ambient sound tag vector acquisition module 602, target spectrogram acquisition module 603, and target ambient sound acquisition module 604 shown are not described again here to avoid repetition.

[0120] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When executed by a processor, the computer program implements the environmental sound generation method described in the above embodiment, for example... Figure 2 As shown in S201-S204, or Figures 3 to 5 As shown, to avoid repetition, it will not be described again here. Alternatively, when the computer program is executed by the processor, it implements the functions of each module / unit in this embodiment of the environmental sound generation device, for example... Figure 6 The functions of the ambient sound tag data acquisition module 601, ambient sound tag vector acquisition module 602, target spectrogram acquisition module 603, and target ambient sound acquisition module 604 shown are not described again here to avoid repetition.

[0121] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0122] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0123] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. An ambient sound generation method, characterized by, The method comprises the following steps: acquiring environmental sound label data, wherein the environmental sound label data comprises environmental audio data and time identifiers corresponding to the environmental audio data; vectorizing the environmental audio data and the time identifiers to obtain environmental sound label vectors; an environmental sound generation model constructed based on a generative adversarial network comprises N generation modules connected in sequence and N head modules connected to each generation module; using a current generation module to process current input information of a current frame to obtain a first generated spectrum corresponding to the current frame, and processing the first generated spectrum to obtain a first up-sampled spectrum corresponding to the current frame; the current input information is an environmental sound label vector or a first up-sampled spectrum output by a previous generation module; using a current head module to process the first generated spectrum corresponding to the current frame to obtain a second generated spectrum corresponding to the current frame, and performing up-sampling processing on the second generated spectrum to obtain a second up-sampled spectrum corresponding to the current frame; based on the N second up-sampled spectrums, a target spectrum graph is obtained; using a target vocoder to perform conversion processing on the target spectrum graph to obtain a target environmental sound corresponding to each time identifier.

2. The ambient sound generation method of claim 1, wherein, The vectorizing the environmental audio data and the time identifiers to obtain environmental sound label vectors comprises: vectorizing the environmental audio data to obtain an audio type vector; vectorizing the time identifiers to obtain a time identifier vector; splicing the audio type vector and the time identifier vector to obtain an environmental sound label vector.

3. The ambient sound generation method of claim 1, wherein, Before the acquiring environmental sound label data, the environmental sound generation method further comprises: acquiring training environmental sound data; performing conversion processing on the training environmental sound data to obtain training spectrum graphs; using a generator in a generative adversarial network to perform generation processing on the training environmental sound data to obtain generated spectrum graphs; using a discriminator in the generative adversarial network to perform discrimination processing on the generated spectrum graphs and the training spectrum graphs to obtain a model loss value; constructing an environmental sound generation model constructed based on a generative adversarial network according to the model loss value.

4. The ambient sound generation method of claim 1, wherein, Before the acquiring environmental sound label data, the environmental sound generation method further comprises: acquiring an original vocoder; fine-tuning the original vocoder according to training environmental sound data to obtain a target vocoder.

5. An ambient sound generating apparatus, characterized by, The method comprises the following steps: an environmental sound label data acquisition module is configured to acquire environmental sound label data, wherein the environmental sound label data comprises environmental audio data and time identifiers corresponding to the environmental audio data; an environmental sound label vector acquisition module is configured to vectorize the environmental audio data and the time identifiers to obtain environmental sound label vectors; The target spectrum obtaining module, wherein the ambient sound generation model constructed based on the generative adversarial network comprises N generation modules connected in sequence and N head modules connected with each generation module, the target spectrum obtaining module is configured to use a current generation module to process current input information of a current frame to obtain a first generated spectrum corresponding to the current frame, process the first generated spectrum to obtain a first up-sampled spectrum corresponding to the current frame, the current input information is an ambient sound label vector or a first up-sampled spectrum output by a previous generation module; use a current head module to process the first generated spectrum corresponding to the current frame to obtain a second generated spectrum corresponding to the current frame, and perform up-sampling processing on the second generated spectrum to obtain a second up-sampled spectrum corresponding to the current frame; and obtain a target spectrum based on N second up-sampled spectrums; The target ambient sound obtaining module is configured to use a target vocoder to perform conversion processing on the target spectrum to obtain a target ambient sound corresponding to each time identifier.

6. The ambient sound generating apparatus of claim 5, wherein, The ambient sound label vector obtaining module is configured to include: An audio type vector obtaining unit configured to perform vectorization processing on the ambient audio data to obtain an audio type vector; A time identifier vector obtaining unit configured to perform vectorization processing on the time identifier to obtain a time identifier vector; An ambient sound label vector obtaining unit configured to perform splicing processing on the audio type vector and the time identifier vector to obtain an ambient sound label vector.

7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the ambient sound generation method of any one of claims 1 to 4.

8. A computer-readable storage medium storing a computer program, the computer-readable storage medium comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 7. The computer program is executed by the processor to implement the ambient sound generation method of any one of claims 1 to 4.

Citation Information

Patent Citations

  • Audio signal generation method, device and equipment and storage medium

    CN112712812A

  • Event identification method and device, equipment and storage medium

    CN113239872A