Audio generation method and system

JP2023017710A5Active Publication Date: 2025-05-28SONY COMP ENTERTAINMENT EURO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2022109692
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-07-16
Filing Date
2022-07-07
Publication Date
2025-05-28
Estimated Expiration
2042-07-07

AI Technical Summary

Technical Problem

The generation of diverse audio assets for video games is time-consuming, costly, and intellectually demanding, with existing computer solutions being slow and expensive, and it is difficult to create audio assets that vary subtly while maintaining thematic consistency.

Method used

A method using a single-painted image generation model to convert input audio assets into graphic expressions, which are then processed to generate output audio assets efficiently, reducing computational power requirements and enabling batch generation of multiple assets.

Benefits of technology

This approach allows for rapid and cost-effective generation of numerous audio assets with varied characteristics, leveraging a single model to learn deformations and reduce computational power, facilitating efficient audio asset creation for video games.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide an audio generation method and a system.SOLUTION: A method for generating audio assets includes a step of receiving a plurality of input audio assets, a step of converting each of the input audio assets into an input graphic representation, a step of generating an input image from each of the input graphic representations, a step of feeding the input image to a generative model and extracting an output graphic representation from each of the output images in order to learn the generative model and generate one or more output images including the output graphic representations, and a step of converting the output graphic representation into an output audio asset.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to audio generation methods and systems, and more particularly to audio assets for use in a video game environment. [Background technology]

[0002] As video games have grown in size in recent years, generating audio content has become a challenge. Audio designers are required to create increasingly diverse sounds and audio assets alongside the game. For example, in the field of video games, a vast library of audio assets related to sound effects is required, particularly to represent sound effects. To meet the needs of video game events, each audio file requires similar assets with slight variations. For example, a footstep audio asset may require multiple variations of a footstep sound to mimic different real-life footstep sounds and to account for factors that affect the acoustic characteristics (e.g., volume, pitch, tone, and timbre) of the footsteps due to in-game actions (e.g., a player running, crawling, etc.). During the generation of such sounds, each such asset is typically manually generated to provide suitable variations on the base audio asset. This is time-consuming, costly (both computationally and financially), and places a significant intellectual burden on the audio creator.

[0003] Furthermore, in some applications (e.g., video games), it is desirable to be able to generate new audio assets on the fly. Such a process is difficult to implement because audio after a game's release cannot be created by a sound engineer and must be created by a computer process. In such cases, it is usually difficult to create audio assets that vary enough from the original assets so that the overall theme is recognizable to the audience as being close enough to the original assets. While attempts have been made to generate audio assets using computers, these are usually complex and heavy processes, making the process slow and expensive. Such computer solutions also only produce a single result. This means that each new audio asset must be created individually, further increasing the creation time and process cost. Summary of the Invention [Problem to be solved by the invention]

[0004] The present invention aims to alleviate at least some of these problems. [Means for solving the problem]

[0005] According to an aspect of the present disclosure, a method for generating audio assets is provided, the method comprising receiving a plurality of input audio assets, converting the input audio assets into input graphical representations, generating an input multi-channel image by stacking each input graphical representation into a separate channel of an image, feeding the input multi-channel images into a single-image generative model to train the generative model, separating output graphical representations from each output multi-channel image, and converting each output graphical representation into an output audio asset, wherein each output multi-channel image includes the output graphical representation.

[0006] Using a graphical representation of the audio (e.g., spectral images) to train a generative model allows for relatively easy and automated generation of audio assets. Furthermore, a batch approach using multi-channel graphical representations for audio generation reduces the computational power required and allows multiple audio assets to be generated quickly.

[0007] Preferably, the generative model is a single-image generative model. Generating a single image using a spectral image to generate new sounds can significantly reduce the amount of data and the required computational power compared to other generative models. A single-image generative model is trained with a single input image to generate new variations of the input image. This is typically achieved using a fully convolutional discriminator with a limited receptive field (e.g., a patch discriminator) and an incremental growth architecture. One practical problem with such single-image models is that they must be trained each time a new image is generated. In other words, if new variations of two different images are required, two different models are trained (for each image). This is typically time-consuming and expensive to implement and maintain. According to the present invention, a novel approach to batch sound generation using generative models can be used to rapidly generate multiple audio assets. This allows a single-image generative model to be trained with a small database of single-channel graphic representations, resulting in a single model that can create different training sound variations. This allows many new audio assets to be easily generated with relatively little computational power.

[0008] The present invention can use any number of different types of graphical representations of audio assets. For example, audio assets may be converted to / from an audio wavelet representation or a spectral image representation. Preferably, audio assets are converted to / from a spectral image. In this case, converting each audio asset to a graphical representation may include performing a Fourier transform on each audio asset and plotting the intensity in the frequency domain to generate a spectral image as a graphical representation. Spectral images are advantageous in that they allow the audio asset to be represented in frequency space (which represents sound signature information). A single spectral image can represent, for example, mono audio and is obtained by taking the short-time Fourier transform of the audio. Spectral images typically have one channel (when using amplitude or complex representations) or two channels (when using amplitude and phase). It has been found that a one-channel spectral image provides a particularly good representation of audio in the frequency domain. This can be used to quickly and efficiently convert between graphical and audio representations of audio assets without significant loss of detail. When dealing with multi-channel audio (e.g., stereo or ambisonics and 3D audio), multiple channels can also be used.

[0009] Spectral images are typically obtained and decomposed using a Fourier transform (and associated inverse transform), although any suitable function may be used for conversion to and from a spectral image. For example, decomposing each output multi-channel image may include separating an output graphic image from each channel of the multi-channel image and performing an inverse Fourier transform on each output graphic image to obtain one or more output audio assets from each output graphic image. Alternatively or additionally, other inverse transforms (e.g., a wavelet transform instead of a Fourier transform) may also be used on the spectral images to generate a graphic representation of the audio assets.

[0010] The single-image generation model may be a generative adversarial network (GAN) with a patch discriminator. The patch discriminator may be a type of GAN discriminator that identifies only structural loss at the scale of image patches and distinguishes whether each patch in the output image is real or fake. The patch discriminator may operate convolutionally across the image and average all responses to provide the final output of the discriminator. If a GAN is used, generating one or more output multi-channel images may include training the GAN with the input multi-channel images.

[0011] Typically, the output multi-channel image may include an output graphical representation in each channel of the multi-channel image, where each graphical representation may be a spectral image with a single channel.

[0012] The above techniques are particularly suitable for use in video game applications, where a large number of audio assets are required and where a large number of similar but slightly different sounds are particularly useful. Receiving a plurality of input audio assets may include receiving video game information from a video game environment. Generating one or more output multi-channel images may include feeding the video game information into a single image generation model, whereby the output multi-channel images are influenced by the video game information.

[0013] In some embodiments, the input audio assets may be received directly from a microphone input, i.e., receiving a plurality of input audio assets may include receiving an input audio clip from a microphone source.

[0014] According to a further aspect of the present disclosure, there is provided a computer program comprising computer-executable instructions that cause a computer to perform a method comprising one or more of the features described above.

[0015] It will be appreciated that the methods described herein may be performed using conventional or (in addition to or instead of) specialized hardware to which suitable software instructions can be applied.

[0016] Implementation using existing parts of conventional equivalent devices can be in the form of a computer program product having a processor capable of executing instructions recorded on a non-transitory computer-readable medium (e.g., floppy disk, optical disk, hard disk, solid state disk, PROM, RAM, flash memory, or a combination of these recording media), or can be in hardware (e.g., ASIC (application specific integrated circuit), FPGA (field programmable gate array), or other configurable circuitry suitable for conventional devices). Such a computer program can be transmitted via a data signal over a network (e.g., Ethernet, a wireless network, the Internet, or a suitable combination of these networks).

[0017] According to a further aspect of the present disclosure, a system for generating audio assets is provided. The system includes an asset input unit, an image generation unit, and an asset output unit. The asset input unit is configured to receive a plurality of input audio assets, convert the input audio assets into input graphical representations, and generate an input multi-channel image by stacking each input graphical representation into a separate channel of the image. The image generation unit implements a generative model for generating one or more output images based on the input images. The output images include the output graphical representations. The asset output unit is configured to separate the output graphical representations from each output multi-channel image and convert each output graphical representation into an output audio asset.

[0018] It is to be understood that both the foregoing general description and the following detailed description are exemplary, but are not restrictive, of the invention. [Brief explanation of the drawings]

[0019] A more complete understanding of the present disclosure and its many advantages will be obtained by reading the following detailed description in conjunction with the accompanying drawings. [Figure 1] FIG. 1 is a schematic diagram illustrating an example workflow for batch generation of audio files. [Figure 2] 1 is a flow chart illustrating a schematic example of a method according to an embodiment of the present disclosure. [Figure 3] FIG. 1 is a schematic diagram illustrating an example of conversion between audio and image according to an embodiment of the present disclosure. [Figure 4] 1 is a schematic diagram illustrating an example of a system according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0020] Referring now to the drawings (in which like or corresponding parts are similarly numbered), the present invention provides a method for efficiently and effectively generating a number of audio assets. The technique generally includes the steps of receiving an audio clip, converting the audio clip into a graphical representation (e.g., in the form of a spectral image), training a single-image generation model of the spectral images to generate new, transformed spectral images, and converting the new, transformed spectral images into new, transformed audio assets. FIG. 1 illustrates this general approach diagrammatically. It outlines how an input set of audio files 101 are converted into a batch of spectral images 102, combined into a single multi-channel image, passed through a neural network 104 (which is trained to generate new multi-channel images 105), separated into a batch of output spectral images 106, and converted into new audio files 107. As illustrated diagrammatically in FIG. 1, the process takes in a batch of audio samples 101 and outputs a new audio file 107 (which is generally different from the input batch 101).

[0021] One aspect of the present disclosure is a method for generating audio assets. A flowchart of an exemplary method is shown in Figure 2. The method includes the following steps:

[0022] Step 201: Receive an input audio asset.

[0023] In a first step, an input audio asset is prepared and received for processing. The input audio asset may be received by a processor that governs the method and the generation of new audio assets. The input audio asset may be received in a memory that is accessed by such processor, allowing the processor to retrieve the audio asset when needed.

[0024] The input audio asset is an audio asset that forms the basis for generating new audio assets using the above-described generative techniques. A generative model is trained on this asset. This method is typically used to generate new audio assets that are recognizable variations of the input audio asset. However, in other embodiments, this method may be used to generate new audio assets that are unrecognizably different from the input audio asset. The variation and degree of variation between the output asset and the input asset may be user-controllable. For example, at this step, the method may include receiving input control information for controlling the output of the method. To vary the degree of variation between the output audio file and the input audio file, one parameter in the input control information may be a variation value readable by a processor executing the method. In particular, the variation value may relate to one or more of the tone, frequency, duration, pitch, timbre, tempo, harshness, loudness, and brightness of the output audio. The input control information may also include a control sound file. This allows the generative model to generate new sounds that are distinctly influenced by the control sound file. For example, input audio files may all have a first tempo (or all files in a batch may have different tempos), and a control sound file with a second tempo (e.g., 100 bpm) may be input along with control information. This results in output sound files that are similar to the input files but all have the second tempo (e.g., 100 bpm). If the method is performed in a video game application, the input control information may be received from the video game environment or one or more events within the video game. The video game information may be branched into the input control information (or may constitute part or all of the input control information). Various control information described herein may be input to the system prior to training. For example, a generative model trained on a single image without control may be saved for later use.When the trained model is accessed to generate new images, control images may be input as input noise vectors or other components to influence the generation by the generative model.

[0025] Input audio assets can be selected, for example, from a predetermined sound library, to generate a particular output or set of outputs. Input audio assets can be selected, for example, from a predetermined library. Specific audio assets for input to the method can be selected based on one or more selection criteria. For example, a large audio asset library can be accessed, and a subset can be selected based on a set of requirements (e.g., desired ambiance, volume level, pitch, duration). Alternatively, input audio assets can be randomly generated based on a predetermined processing or generation method. Selection from a library or database of audio assets can be based on input from another process. In a video game application, an event can be triggered by an object or player in the video game environment. This event can output a signal that is received and used to select an input audio asset. In some implementations, the number of audio assets in the plurality of input audio assets can be controlled by a signal related to an event in the game.

[0026] In some embodiments, the input audio assets may be received directly from an external input, such as a microphone. The input audio assets may be received in real time, i.e., the input audio assets may be received and processed using an on-the-fly method to return output audio assets. Typically, each audio asset in the plurality of audio assets is different from the others. However, in some embodiments, one or more audio assets may be duplicate files.

[0027] Step 202: Convert the input audio asset into an input graphical representation.

[0028] As used herein, "graphical representation" refers to a visual form of data that records and characterizes the characteristics of each audio asset without losing the audio information. In other words, an audio asset is converted into a form in which the acoustic characteristics of the asset are recorded in visual form. Conversely, a conversion from visual form back to the original sound is also possible.

[0029] In this example, each input audio asset is converted into an input spectral image. The input spectral image provides a visual representation of audio information by plotting the frequency and amplitude of the audio against time. Figure 3 shows an example of the process for converting each input audio asset into an input spectral image. A mono audio sample 301 is an audio file with duration L and dimensions 1 x L. A conversion process is performed on the audio sample 301 to convert it from an acoustic form to a graphical representation (in this example, a single-channel grayscale spectrum). In this example, a short-time Fourier transform (STFT) is performed to convert the audio sample 301 into a log-magnitude (log-mag) spectral image 302 with dimensions 1 x 1 x h x w. This spectral image plots the frequency and amplitude of each frequency against the duration of the sound. Other Fourier forms and variations of the STFT may be used to obtain similar spectral images. This spectral image 302 can be converted back to an audio file by performing an inverse short-time Fourier transform (ISTFT) to obtain a reconstructed audio sample 303. When applying an inverse transform to obtain a reconstructed audio file, one can also apply, for example, the Griffin-Lim algorithm to reconstruct phase from an amplitude spectral image. This allows for an inverse Fourier transform to be performed on both phase and amplitude. Ideally, after the input 301 is converted to a spectral image 302 and then converted back to audio format 303, the input audio 301 and the reconstructed audio 303 are identical. That is, it is desirable that no audio information is lost when converting between acoustic and visual formats.

[0030] To obtain multiple input spectral images, the STFT technique, generally shown in Figure 3, is applied in step S202 to convert each input audio asset into a spectral image, although other transforms such as wavelet transforms may be used as well.

[0031] Step 203: Generate an input multi-channel image by stacking each input graphical representation into a separate channel of the image.

[0032] In this step, the multiple input spectral images obtained after steps S201 and S202 are combined into a multi-channel image. This step converts the batch 102 in FIG. 1 into a single image 103. In this example, each spectral image is single-channel, showing spectral features that display only either the amplitude or a composite representation of the Fourier transform. However, in other examples, the input spectral images may be multi-channel images (e.g., dual-channel images if both amplitude and phase are selected). The number of channels in the input audio file affects the number of channels in the spectral image.

[0033] The input spectral images may then be stacked together to form a multi-channel image, with each input spectral image assigned to a different channel within the multi-channel image. A simple example of an input batch of three input audio files can be illustrated using an RGB image. Each of the single-channel spectral images for the three input files can be placed in a separate channel. To generate an RGB (three-channel) image with each spectrogram stacked in a separate channel, the first spectrogram is assigned to the red channel, the second spectrogram to the green channel, and the third spectrogram to the blue channel. This concept can be extended to any number of spectral images stacked in any number of channels of a multi-channel image.

[0034] In this manner, a batch of input log-mag spectral images 102 is generated. In another embodiment, the audio files are converted into a spectral image format before being collected into the batch 102.

[0035] Step 204: Feed the input multi-channel image into the single-image generative model to train a generative model, and generate one or more output multi-channel images, each of which includes an output graphical representation.

[0036] The multi-channel image obtained after steps S201, S202, and S203 is fed to a generative model. This generative model is configured to generate a new multi-channel image that is a variation of the input multi-channel image. The generative model (or the neural network included therein) is trained on the multi-channel image. This model is typically configured to generate a multi-channel image that resembles the input spectral image. As mentioned above, the similarity or degree of change of the output spectral image to the input spectral image can be controlled via input control information received by the system. All input control information received in step S201 can be fed to this step through the generative model. This affects the performance and output of the generative model when generating the output spectral image.

[0037] The generative model used in this step is typically an image generation model. Such an image generation model typically comprises a generative adversarial network (GAN) with one or more generator neural networks and one or more discriminator networks. The discriminator network in such a model typically takes patch images generated by the generator network, identifies loss of structure at the scale of a small image within a larger image, convolves the entire image to distinguish whether each patch is real or fake, and averages all responses to provide the overall discriminator output. Examples of single-image generative models include SinGAN and CosSinGAN. Such generative models are particularly well-suited for the methods and systems described herein because they can only take a single image as training data and, once trained, can use the patch discriminator to generate images of any size. While the present invention can be generally described using single-image generative models, batch processing methods using other generative techniques for audio generation (e.g., variational autoencoders (VAEs), autoregressive models, and other neural network and GAN techniques) can also be applied.

[0038] If input control information is received in step S201 (or any other step), this information can be fed to the generative model to control aspects of the output image. The input control information can be fed to the generator of the GAN to influence how the generator generates images or patch images. For example, the input control information can first be converted into a noise vector. This noise vector is used as an input noise vector to the generator. Alternatively, or in combination, the input control information can be fed to a discriminator to influence how the loss is calculated (and / or output) at each step, for example. This technique can be applied to pre-trained networks and generative models by loading a trained model stored in memory and inputting the control information when generating new audio using the model.

[0039] Within each channel of this output multi-channel image, there is a spectral image (or other graphical representation of the audio asset). Any output multi-channel image obtained in this step can be sent to and stored in a memory unit for long-term storage or random access. The result of the method can be stored in a compressed file size until an audio format is needed to perform step S405 to obtain the audio file.

[0040] Step 205: Separate the output graphical representation from each output multi-channel image and convert each output graphical representation into an output audio asset.

[0041] In this step, a single-channel image is extracted from each channel of the multi-channel image to obtain multiple output multi-channel images. For example, if the output multi-channel image contains three channels, a simple one-to-one extraction results in three single-channel spectral images. If the input spectral image is a multi-channel spectral image, grayscale images from several channels of the output multi-channel image may be combined to form the output spectral image. For example, if the input spectral image is a two-channel spectral image, single-channel images from pairs of channels of the output multi-channel image may be combined to form the output spectral image.

[0042] While spectral images within a multi-channel image are advantageous, alternatively, only one spectral image may be represented using a single image, either as a single-channel grayscale spectral image or as a color spectral image spanning multiple channels (e.g., with different signal intensities represented across a color range). Stacking is therefore particularly advantageous, but not required.

[0043] Once the output spectral images are extracted, each spectral image can be converted into an audio file by performing an inverse transform (e.g., an inverse ISTFT or inverse wavelet transform). The inverse transform converts each of the output spectral images from a visual or graphical representation into an audio file. As a result, each spectral image is converted into a newly generated audio file. Each of these audio files can be sent to and stored in a memory unit for long-term storage or random access. Alternatively, each of these audio files can be sent to and processed by a processor for immediate use (e.g., playback in a video game environment).

[0044] In some embodiments, layered sounds can be generated in this step. The audio clips obtained by converting each output spectral image in this step are stacked on top of each other to generate a layered sound file. The audio clips can be simply stacked on top of each other so that they play simultaneously. Alternatively, in another example, all or some of the clips can be offset in time for staggered playback. The time difference between clips in such layered sounds can be variable or predetermined. For example, in the case of footsteps, the input sound asset (training sound) may include (i) the sound of a heel striking the ground, (ii) the sound of a toe striking the ground, and (iii) Foley sounds. A generative model trained on these input sound assets can output new heel, toe, and Foley sounds. These can be combined into a layered sound to generate a new overall footstep sound asset. Because the heel typically strikes the ground first, the layered sound can have the toe sound delayed from (but overlapping in terms of duration with) the heel sound. The same applies to Foley sounds.

[0045] Once trained, the generative model used may be stored and used "offline" to generate new sounds. While training a single-image generative model takes some time, once the model is trained, new sounds can be generated from this model very quickly. Thus, after a generative model is trained for a certain sound or sound type, it can be stored in memory (accessible for quickly generating new sounds similar to the trained sound). For example, a generative model can be trained with one or more training sounds for footsteps, according to the methods described above. This "footstep model" can then be stored and used, for example, in a video context. Each time a character in a video game moves around (e.g., in response to a user's input control), a new footstep sound can be generated from the model and video playback. This allows each character's footstep to sound slightly different. An event in the video game environment can trigger a signal to the generative model to generate a new sound of a certain type. Multiple different generative models may be stored in memory or within a processor to generate all the different types of sounds. When an "offline" generative model is used to generate layered sounds, the signal requesting the generation of a new sound may include information about the delay between the various sounds. For example, in the case of layered sounds of footsteps, the delay between the heel and toe sounds may be independent of the speed at which the character moves within the video game environment. Further data, such as video game data, may be sent to the trained generative model to influence the outcome as the generative model reacts to the situation.

[0046] A further aspect of the present disclosure provides a system. FIG. 4 shows a schematic representation of this system. The system 40 includes a memory, an asset input unit, an image generation unit, and an asset output unit. Each of the asset input unit, the image generation unit, and the asset output unit may be implemented in a single processor or in separate processors. Alternatively, these units may be located in separate remote memories and accessed (and processed) by a processor connected to the main memory. In this example, each unit is implemented in a processor 42.

[0047] The asset input unit 43 is configured to receive a plurality of input audio assets in the manner described in step S201. The image generation unit 44 is configured to receive an input multi-channel image from the asset input unit 43 and access a generative model to generate a new multi-channel image based on the input multi-channel image. In particular, the image generation unit 44 is configured to apply the generative model (typically a neural network-based machine learning model that trains on the input multi-channel image) to generate a new image in the manner described in step S203. The image generation unit 44 generates an output multi-channel image including an output graphical representation in each channel of the output multi-channel image. The asset output unit 45 is configured to receive the output multi-channel image from the image generation unit 44 and extract the output graphical representation in each channel. The asset output unit 45 is also configured to convert each of the extracted graphical representations into output audio assets to form a plurality of output audio assets.

[0048] In some embodiments, the system may further comprise a Fourier transform unit. The Fourier transform unit is accessed by either or both of the asset input unit 43 and the asset output unit and converts audio files to and from graphic files. The Fourier transform unit is configured to perform a Fourier transform operation (e.g., STFT) on the audio assets to convert the audio files into graphic representations (e.g., spectral images). The Fourier transform unit is also configured to perform an inverse Fourier transform operation (e.g., ISTFT) to convert the graphic representations (e.g., spectral images) into audio assets.

[0049] In some embodiments, the system may further include a video game data processing unit. The video game data processing unit is configured to process data extracted from (or related to) the video game environment and feed this data to at least one of the asset input unit, the image generation unit, and the asset output unit. In one example, the video game data processing unit generates video game information based on the virtual environment and sends this video game information to the image generation unit 44. The image generation unit executes the generative model using the video game information as one of the inputs (e.g., using the video game information as a conditional input to the generative model used). In another embodiment, the video game data processing unit simply receives video game information from a separate video game processor and sends this video game information to one or more other units in the system.

[0050] The foregoing discussion discloses and describes merely exemplary embodiments of the present invention. Those skilled in the art will recognize that the present invention can be embodied in other specific forms without departing from the spirit or essential characteristics thereof. Accordingly, the disclosure of the present invention is intended to be illustrative and not limiting of the scope of the present invention and the claims that follow. The present disclosure, including any identifiable variations of the above teachings, defines in part the scope of the claim terms. The subject matter of the invention is not dedicated to the public.

Claims

1. A method for generating an audio asset, comprising: receiving a plurality of input audio assets; converting each of the input audio assets into an input graphic representation; generating an input image from each of the input graphic representations; feeding the input images into a generation model for learning the generation model and generating one or more output images; wherein the output images include an output graphic representation; extracting the output graphic representation from each of the output images; and further comprising converting the output graphic representation into an output audio asset.

2. The method according to claim 1, wherein the input image is a multi-channel image generated by stacking each of the input graphic representations into individual channels of the image, and the generated output image is an output multi-channel image.

3. The step of converting each of the input audio assets into an input graphic representation includes: performing a Fourier transform on each audio asset; plotting the intensity in the frequency domain to generate a spectral image as the graphic representation.

4. The method according to claim 2, further comprising separating the output graphic representation from each of the output multi-channel images, and performing an inverse Fourier transform on each of the output graphic representations to obtain one or more output audio assets from each of the output graphic representations.

5. The method according to claim 1, wherein each of the output graphic representations is a spectral image.

6. The generation model is a single-image generation model comprising an adversarial generation network (GAN) with a generator and a patch discriminator, and the method according to claim 1 further comprises learning the GAN with the input images.

7. The method according to claim 2, wherein the output multi-channel image includes an output graphic representation therein.

8. The step of converting the output graphic representation into an output audio asset includes generating one or more layered output audio assets. ​ ​ ​ The method according to claim 1, wherein each of the layered output audio assets includes one or more audio assets extracted from the output graphic representation.

9. The method according to claim 8, wherein the audio assets within the layered output audio assets are temporally shifted so that a time delay occurs.

10. The step of receiving the plurality of input audio assets includes the step of receiving video game information from a video game environment, The method according to claim 1, wherein the step of feeding the input image into a generation model includes the step of feeding the video game information into a single image generation model such that the output image is affected by the video game information.

11. The method according to claim 1, further comprising the step of storing the learned generation model in a memory accessed to generate further audio assets.

12. A computer program comprising computer-executable instructions for causing a computer to execute the method according to claim 1.

13. A system for generating audio assets, an asset input unit, an image generation unit, and an asset output unit, wherein the asset input unit is configured to receive a plurality of input audio assets, convert each of the input audio assets into an input graphic representation, and generate an input image from each of the input graphic representations, the image generation unit implements a generation model for generating one or more output images based on the input images, the output image includes an output graphic representation, and the asset output unit is configured to separate the output graphic representation from the output image and convert the output graphic representation into an output audio asset.

14. The input image is a multi-channel image generated by stacking each of the input graphic representations into individual channels of the image using the asset output unit, The system according to claim 13, wherein the generated output image is an output multi-channel image.

15. Further comprising a conversion unit configured to perform a Fourier transform and an inverse Fourier transform to convert an audio file and a graphic file into each other. The asset input unit is configured to access the conversion unit to convert each of the input audio assets into an input graphic representation. The system according to claim 13 or 14, wherein the asset output unit is configured to access the conversion unit to convert each of the output graphic representations into an output audio asset. **Claim 16** Further comprising a video game data processing unit. The video game data processing unit processes video game information extracted from or related to a video game environment. The video game information is configured to be fed to at least one of the asset input unit, the image generation unit, and the asset output unit. The system according to claim 13, wherein the image generation unit is configured to execute the generation model at least partially based on the video game information.