Sound file generation device and method for generating sound file

The audio file generation device and method automate the creation of sound files by using AI or templates to process sound source files, addressing the inefficiency of manual sound editing and ensuring high-quality, metadata-tagged sound effects.

WO2026069674A1PCT designated stage Publication Date: 2026-04-02SONY INTERACTIVE ENTERTAINMENT LLC
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

The manual process of creating sound effects from sound source files is time-consuming and labor-intensive, especially when dealing with large numbers of sounds, requiring significant effort from sound creators.

Method used

An audio file generation device and method that utilizes an AI asset generation model or sound waveform templates to automatically extract and process sound data from sound source files, applying fade processes based on sound type to generate sound files, optionally with metadata, reducing manual effort.

Benefits of technology

Automated generation of sound files significantly reduces the time and effort required, enabling efficient production of high-quality sound effects with natural transitions and accurate metadata tagging.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024034971_02042026_PF_FP_ABST
    Figure JP2024034971_02042026_PF_FP_ABST
Patent Text Reader

Abstract

A sound source file acquisition unit 12 acquires a sound source file in which sound is recorded. A sound file generation unit 14 generates a sound file including one sound from the acquired sound source file.
Need to check novelty before this filing date? Find Prior Art

Description

Sound File Generation Device and Sound File Generation Method

[0001] The present disclosure relates to a technique for generating sound files of sound effects used in games and the like.

[0002] In games, various types of sound effects are used to enhance the sense of immersion. For example, by outputting sound effects such as the footsteps and movement sounds of player characters, the voices of animals, gunshots, etc. at appropriate timings according to the game scenes, the sense of immersion can be enhanced. Currently, sound effects are produced by various methods, but by recording the sounds actually emitted in a studio or natural environment and using them as sound effects, real-world sounds can be provided to users.

[0003] For example, considering the sound effect of footsteps, the footsteps when walking on grass are different from those when walking through a puddle, and the footsteps when walking wearing sandals are also different from those when walking wearing leather shoes. When producing sound effects in a studio, a sound creator creates an environment similar to grass or a puddle on the studio floor, steps on it many times wearing sandals or leather shoes, and records the footsteps in various situations. The sound creator may also actually walk on the grass or puddle in the natural environment and record the footsteps. After recording multiple footsteps in one environment, the sound creator performs an editing operation of manually cutting out the sound data of the footsteps one by one from the source sound file to generate multiple footstep files to be used as sound effects.

[0004] FIG. 1(a) shows an example of a source sound file in which multiple footsteps are recorded. In this source sound file, the footsteps when a sound creator wearing leather shoes stepped on the puddle created on the studio floor six times are recorded. In an actual source sound file, the footsteps when stepping on it dozens of times may be recorded. The sound creator plays the source sound file on an editing device (personal computer) and manually cuts out multiple footsteps while listening to the output footsteps.

[0005] Figure 1(b) shows the six footstep data extracted from the sound source file. In the example shown in Figure 1(b), sound 7 is included between footstep 4 and footstep 5, but the sound creator listened to sound 7 and confirmed that it was noise and not a footstep, so sound 7 was not extracted as a footstep.

[0006] The sound creator applies a fade-in process to the beginning and end of each extracted footstep data and saves them as separate footstep files. As part of the fade-in process, the sound creator applies a fade-in process to the beginning of the footstep data so that the volume level at the beginning starts at a minimum value (e.g., zero), and applies a fade-out process to the end of the footstep data so that the volume level at the end of the footstep data ends at a minimum value (e.g., zero). At this time, metadata indicating the type of sound, for example, information indicating that it is the sound of footsteps walking through a puddle while wearing hard shoes, may be added to each footstep file. These multiple footstep files may be randomly combined in chronological order in a game scene in which a player character wearing hard shoes is walking through a puddle and used as a sound effect for footsteps.

[0007] As described above, the editing process of creating audio files from sound source files is performed manually, which requires a great deal of time and effort. In particular, if the sound source file contains a large number of sounds, the time and effort required becomes enormous. Therefore, this disclosure aims to provide a technology for automatically generating audio files from sound source files in order to reduce the burden on sound creators.

[0008] Certain aspects of this disclosure relate to an audio file generation device. The audio file generation device comprises an audio source file acquisition unit that acquires an audio source file on which sound is recorded, and an audio file generation unit that generates an audio file containing one sound from the acquired audio source file.

[0009] Another aspect of this disclosure relates to a method for generating an audio file. The method for generating an audio file comprises the steps of: acquiring an audio source file on which sound is recorded; and generating an audio file containing one sound from the acquired audio source file.

[0010] Furthermore, any combination of the above components, as well as any conversion of the expressions of this disclosure between methods, apparatus, systems, recording media, computer programs, etc., are also valid forms of this disclosure.

[0011] This figure shows an example of an audio file containing multiple footsteps. It also shows the functional blocks of the sound file generation device. Another example of an audio file is shown. Finally, an example of sound data extracted from an audio file is shown. The image also shows an example of the sound file editing screen.

[0012] Figure 2 shows the functional blocks of the sound file generation device 1, which automatically generates sound files containing one sound from a sound source file on which sounds are recorded. The sound source file 2 contains multiple sounds to be extracted as sound effects, recorded at different timings, and multiple sounds of the same type are recorded without temporal overlap. For example, the sound source file 2 may contain multiple footsteps. In another example, the sound source file 2 may contain multiple different types of sounds recorded at different timings, for example, footsteps and gunshots may be recorded without temporal overlap. The processing device 10 has the function of automatically generating multiple sound files containing one sound by cutting out sections from the sound source file 2 where the waveform of the sound to be extracted is recorded for each sound.

[0013] Figure 3 shows an example of sound source file 2. Sound source file 2 is created by recording sounds actually emitted in a studio or natural environment. Sound source file 2 may be created in a single recording session. The following explanation describes a case where multiple footsteps are recorded in sound source file 2, but the types of sounds recorded may include animal noises, gunshots, clapping, etc.

[0014] The sound file generation device 1 comprises a processing unit 10, a recording device 20, an input device 30, and an output device 32. The processing unit 10 has a function to automatically generate sound effect files and comprises a sound source file acquisition unit 12, a sound file generation unit 14, and an editing processing unit 16. The recording device 20 is a large-capacity recording device such as an HDD (hard disk drive) or SSD (solid state drive), and may be an internal recording device or an external recording device connected to the processing unit 10 by USB (Universal Serial Bus), etc. The recording device 20 records an asset generation model 22, which is a machine learning model. The recording device 20 also has a sound file recording unit 24 that records sound effect files generated by the processing unit 10. The input device 30 is an input interface operated by a sound creator and may include a keyboard or mouse. The output device 32 includes a display device that displays an editing screen for editing the sound effect files generated by the processing unit 10, and a speaker that plays back the sound effect files generated by the processing unit 10.

[0015] The functions of the components in the processing unit 10 may be realized in a circuit or processing circuitry, including a general-purpose processor, an application-specific processor, an integrated circuit, an ASIC (Application Specific Integrated Circuit), a CPU (a Central Processing Unit), a conventional circuit, and / or a combination thereof, configured or programmed to realize the functions described herein. A processor is considered to be a circuit or processing circuitry that includes transistors and other circuits. A processor may also be a programmed processor that executes a program stored in memory.

[0016] In this specification, circuits, units, and means are hardware programmed to perform or execute the functions described herein. Such hardware may be any hardware disclosed herein, or any hardware known to be programmed to perform or execute the functions described herein.

[0017] If the hardware is a processor that is considered to be of the type of circuit, the circuit, means, or unit may be a combination of hardware and software used to constitute the hardware and / or processor.

[0018] In the processing unit 10, the sound source file acquisition unit 12 acquires the sound source file 2, and the sound file generation unit 14 generates multiple sound files, each containing one sound, from the acquired sound source file 2. The sound source file 2 shown in Figure 3 records seven footsteps, and therefore the sound file generation unit 14 generates seven footstep files, each containing one footstep, from the sound source file 2.

[0019] To automatically generate multiple sound files from sound source file 2, the sound file generation device 1 may be equipped with AI (artificial intelligence) that has an automatic sound file generation function. The sound file generation device 1 of this embodiment includes an asset generation model 22 that has been trained by machine learning on sound files that have been manually generated by a sound creator in the past. The asset generation model 22 is a trained model that has been trained to extract regularities (features) of sound files by machine learning on multiple sound files. The asset generation model 22 may perform machine learning using training sound files and information indicating the type of the sound file (type information) as training data.

[0020] The type information for learning sound files may be defined by one or more parameters. The type information consists of a major category and a minor category, with the major category parameter defining the "type of sound." For example, if the sound file contains footstep data, the major category parameter is "footsteps," and if the sound file contains gunshot data, the major category parameter is "gunshot." The minor category parameters are used to subdivide the type of sound defined in the major category, and for example, they define information indicating the circumstances under which the sound was emitted. For example, the minor category parameters for footsteps may include "type of ground" and "hardness of shoes," and the minor category parameters for gunshots may include "type of gun." Thus, the content and number of minor category parameters may be determined as appropriate according to the major category parameters. Note that the type information must include the major category parameters, but it does not have to include the minor category parameters.

[0021] Examples of footstep sound file type information are shown below in the order of "major category," "minor category," and "minor category." ・Footsteps, gravel, hard shoes ・Footsteps, grass, hard shoes ・Footsteps, puddle, hard shoes ・Footsteps, gravel, soft shoes ・Footsteps, grass, soft shoes ・Footsteps, puddle, soft shoes In this way, footstep sound file type information may be defined by a parameter indicating the major category, "type of sound," a parameter indicating the minor category, "type of ground," and a parameter indicating the "hardness of shoes." Note that these parameters are examples, and for example, "type of ground" could be concrete, soil, stone, etc. Also, instead of "hardness of shoes," "type of shoes" may be defined as a minor category parameter. In this case, "type of shoes" may include sandals, leather shoes, sneakers, military boots, etc.

[0022] Figure 1(b) shows multiple footstep data manually extracted from a sound source file by a sound creator. The sound creator applies a fade-in process to the beginning of the extracted footstep data so that the volume level at the beginning of the footstep data starts at a minimum value (e.g., zero), and a fade-out process to the end of the footstep data so that the volume level at the end of the footstep data ends at a minimum value (e.g., zero), thereby generating a footstep file. By applying a fade process to the extracted footstep data, when multiple footstep files are played sequentially in a game, they will sound like different footsteps to the game player without any sense of incongruity.

[0023] To make the footsteps sound natural, the sound creator sets different fade times depending on the type of footstep. For example, the fade time for footsteps walking on gravel, footsteps walking on grass, and footsteps walking through puddles are all different. It is also common for the fade times for footsteps, gunshots, and animal sounds to be set differently. The asset generation model 22 is trained to extract regularities (features) of sound waveforms corresponding to the type of sound file by performing machine learning using training sound files manually generated by the sound creator and type information indicating the type of sound file as training data. Therefore, the asset generation model 22 understands the fade-in and fade-out features corresponding to the type of sound file.

[0024] The asset generation model 22 may be trained to, upon input of a sound source file 2, extract each sound data contained in the sound source file 2, identify the type of extracted sound data, apply a fade process according to the type of sound data, and output an audio file containing the faded sound data and information indicating the type of sound data as metadata. If the sound source file 2 contains N sound data, the asset generation model 22 is configured to output N audio files upon input of the sound source file 2. As will be described later, the audio file generation unit 14 may have the function of extracting each sound data contained in the sound source file 2, identifying the type of extracted sound data, applying a fade process according to the type of sound data, and generating an audio file containing the faded sound data and information indicating the type of sound data as metadata, without using the AI ​​asset generation model 22.

[0025] Figure 4 shows an example of sound data extracted from sound source file 2. When the sound file generation unit 14 receives sound source file 2 from sound source file acquisition unit 12, it inputs sound source file 2 to the asset generation model 22. The asset generation model 22 extracts footsteps a to footsteps g from sound source file 2, processes the extracted footsteps data, and generates multiple footsteps files. The asset generation model 22 identifies the type of footstep from the footstep waveform data, and processes the footsteps data by applying fade-in and fade-out processing to the beginning and end of the footsteps data according to the identified type of footstep. The asset generation model 22 generates a footsteps file containing the processed footsteps data and information indicating the type of footstep (type information) as metadata, and outputs it to the sound file generation unit 14. The sound file generation unit 14 may also add information indicating the type of footstep (type information) as metadata to the footsteps file output by the asset generation model 22.

[0026] In the example shown in Figure 4, the asset generation model 22 does not extract the waveforms of sound i, sound j, and sound k as sound data to be used as sound effects. This means that the asset generation model 22 has never learned the waveform data of sound i, sound j, and sound k in the past and does not recognize them as sound data (assets) that should be extracted. For example, if the asset generation model 22 has learned the sound data of applause and detects that the waveforms of sound i, sound j, and sound k are applause sound data, the asset generation model 22 may extract sound i, sound j, and sound k as applause sound data, apply a fade process corresponding to the applause sound data, and generate three applause sound files.

[0027] In the embodiments described above, the asset generation model 22 is a trained model that has learned from multiple types of sound data and has the function of identifying the type of sound contained in the input sound source file 2. In another embodiment, the asset generation model 22 may be a trained model that has learned from each type of sound. For example, for an asset generation model 22 that learns footsteps, an asset generation model 22 may be created for each combination of "type of ground" and "hardness of shoes". In this case, an asset generation model 22 may be created for each combination of gravel, hard shoes; grass, hard shoes; puddle, hard shoes; gravel, soft shoes; grass, soft shoes; puddle, soft shoes, meaning that six asset generation models 22 may be created.

[0028] In this alternative embodiment, if the sound source file 2 contains a recording of footsteps made while walking through a puddle with hard shoes, the sound file generation unit 14 generates a footstep file using an asset generation model 22 that has learned the sound of footsteps made while walking through a puddle with hard shoes. In many cases, the sound creator knows the environment in which the footsteps contained in the sound source file 2 were recorded, so they can specify the asset generation model 22 corresponding to the footsteps and have the sound file generation unit 14 automatically generate the sound file. By training the asset generation model 22 for each type of sound in this way, if the type of sound contained in the sound source file 2 is known in advance, a dedicated asset generation model 22 can be used, making it possible to generate footstep files with high accuracy.

[0029] The above describes a case in which the sound file generation unit 14 automatically generates sound files using the AI ​​asset generation model 22. However, it is also possible to automatically generate sound files without using the asset generation model 22. Below, we will describe a method by which the sound file generation unit 14 automatically generates sound files without using AI.

[0030] In this case, the sound file generation unit 14 may maintain one or more representative sound waveform templates and fade times for each type of sound. The sound file generation unit 14 identifies sections of sound waveforms with amplitudes exceeding a predetermined threshold in the sound source file 2 and extracts sound data from the sound source file 2. The predetermined threshold is set to an amplitude value that can distinguish between sounds to be extracted as sound effects and other noise (for example, white noise), and the sound file generation unit 14 can extract some sound data (at this point, it is not known what kind of sound it is) by identifying sections of sound waveforms with amplitudes exceeding the predetermined threshold.

[0031] The sound file generation unit 14 identifies the type of sound by searching for a sound waveform template that matches or approximates the extracted sound waveform. For example, the sound file generation unit 14 may perform pattern matching between the extracted sound waveform and multiple sound waveform templates to identify the matching or most similar sound waveform template. For example, if the extracted sound waveform matches or approximates the sound waveform template of footsteps when walking on grass in hard shoes, the sound file generation unit 14 identifies that the extracted sound waveform represents the sound of footsteps when walking on grass in hard shoes.

[0032] The sound file generation unit 14 then applies a fade process to the sound data to generate a sound file. At this time, the sound file generation unit 14 may apply a fade process according to the type of sound to the beginning and end portions of the extracted sound data to generate the sound file. Alternatively, the sound file generation unit 14 may add a fade time according to the type of sound to the beginning of the extracted sound data and to the end of the extracted sound data, respectively, and generate a sound file with a fade-in process applied to the beginning fade time and the fade-in process applied to the end fade time. By applying a fade process according to the type of sound, when the sound file is played, the game player will be able to hear it as a natural sound effect.

[0033] Furthermore, if the type of sound contained in the sound source file 2 is known, the sound file generation unit 14 identifies the section of the sound waveform with an amplitude exceeding a predetermined threshold, extracts the sound data from the sound source file 2, and then identifies the type of sound without the need for pattern matching. Therefore, the sound file generation unit 14 can generate a sound file by applying a fade process according to the type of sound.

[0034] As described above, the sound file generation unit 14 automatically generates multiple sound files from a single sound source file 2. The sound file generation unit 14 then records the generated sound files in the sound file recording unit 24.

[0035] The editing processing unit 16 supports editing of automatically generated sound files. After multiple sound files have been generated, it is preferable that the editing processing unit 16 plays the multiple sound files and outputs sound effects from the output device 32, allowing the sound creator to check the quality of the sound files.

[0036] Figure 5 shows an example of an audio file editing screen. The editing processing unit 16 displays an editing screen on the output device 32 for editing multiple audio files generated by the audio file generation unit 14. On the editing screen, a list of multiple audio files generated by the audio file generation unit 14 is displayed. When the sound creator selects an audio file by operating the input device 30, the sound waveform of the selected audio file is displayed, and playback of the sound effect begins. In this example, the footsteps file d is selected, and the editing processing unit 16 displays a progress bar indicating the progress of playback.

[0037] Below the sound waveform, editing tools are displayed for editing amplitude, fade-in, and fade-out. After listening to the automatically generated sound effect, the sound creator can adjust various parameters by operating the input device 30 and moving the pointers shown in the editing tools if they determine that adjustments are necessary. In this example, the sound creator can increase the volume by moving the pointer shown in the amplitude editing tool to the right and decrease the volume by moving it to the left. The sound creator can also lengthen the fade time by moving the pointers shown in the fade-in and fade-out editing tools to the right and shorten the fade time by moving them to the left.

[0038] The present disclosure has been described above based on embodiments. These embodiments are illustrative, and it will be understood by those skilled in the art that various modifications are possible in combinations of their components and processing processes, and that such modifications are also within the scope of the present disclosure. In the embodiments, it was explained that multiple sounds are recorded in the sound source file 2 at different timings, but the sound source file 2 may contain only one sound.

[0039] This disclosure may include the following embodiments: [Item 1] An information processing device comprising a circuit configured to, wherein the circuit acquires a sound source file on which sound is recorded, and generates a sound file containing one sound from the acquired sound source file. [Item 2] The sound file generation device according to Item 1, wherein the circuit extracts sound data containing one sound from the sound source file. [Item 3] The sound file generation device according to Item 2, wherein the circuit processes the extracted sound data to generate a sound file. [Item 4] The sound file generation device according to Item 3, wherein the circuit applies a fade process to the sound data to generate a sound file. [Item 5] The sound file generation device according to Item 4, wherein the circuit applies a fade process to the beginning and end portions of the sound data according to the type of sound to generate a sound file. [Item 6] The sound file generation device according to Item 5, wherein the circuit includes information indicating the type of sound as metadata for the sound file in the sound file. [Item 7] The sound file generation device described in Item 1, wherein the sound source file contains multiple sounds of the same or different types recorded at different timings. [Item 8] The sound file generation device described in Item 7, wherein the circuit extracts multiple sound data containing one sound from the sound source file, processes the extracted multiple sound data, and generates multiple sound files. [Item 9] The sound file generation device described in Item 1, wherein the circuit inputs the sound source file into a trained model created by machine learning using multiple manually generated sound files as training data. [Item 10] A method for generating multiple sound files, comprising acquiring a sound source file in which multiple sounds are recorded at different timings, and generating multiple sound files containing one sound from the acquired sound source file.[Item 11] A recording medium that records a program executed on a computer, wherein the program enables the computer to perform the following functions: acquire a sound source file in which multiple sounds are recorded at different timings, and generate multiple sound files, each containing one sound, from the acquired sound source file.

[0040] This disclosure can be used in the technical field of generating an audio file containing one audio data from an audio source file.

[0041] 1...Sound file generation device, 2...Sound source file, 10...Processing device, 12...Sound source file acquisition unit, 14...Sound file generation unit, 16...Editing processing unit, 20...Recording device, 22...Asset generation model, 24...Sound file recording unit, 30...Input device, 32...Output device.

Claims

1. A sound file generation device comprising: a sound source file acquisition unit that acquires a sound source file on which sound is recorded; and a sound file generation unit that generates a sound file containing one sound from the acquired sound source file.

2. The sound file generation device according to claim 1, characterized in that the sound file generation unit extracts sound data containing one sound from a sound source file.

3. The sound file generation device according to claim 2, characterized in that the sound file generation unit processes the extracted sound data to generate a sound file.

4. The sound file generation device according to claim 3, characterized in that the sound file generation unit generates a sound file by applying a fade process to sound data.

5. The sound file generation device according to claim 4, characterized in that the sound file generation unit generates a sound file by applying a fade process to the beginning and end portions of the sound data according to the type of sound.

6. The sound file generation device according to claim 5, characterized in that the sound file generation unit includes information indicating the type of sound as metadata for the sound file.

7. The sound file generating device according to claim 1, characterized in that the sound source file contains multiple sounds of the same or different types recorded at different timings.

8. The sound file generation device according to claim 7, characterized in that the sound file generation unit extracts multiple sound data, each containing one sound, from a sound source file, processes the extracted multiple sound data, and generates multiple sound files.

9. The sound file generation device according to claim 1, characterized in that the sound file generation unit generates multiple sound files by inputting sound source files into a trained model created by machine learning using multiple sound files previously generated manually as training data.

10. A method for generating an audio file, comprising the steps of: obtaining an audio source file on which sound is recorded; and generating an audio file containing one sound from the obtained audio source file.

11. A program to implement the following functions for a computer: acquiring sound source files containing recorded sounds, and generating a single sound file from the acquired sound source files.

Citation Information

Patent Citations

  • Digital data editing apparatus and method therefor

    JP1999016332A

  • Information processing apparatus, sound material segmentation method, and program

    JP2010134231A

  • Information processing device, information processing method, and information processing program

    WO2023032270A1

  • Information processing device, information processing method, program, and information processing system

    WO2023127422A1

  • Information processing device, information processing method, and program

    WO2023218993A1