Sound file generation device and sound file generation method
The sound file generation device automates the separation and metadata assignment of ambient and sound effect data, addressing the inefficiency of manual sound source location setting in games, enhancing immersion through automated audio file creation.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SONY INTERACTIVE ENTERTAINMENT LLC
- Filing Date
- 2024-10-16
- Publication Date
- 2026-04-23
Smart Images

Figure JP2024036762_23042026_PF_FP_ABST
Abstract
Description
Sound File Generation Device and Sound File Generation Method
[0001] The present disclosure relates to a technique for generating sound files such as sound effects used in games and the like.
[0002] A surround sound space with speakers arranged to surround the user provides the user with a sense of the spread of the sound field, as well as the three-dimensional and immersive feeling of the sound. 5.1-channel and 7.1-channel audio systems are widespread, but in recent years, in order to improve the three-dimensional feeling of the three-dimensional audio space, an audio system that installs speakers on the ceiling or the like and emits sound from above has also been developed. Patent Document 1 discloses a technique for localizing the sound image of an agent sound at a second localization position different from the first localization position where the sound image of a game sound is localized.
[0003] Conventionally, virtual surround technology has been developed to realize a virtual multi-channel audio space using a small number of speakers. In virtual surround technology, surround sound is pseudo-reproduced by localizing virtual sound sources. Also, in a video game system, in order to enhance the immersion in the game, a technique for estimating the head-related transfer function (HRTF) of the user and providing optimized stereo sound to the user through headphones has been put into practical use.
[0004] Japanese Unexamined Patent Application Publication No. 2019-208185
[0005] During game play by the user, various types of environmental sounds are output. Environmental sound is the sound heard from the surrounding space, and in games, for example, various environmental sounds such as the sound of a forest, the sound of water, the sound of a big city, the sound of a park, the sound of rain, and the sound of waves are prepared. By using the raw sound recorded in a natural environment as the environmental sound, a real sound of the real world can be provided to the user. By outputting appropriate environmental sounds at a timing according to the game scene during game play by the user, the immersion in the game can be enhanced.
[0006] In object-based spatial audio, sound sources are given metadata such as location information, and the playback device calculates in real time what kind of sound to output from multiple speakers (panning information) based on the location information of the sound source, thereby outputting three-dimensional sound. During gameplay, if a bird sings above the head of the player character controlled by the user, the game device uses the location information of the bird (the sound source) to localize the sound source above the user, so that the bird's song is heard from above. Similarly, if a car honks its horn near the player character, the game device uses the location information of the car (the sound source) to localize the sound source at approximately the same height as the user (for example, the user's ears), so that the horn sound is heard from around the user. In this way, in games, localizing the sound images of sounds emitted by various objects to the appropriate positions can enhance immersion in the game.
[0007] To implement object-based audio, game programmers need to set the location information of sound sources as metadata in the sound files. However, this setting process is done manually and requires a great deal of time and effort. Therefore, this disclosure aims to provide a technology that automatically sets the location information of sound sources in sound files in order to reduce the burden on game programmers.
[0008] A sound file generation device according to one aspect of the present disclosure includes a sound data acquisition unit that acquires recording data of ambient sounds, a sound source separation unit that separates the recording data into ambient sound data and sound effect data, and a sound file generation unit that identifies location information corresponding to the separated sound effects and generates a sound effect file including the sound effect data and location information.
[0009] Another embodiment of the present disclosure of a sound file generation method includes the steps of: acquiring recording data of ambient sounds; separating the recording data into ambient sound data and sound effect data; identifying location information corresponding to the separated sound effects; and generating a sound effect file including the sound effect data and location information.
[0010] Furthermore, any combination of the above components, as well as any conversion of the expressions of this disclosure between methods, apparatus, systems, recording media, computer programs, etc., are also valid as aspects of this disclosure.
[0011] This diagram shows the functional blocks of the sound file generation device. This diagram shows an example of recorded data. This diagram shows a flowchart of the sound file generation procedure. (a) shows an example of separated ambient sound data, and (b) shows an example of separated sound effect data. This diagram shows an example of ambient sound data for a predetermined time. This diagram shows an example of extracted sound effect data. This diagram shows the correspondence relationships held in the position information holding unit. This diagram shows an example of the user's surround sound space. This diagram shows the hardware configuration of the information processing device. This diagram shows the functional blocks of the information processing device. This diagram shows an example of a displayed game image.
[0012] <Sound File Generation Function> Figure 1 shows the functional block of a sound file generation device that automatically generates sound files from recorded data containing ambient sounds. The recorded data 104 was recorded outdoors by a sound creator and contains ambient sounds. The sound file generation device 100 of this embodiment generates ambient sound files from the recorded data 104 for incorporation into game software.
[0013] The recorded data 104 contains not only ambient sounds but also sounds other than ambient sounds. The sound file generation device 100 has a function to separate sound data by source, and separates the recorded data 104 into sound data (sound signals) of ambient sounds and sound data (sound signals) of sounds other than ambient sounds, and generates separate sound files for each. In this embodiment, the sound file generation device 100 does not discard the sound data of sounds other than ambient sounds, but effectively utilizes it as sound data for sound effects superimposed on the ambient sounds.
[0014] For example, when recording ambient sounds in a forest, the sounds of birds, frogs, and other animals may also be recorded. In this embodiment, the recording data 104 may include ambient sounds recorded in the forest and animal sounds. Although animal sounds are sometimes included and referred to as forest ambient sounds, in this embodiment, for the following reasons, animal sounds are treated as sounds different from forest ambient sounds.
[0015] In game software that uses many types of ambient sounds, the number of ambient sound files becomes enormous. Therefore, a limit is placed on the data size of the ambient sound files, and the duration of each ambient sound data is set to be as short as possible. As a result, the user's game device outputs ambient sounds by looping (repeating) one ambient sound data. In this case, if a prominent (distinctive) sound remains in the ambient sound data, that prominent sound will be output periodically. For example, if the ambient sound data for a forest has a duration of 30 seconds, and the distinctive calls of birds or frogs remain in the ambient sound data, those calls will be output every 30 seconds. If the user notices that the ambient sounds are being looped, it may impair their immersion in the game, which is undesirable.
[0016] Therefore, the sound file generation device 100 of this embodiment separates the recorded data 104 into ambient sound data and sound data such as bird and frog calls, so that no distinctive or noticeable sounds remain in the ambient sound data. The sounds to be removed from the ambient sound are not limited to animal calls, but also include sounds that are emitted suddenly at a relatively loud volume, such as car horns and human shouts. For example, if the recorded data 104 is sound data that records ambient sounds underwater, and the recorded data 104 contains dolphin calls, the sound file generation device 100 will not treat the dolphin calls as part of the ambient sound, but will treat the dolphin calls as sounds that should be removed from the ambient sound.
[0017] Figure 2 shows an example of recording data 104. Recording data 104 is created by recording sounds in a natural environment. In the following, we will explain the case in which recording data 104 includes forest ambient sounds and bird calls, but sounds other than forest ambient sounds, such as frog croaks or car horns, may also be recorded. In recording data 104, forest ambient sounds are recorded continuously, but bird calls are recorded at the moment a bird sings during the recording.
[0018] The sound file generation device 100 comprises a processing device 102 and a sound file recording device 120. The processing device 102 has the function of separating the sound source from the recorded data 104 and automatically generating ambient sound files and sound effect files, and comprises a sound data acquisition unit 110, a sound source separation unit 112, and a sound file generation unit 114. The sound file generation unit 114 has a first generation unit 116 that generates ambient sound files and a second generation unit 118 that generates sound effect files. The sound file recording device 120 comprises an ambient sound file recording unit 122 that records ambient sound files generated by the first generation unit 116, a sound effect file recording unit 124 that records sound effect files generated by the second generation unit 118, and a location information holding unit 126 that holds location information of the sound source. The location information holding unit 126 holds the type of sound information and the location information of the sound source that emits the sound in association. The audio file recording device 120 is a high-capacity recording device such as an HDD (hard disk drive) or an SSD (solid state drive), and may be an internal recording device or an external recording device connected to the processing unit 102 by USB (Universal Serial Bus) or the like.
[0019] The functions of the components in the processing unit 102 may be realized in a circuit or processing circuitry, including a general-purpose processor, an application-specific processor, an integrated circuit, an ASIC (Application Specific Integrated Circuit), a CPU (a Central Processing Unit), a conventional circuit, and / or a combination thereof, configured or programmed to realize the functions described herein. A processor is considered to be a circuit or processing circuitry that includes transistors and other circuits. A processor may also be a programmed processor that executes a program stored in memory.
[0020] In this specification, circuits, units, and means are hardware programmed to perform or execute the functions described herein. Such hardware may be any hardware disclosed herein, or any hardware known to be programmed to perform or execute the functions described herein.
[0021] If the hardware is a processor that is considered to be of the type of circuit, the circuit, means, or unit may be a combination of hardware and software used to constitute the hardware and / or processor.
[0022] Figure 3 shows a flowchart of the sound file generation procedure in the embodiment. When the sound data acquisition unit 110 acquires recording data 104 input from an external source (S10), it provides it to the sound source separation unit 112. The sound source separation unit 112 has the function of separating the recording data 104, which is a mixture of multiple sound sources, and separates the recording data 104 into sound data of ambient sounds and sound data of sounds other than ambient sounds, and identifies information indicating the type of separated sound (type information). In the embodiment, the sound file generation device 100 extracts sounds other than ambient sounds as sound effects, so the sound source separation unit 112 separates the recording data 104 into ambient sound data and sound effect data, and identifies the type information of ambient sounds and the type information of sound effects (S12).
[0023] Figure 4(a) shows an example of separated ambient sound data, and Figure 4(b) shows an example of separated sound effect data. The sound source separation unit 112 may separate the recording data 104 using a machine learning-trained sound source separation model. This sound source separation model is a trained model that has been trained to extract regularities (features) of ambient sound data and sound effect data by machine learning from multiple ambient sound data and sound effect data.
[0024] The sound source separation model may be a neural network trained to extract regularities (features) of sound signals corresponding to the types of ambient sounds and sound effects by performing machine learning using training recording data, ambient sound data and sound effect data separated from the recording data, and type information of the ambient sound data and sound effect data as training data. When the sound source separation model is input to the recording data 104, it is trained to output ambient sound data and sound effect data separated from the recording data 104, as well as type information of the ambient sound and type information of the sound effect. If multiple types of sound effects are recorded in the recording data 104, the sound source separation model outputs multiple types of sound effect data and type information of each sound effect.
[0025] The sound source separation unit 112 inputs the recording data 104 into the sound source separation model and receives the output of the sound source separation model. As a result, the sound source separation unit 112 receives the ambient sound data shown in Figure 4(a) and the sound effect data shown in Figure 4(b), as well as information indicating that the ambient sound is a forest ambient sound and information indicating that the sound effect is a bird's song. The sound source separation unit 112 supplies this sound data and type information to the sound file generation unit 114.
[0026] In the above example, the sound source separation model has the function of separating the recording data 104 into sound sources and simultaneously identifying the type of each sound source. However, in another example, the sound source separation model may consist of a first model that separates the recording data 104 into sound sources and a second model that identifies the type of separated sound source. When this two-stage model configuration is adopted, there is the advantage that an existing machine learning model (AI model) can be used for the first model. The second model may be an AI classification model that has been trained to output sound type information when sound data (ambient sound data or sound effect data) is input. In this way, the sound source separation unit 112 may separate the recording data 104 into ambient sound data and sound effect data and identify the type of each sound using this two-stage model configuration.
[0027] In the sound file generation unit 114, the first generation unit 116 generates an ambient sound file containing the separated ambient sound data and records the generated ambient sound file in the ambient sound file recording unit 122. The second generation unit 118 generates a sound effect file containing the separated sound effect data and records the generated sound effect file in the sound effect file recording unit 124.
[0028] <Generating ambient sound files> First, the process for generating ambient sound files will be explained. As mentioned above, there are limitations on the data size of ambient sound files, and the duration of ambient sound data must be set as short as possible. The first generation unit 116 extracts ambient sound data of a predetermined duration from the ambient sound data separated from the recording data 104 in order to reduce the data size.
[0029] Figure 5 shows an example of ambient sound data for a predetermined duration that is extracted. In this example, the first generation unit 116 extracts 30 seconds of ambient sound data. Note that the longer the ambient sound data, the less likely the user is to notice that it is being looped when it is played back on a game device. Therefore, if the data size constraints allow, the first generation unit 116 may extract ambient sound data longer than 30 seconds.
[0030] The first generation unit 116 generates an ambient sound file by applying a fade process to the extracted ambient sound data (S14). Specifically, the first generation unit 116 performs a fade-in process on the beginning of the ambient sound data so that the volume level at the beginning starts at a minimum value (for example, zero), and performs a fade-out process on the end of the ambient sound data so that the volume level at the end ends at a minimum value (for example, zero). By applying a fade process to the ambient sound data, the first generation unit 116 ensures that the ambient sound data is not perceived as unnatural by the user when it is looped on the user's game device.
[0031] The first generation unit 116 also includes information indicating the type of ambient sound as metadata in the ambient sound file. In this embodiment, since the ambient sound is the sound of a forest, the first generation unit 116 may include information indicating that the ambient sound is the sound of a forest as metadata in the ambient sound file. The first generation unit 116 records the generated ambient sound file in the ambient sound file recording unit 122.
[0032] <Sound effect file generation process> Next, the sound effect file generation process will be explained. The second generation unit 118 extracts sound effect data separated from the recording data 104. Figure 6 shows an example of the extracted sound effect data. In this example, the second generation unit 118 extracts three sound effect data within the range enclosed by the dashed line. Based on the type information of the extracted sound effect data, the second generation unit 118 identifies the location information of the sound source of the sound effect.
[0033] Referring to Figure 1, the position information holding unit 126 holds information on the type of sound and the position information of the sound source that emits the sound in association. Figure 7 shows the correspondence held by the position information holding unit 126. In this correspondence, the position information indicates the position of the sound source relative to the player character, and in the embodiment, it indicates the height relative to the height position of the player character. In the position information illustrated in Figure 7, "above" indicates that the sound source is located above the player character, "below" indicates that the sound source is located below the player character, and "sideways" indicates that the sound source is located at the same height as the player character. The height position of the player character may be the height of the player character's ears (head), and the positional relationships of "above," "below," and "sideways" may be defined with respect to the position of the player character's ears (head).
[0034] The correspondence shown in Figure 7 may be created manually by the game programmer, or it may be created by analyzing the recorded data. Furthermore, both "up" and "down" may have multiple height positions. For example, for "up," multiple height positions such as "slightly up," "up," and "quite up" may be set. Note that in the example shown in Figure 7, the position information only indicates the height relative to the player character, but in another example, it may also indicate the orientation and distance relative to the player character.
[0035] The second generation unit 118 identifies location information corresponding to the type of sound effect based on the correspondence relationships held in the location information holding unit 126 (S16). If the sound effect is a bird's chirp, the second generation unit 118 refers to the correspondence relationships held in the location information holding unit 126 and identifies that the location information is "up". Once the second generation unit 118 identifies the location information of the sound effect, it generates a sound effect file containing the sound effect data and the location information (S18). At this time, the second generation unit 118 may include the location information "up", which indicates the height position, as metadata in the sound effect file.
[0036] If the angle relative to the player character's frontal direction is defined as the "azimuth angle," the second generation unit 118 may include positional information indicating the bird's height and azimuth angle as metadata in the sound effect file. In this case, the second generation unit 118 may set the azimuth angle arbitrarily, for example, by using a random function to determine the azimuth angle.
[0037] Furthermore, if the recording data 104 is recorded using multiple microphones, and the position of each sound source can be analyzed from the data recorded through the multiple microphones, the second generation unit 118 may identify the position information of each sound source from the analyzed position. The second generation unit 118 may identify the position information of each sound source from data collected by, for example, an ambisonic microphone. For example, if bird calls are recorded, the second generation unit 118 may analyze and derive the position of the bird, which is the sound source, and include the position information indicating the bird's position as metadata in the sound effect file. In this case, the second generation unit 118 may include the position information indicating the bird's height and azimuth angle as metadata in the sound effect file.
[0038] The second generation unit 118 does not have to include the results of its analysis of the sound source's location (location information) as metadata in the sound effect file. The second generation unit 118 may include location information different from the analyzed location information as metadata in the sound effect file. For example, the second generation unit 118 may change the analyzed height position and azimuth angle of the bird within a predetermined range and include location information indicating the changed height position and azimuth angle of the bird as metadata in the sound effect file.
[0039] The second generation unit 118 may apply a fade-in process to the sound effect data. Specifically, the second generation unit 118 may apply a fade-in process to the beginning portion and a fade-out process to the end portion. The second generation unit 118 includes information indicating the type of sound effect as metadata in the sound effect file. In this embodiment, since the sound effect is a bird's song, the second generation unit 118 may include information indicating that the sound effect is a bird's song as metadata in the sound effect file. The second generation unit 118 records the generated sound effect file in the sound effect file recording unit 124.
[0040] As described above, the sound file generation device 100 generates ambient sound files and sound effect files. The sound file generation device 100 can generate ambient sound files and sound effect files from a single recording data 104 by utilizing sound data of sounds different from ambient sounds separated from the recording data 104 as sound effect data. The ambient sound files and sound effect files generated in this way may be incorporated as audio files into game software and played back during the user's gameplay.
[0041] <Audio File Playback Function> Figure 8 shows an example of a surround sound space for a user playing a game. The information processing system 1 includes an information processing device 10, which is a user terminal device, a plurality of speakers 3a to 3j (hereinafter simply referred to as "speakers 3" unless otherwise specified) that realize a multi-channel sound space, a display device 4 such as a television, and a camera 7 that captures the space in front of the display device 4. In this embodiment, the information processing device 10 operates as an audio file playback device that outputs game sounds such as ambient sounds and sound effects during the user's gameplay. The camera 7 may be a stereo camera. The information processing device 10 is connected to the plurality of speakers 3, the display device 4, and the camera 7.
[0042] In this embodiment, the information processing device 10 is a game device that executes a game program, and is connected wirelessly or via a wired connection to a game controller operated by the user, and reflects the operation information supplied from the game controller in the game processing. The information processing device 10 supplies game images to the display device 4 and supplies game sounds to a plurality of speakers 3. In this embodiment, the information processing device 10 has the function of executing a game program, but in a modified example, the information processing device 10 may not have the function of executing a game program, but may be a terminal device that transmits operation information input to the game controller to a cloud server and receives game images and sounds from the cloud server.
[0043] The user sits in front of the display device 4, which displays game images, and plays the game. The information processing device 10 acquires images of the user taken by the camera 7 and determines the user's position in the room. A center speaker 3a is positioned directly in front of the user, and front speakers 3b and 3c, and subwoofers 3k and 3l, which are speakers responsible for ultra-low frequencies, are positioned to the left and right of the front of the user, flanking the display device 4. Surround speakers 3d and 3e are positioned to the left and right of the user. A surround back speaker 3f is positioned behind the user, and surround back speakers 3g and 3h are positioned to the left and right of the rear. Height speakers 3i and 3j are also positioned on the ceiling.
[0044] In the surround sound space of the embodiment, by installing the height speakers 3i and 3j at a position higher than other speakers, it is possible to achieve highly accurate three-dimensional sound image localization. The surround sound space shown in FIG. 8 is an example of the user's game play environment. In order for the information processing apparatus 10 to localize sound images such as game sounds at desired positions within the room space, it grasps the positions and orientations of the speakers 3 in the room.
[0045] In the embodiment, a surround sound space is formed by arranging a plurality of speakers 3 so as to surround the user, but a virtual multi-channel sound space using two speakers may be realized using virtual surround technology. When realizing a virtual surround sound space, the information processing apparatus 10 localizes virtual sound sources using a head-related transfer function (HRTF). Further, the information processing apparatus 10 may provide stereophonic sound optimized for the user to the user through headphones based on the result of estimating the head-related transfer function (HRTF) of the user individual.
[0046] FIG. 9 shows the hardware configuration of the information processing apparatus 10. The information processing apparatus 10 includes a main power button 20, a power-on LED 21, a standby LED 22, a system controller 24, a clock 26, a device controller 30, a media drive 32, a USB module 34, a flash memory 36, a wireless communication module 38, a wired communication module 40, a subsystem 50, and a main system 60.
[0047] The main system 60 includes a main CPU (Central Processing Unit), a memory which is the main storage device and a memory controller, a GPU (Graphics Processing Unit), and audio dedicated hardware. The GPU is mainly used for the arithmetic processing of game programs. The audio dedicated hardware calculates in real time the sounds output from each of the plurality of speakers 3 based on the position information of the sound sources in order to realize object-based stereophonic sound, and outputs three-dimensional sound. These functions may be configured as a system on chip and formed on one chip. The main CPU has a function of executing a game program recorded in an auxiliary storage device (not shown) or a ROM medium 44.
[0048] The subsystem 50 includes a sub CPU, a memory which is the main storage device and a memory controller, etc., does not include a GPU, and does not have a function of executing a game program. The number of circuit gates of the sub CPU is less than that of the main CPU, and the operating power consumption of the sub CPU is less than that of the main CPU. The sub CPU operates even while the main CPU is in the standby state, and its processing function is restricted in order to keep the power consumption low.
[0049] The main power button 20 is an input unit for receiving an operation input from the user, is provided on the front surface of the housing of the information processing device 10, and is operated to turn on or off the power supply to the main system 60 of the information processing device 10. The power-on LED 21 lights up when the main power button 20 is turned on, and the standby LED 22 lights up when the main power button 20 is turned off.
[0050] The system controller 24 detects the pressing of the main power button 20 by the user. When the main power button 20 is pressed while the main power is off, the system controller 24 acquires the pressing operation as an "on instruction", while when the main power button 20 is pressed while the main power is on, the system controller 24 acquires the pressing operation as an "off instruction".
[0051] Clock 26 is a real-time clock that generates current date and time information and supplies it to the system controller 24, subsystem 50, and main system 60. The device controller 30 is configured as an LSI (Large-Scale Integrated Circuit) that performs information transfer between devices, similar to a southbridge. As shown in the figure, devices such as the system controller 24, media drive 32, USB module 34, flash memory 36, wireless communication module 38, wired communication module 40, subsystem 50, and main system 60 are connected to the device controller 30. The device controller 30 absorbs the differences in electrical characteristics and data transfer speeds of each device and controls the timing of data transfer.
[0052] The media drive 32 is a drive device that drives a ROM medium 44 containing application software such as games and license information, and reads programs and data from the ROM medium 44. The ROM medium 44 may be a read-only recording medium such as an optical disc, magneto-optical disc, or Blu-ray disc.
[0053] The USB module 34 is a module that connects to external devices via a USB cable. The USB module 34 may also be connected to an external auxiliary storage device and a camera 7 via a USB cable. The flash memory 36 is an auxiliary storage device that constitutes the internal storage. The wireless communication module 38 communicates wirelessly with, for example, the game controller 6 using a communication protocol such as Bluetooth® protocol or IEEE802.11 protocol. The wired communication module 40 communicates with external devices via a wired connection and connects to an external network such as the internet via an access point.
[0054] Figure 10 shows the functional blocks of the information processing device 10 in the embodiment. The information processing device 10 includes a game processing unit 200, which has a game execution unit 210, a game image output unit 220, a game sound localization processing unit 230, and a game sound output unit 240.
[0055] The functions of the components in the information processing device 10 may be realized in a circuit or processing circuitry, including a general-purpose processor, an application-specific processor, an integrated circuit, an ASIC (Application Specific Integrated Circuit), a CPU (a Central Processing Unit), a conventional circuit, and / or a combination thereof, configured or programmed to realize the functions described herein. A processor is considered to be a circuit or processing circuitry that includes transistors and other circuits. A processor may also be a programmed processor that executes a program stored in memory.
[0056] In this specification, circuits, units, and means are hardware programmed to perform or execute the functions described herein. Such hardware may be any hardware disclosed herein, or any hardware known to be programmed to perform or execute the functions described herein.
[0057] If the hardware is a processor that is considered to be of the type of circuit, the circuit, means, or unit may be a combination of hardware and software used to constitute the hardware and / or processor.
[0058] The game execution unit 210 executes a game program (hereinafter sometimes simply referred to as "the game") to generate game image data and game sound data. The functions shown as the game execution unit 210 are realized by system software, a game program, a GPU that executes rendering processes, dedicated audio hardware, etc. Note that a game is just one example of an application, and the game execution unit 210 may execute applications other than games.
[0059] The game execution unit 210 performs calculations to move the player character in the virtual space based on the operation information input by the user to the game controller 6. Based on the calculation results in the virtual space, the game execution unit 210 generates game image data from the viewpoint position (virtual camera) in the virtual space. The game execution unit 210 also generates game sound data in the virtual space.
[0060] For example, in a first-person shooter (FPS) game, the game execution unit 210 generates image data from the perspective of the player character controlled by the user (first-person perspective) and supplies it to the game image output unit 220. The game image output unit 220 displays the game image data in the display area of the display device 4.
[0061] Figure 11 shows an example of a game image displayed on the display device 4. The game image output unit 220 generates display data from game image data and displays the game image in the display area of the display device 4.
[0062] The game execution unit 210 identifies the game sound file to be played by the sound source (sound source) in the game space in accordance with the game's progress and supplies the game sound data to the game sound localization processing unit 230. At this time, the game execution unit 210 also supplies sound source position information, which indicates the location of the sound source, to the game sound localization processing unit 230. For example, the sound source position information may include the position and direction (orientation) of the player character in the world coordinate system of the game space, and the position information of the sound source. The game sound localization processing unit 230 calculates the relative position of the sound source to the player character from the sound source position information and determines the position in three-dimensional space where the sound image of the game sound is localized.
[0063] The following describes the case where sound effects are output. Sound effect files incorporated into game software contain sound effect data along with positional information indicating the height of the sound source as metadata. The game execution unit 210 then identifies the sound effect file to be played and generates sound source position information, which includes the position and direction (orientation) of the player character in the world coordinate system of the game space, and the position of the sound source, based on the positional information contained in the sound effect file. The game sound positioning unit 230 then supplies the sound effect data and sound source position information to the game sound localization processing unit 230. The game sound localization processing unit 230 converts the relative positional relationship between the player character and the sound source into the positional relationship between the user and the localization position of the sound image in three-dimensional space, and determines the localization position of the sound image of the sound effect.
[0064] If the sound effect is a bird's song, the sound source location information is "above" the player character. Therefore, the game sound localization processing unit 230 determines the localization position of the bird's song sound image above the user in three-dimensional space. The game sound localization processing unit 230 performs sound image localization processing to localize the sound image at the determined localization position. Sound image localization processing is a process that uses surround technology to make the user perceive that the sound is emanating from the localization position. The game sound localization processing unit 230 adjusts the game sound output of multiple speakers 3 to generate a localized sound effect signal. When localizing the sound image at the localization position using virtual surround technology, the game sound localization processing unit 230 performs sound image localization processing using a head-related transfer function and generates a localized sound effect signal. The game sound output unit 240 supplies the localized sound effect signal to each speaker 3. As a result, the user can hear the bird's song from the localization position (above the user).
[0065] If the sound source's location information is "above" the player character, the game sound localization processing unit 230 may determine the localization position of the bird sound image at a position directly above the user in three-dimensional space, or it may determine the localization position of the bird sound image at an upper position other than directly above. If the angle with respect to the front direction of the player character (user) is defined as the "azimuth angle," the game sound localization processing unit 230 may arbitrarily set the azimuth angle to determine the localization position of the bird sound. By setting the azimuth angle randomly, for example, the game sound localization processing unit 230 can enable the user to hear the bird sound from various positions above the user.
[0066] The present disclosure has been described above based on embodiments. These embodiments are illustrative, and it will be understood by those skilled in the art that various modifications are possible in combinations of their components and processing processes, and that such modifications are also within the scope of the present disclosure. In the embodiments, the sound source separation unit 112 separates the recorded data 104 into ambient sound data and sound effect data using a sound source separation model, but in the modifications, the recorded data 104 may be separated into ambient sound data and sound effect data using a multi-channel sound source separation method that uses multiple microphones.
[0067] In the sound file generation device 100, the first generation unit 116 may set the position information of the ambient sound source to omnidirectional and include the position information indicating omnidirectional as metadata in the ambient sound file. As a result, in the information processing device 10, the game sound output unit 240 will output ambient sounds from all directions of the user. The game execution unit 210 may decide to output sound effects at non-periodic timings when outputting ambient sounds. As a result, when the user is listening to ambient sounds output from all directions, they will hear bird calls from above at non-periodic timings, making it less likely that they will notice the ambient sounds are being played in a loop, thus reducing the possibility of impairing immersion in the game.
[0068] This disclosure may include the following embodiments: [Item 1] A sound file generating device comprising a circuit configured to, wherein the circuit acquires recording data of ambient sounds, separates the recording data into ambient sound data and sound effect data, identifies location information corresponding to the separated sound effects, and generates a sound effect file containing the sound effect data and location information. [Item 2] The sound file generating device according to Item 1, wherein the circuit includes location information indicating height as metadata in the sound effect file. [Item 3] The sound file generating device according to Item 2, wherein the circuit includes type information indicating the type of sound effect as metadata in the sound effect file. [Item 4] The sound file generating device according to Item 1, wherein the circuit identifies the type of separated sound effect. [Item 5] The sound file generating device according to Item 1, wherein the circuit identifies the type of separated ambient sound and generates an ambient sound file containing ambient sound data and type information indicating the type of ambient sound. [Item 6] The sound file generation device according to Item 1, wherein the circuit identifies location information corresponding to the type of separated sound effect and generates a sound effect file containing sound effect data and location information. [Item 7] The sound file generation device according to Item 1, wherein the circuit generates a sound effect file containing sound effect data and location information different from the identified location information. [Item 8] A method for generating a sound file, comprising: acquiring recording data of ambient sound; separating the recording data into ambient sound data and sound effect data; identifying location information corresponding to the type of separated sound effect; and generating a sound effect file containing sound effect data and location information. [Item 9] A recording medium on which a program executed on a computer is recorded, wherein the program causes the computer to perform the following functions: acquiring recording data of ambient sound; separating the recording data into ambient sound data and sound effect data; identifying location information corresponding to the type of separated sound effect; and generating a sound effect file containing sound effect data and location information.
[0069] This disclosure can be used in the technical field of generating audio files.
[0070] 1... Information processing system, 3... Speaker, 100... Sound file generation device, 102... Processing device, 104... Recorded data, 110... Sound data acquisition unit, 112... Sound source separation unit, 114... Sound file generation unit, 116... First generation unit, 118... Second generation unit, 120... Sound file recording device, 122... Ambient sound file recording unit, 124... Sound effect file recording unit, 126... Position information holding unit, 200... Game processing unit, 210... Game execution unit, 220... Game image output unit, 230... Game sound localization processing unit, 240... Game sound output unit.
Claims
1. A sound file generation device comprising: a sound data acquisition unit that acquires recording data of ambient sounds; a sound source separation unit that separates the recording data into ambient sound data and sound effect data; and a sound file generation unit that identifies location information corresponding to the separated sound effects and generates a sound effect file containing the sound effect data and location information.
2. The sound file generation device according to claim 1, characterized in that the sound file generation unit includes position information indicating the height position as metadata in the sound effect file.
3. The sound file generation device according to claim 2, characterized in that the sound file generation unit includes type information indicating the type of sound effect as metadata in the sound effect file.
4. The sound file generation device according to claim 1, characterized in that the sound source separation unit identifies the type of sound effect that has been separated.
5. The sound file generation device according to claim 1, characterized in that the sound source separation unit identifies the type of separated ambient sound, and the sound file generation unit generates an ambient sound file containing ambient sound data and type information indicating the type of ambient sound.
6. The sound file generation device according to claim 1, characterized in that the sound file generation unit identifies location information corresponding to the separated sound effect types and generates a sound effect file including sound effect data and location information.
7. The sound file generation device according to claim 1, characterized in that the sound file generation unit generates sound effect data and a sound effect file that includes location information different from the specified location information.
8. A method for generating an audio file, comprising the steps of: acquiring recording data of ambient sounds; separating the recording data into ambient sound data and sound effect data; identifying location information corresponding to the separated sound effects; and generating an audio file containing the sound effect data and location information.
9. A program to enable a computer to perform the following functions: acquire recording data of ambient sounds; separate the recording data into ambient sound data and sound effect data; identify location information corresponding to the separated sound effects; and generate a sound effect file containing the sound effect data and location information.
Citation Information
Patent Citations
Acoustic scene reconstruction device, acoustic scene reconstruction method, and program
JP2020030376A