Audio file generation device, audio file playback device, audio file generation method, and audio file playback method
The device automatically divides and concatenates sound files based on similarity, addressing repetitive playback in games by enhancing variation and immersion through seamless ambient sound generation.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-10-04
- Publication Date
- 2026-04-09
AI Technical Summary
The challenge in games is the large number of environmental sound files, which are looped, causing repetitive playback of characteristic sounds, leading to user detection and the need for manual editing to remove unnecessary sounds, which is time-consuming and laborious.
A device and method that automatically divides audio data into multiple files, derives similarity between segments, and records these similarities to seamlessly concatenate and play back sound files based on similarity, reducing repetitive playback and enhancing variation.
This approach provides varied and seamless ambient sound playback, reducing user detection of looped sounds by ensuring high similarity between concatenated files, thus enhancing immersion without manual editing effort.
Smart Images

Figure JP2024035642_09042026_PF_FP_ABST
Abstract
Description
Sound file generation device, sound file playback device, sound file generating method, and sound file playback method
[0001] The present disclosure relates to a technique for generating and / or playing back sound files such as environmental sounds used in games and the like.
[0002] In games, various types of environmental sounds are output to enhance the sense of immersion. Environmental sound refers to the sound heard from the surrounding space. In games, for example, sounds of forests, underwater, big cities, parks, rain, waves, and various other environmental sounds are prepared. By outputting appropriate environmental sounds at the timing corresponding to the game scene, the sense of immersion in the game can be enhanced. Although environmental sounds are produced by sound creators using various techniques, by recording actual ambient sounds in a natural environment and using them as environmental sounds, real-world sounds can be provided to users.
[0003] In games that use many types of environmental sounds, the number of environmental sound files becomes extremely large, so it is required to reduce the size of each environmental sound file. The duration of the environmental sound file is often set within the range of 1.5 seconds to 30 seconds, depending on the type. Therefore, when outputting environmental sounds, the environmental sound file is repeatedly played (loop played) so that the environmental sound is output for a time longer than the duration of the environmental sound file. Therefore, if a characteristic sound exists in the environmental sound file, that sound will be output every loop cycle (the duration of the environmental sound file), which may be a factor that makes the user notice that it is loop-playing.
[0004] Therefore, the sound creator uses an editing device (personal computer) to play back the recorded environmental sound file and manually remove unnecessary sounds while listening to the output environmental sound. However, it is not easy to completely remove unnecessary sounds, and it is a very time-consuming and laborious task. Therefore, one object of the present disclosure is to provide a technique for increasing the variations of sounds when outputting sounds such as environmental sounds.
[0005] An audio file generation device according to one aspect of the present disclosure includes: an audio data acquisition unit that acquires audio data; an audio file generation unit that divides the acquired audio data into a plurality of divided audio data and generates a plurality of audio files containing the divided audio data; a similarity derivation unit that derives the similarity between two divided audio data; and a recording unit that records the derived similarity between the two divided audio data.
[0006] Another embodiment of the present disclosure of a method for generating sound files includes the steps of: acquiring sound data; dividing the acquired sound data into a plurality of divided sound data to generate a plurality of sound files containing the divided sound data; deriving the similarity between two divided sound data; and recording the derived similarity between the two divided sound data.
[0007] A sound file playback device in yet another aspect of the present disclosure includes a similarity recording unit that records the similarity between two sound data, a sound file playback unit that concatenates and plays multiple sound files based on the similarity between the two sound data recorded in the similarity recording unit, and an output unit that outputs the sound to be played.
[0008] A method for playing an audio file in yet another aspect of the present disclosure includes the steps of: recording the similarity between two audio data; concatenating and playing a plurality of audio files based on the similarity between the two recorded audio data; and outputting the sound to be played.
[0009] Furthermore, any combination of the above components, as well as any conversion of the expressions of this disclosure between methods, apparatus, systems, recording media, computer programs, etc., are also valid forms of this disclosure.
[0010] This figure shows an example of the functional blocks of an audio file generation device. This figure shows an example of audio data. This figure shows a flowchart of the audio file generation procedure. This figure shows an example of divided audio data obtained by dividing audio data. This figure shows an example of the similarity between two divided audio data. This figure shows an information processing system. This figure shows the hardware configuration of the information processing device. This figure shows an example of the functional blocks of the information processing device. (a) to (g) are diagrams to explain the playback order of audio files. This figure shows an example of two audio files played simultaneously. This figure explains the similarity derivation method in a modified example.
[0011] <Sound File Generation Function> Figure 1 shows the functional blocks of a sound file generation device that automatically generates multiple sound files from sound data recorded with ambient sounds. The information processing device 10 of this embodiment has the function of dividing the sound data 104 and automatically generating multiple sound files.
[0012] Figure 2 shows an example of sound data 104. Sound data 104 may be recording data of natural ambient sounds recorded by a sound creator, and may be data from which sounds such as bird calls and car horns have been removed by editing by the sound creator. Alternatively, sound data 104 may be recording data that has not been edited by the sound creator. Sound data 104 has a duration of, for example, about 30 seconds, but may have a duration longer or shorter.
[0013] The sound file generation device 100 comprises a processing device 102 and a recording device 120. The processing device 102 has the function of dividing sound data 104 into a plurality of divided sound data to generate a plurality of sound files and deriving the similarity between two divided sound data, and comprises a sound data acquisition unit 110, a sound file generation unit 112, and a similarity derivation unit 114. The recording device 120 has a sound file recording unit 122 that records the plurality of sound files generated by the processing device 102, and a similarity recording unit 124 that records the similarity between the divided sound data derived by the processing device 102. The recording device 120 is a large-capacity recording device such as an HDD (hard disk drive) or SSD (solid state drive), and may be an internal recording device or an external recording device connected to the processing device 102 by USB (Universal Serial Bus) or the like.
[0014] The functions of the components in the processing unit 102 may be realized in a circuit or processing circuitry, including a general-purpose processor, an application-specific processor, an integrated circuit, an ASIC (Application Specific Integrated Circuit), a CPU (a Central Processing Unit), a conventional circuit, and / or a combination thereof, configured or programmed to realize the functions described herein. A processor is considered to be a circuit or processing circuitry that includes transistors and other circuits. A processor may also be a programmed processor that executes a program stored in memory.
[0015] In this specification, circuits, units, and means are hardware programmed to perform or execute the functions described herein. Such hardware may be any hardware disclosed herein, or any hardware known to be programmed to perform or execute the functions described herein.
[0016] If the hardware is a processor that is considered to be of the type of circuit, the circuit, means, or unit may be a combination of hardware and software used to constitute the hardware and / or processor.
[0017] Figure 3 shows a flowchart of the sound file generation procedure in the embodiment. When the sound data acquisition unit 110 acquires sound data 104 input from an external source (S10), it provides it to the sound file generation unit 112. The sound file generation unit 112 divides the provided sound data 104 into multiple divided sound data, generates multiple sound files containing the divided sound data, and records them in the sound file recording unit 122 (S12).
[0018] Figure 4 shows an example of divided sound data obtained by dividing sound data. In this example, the sound file generation unit 112 divides the sound data 104 into 10 divided sound data A to J. It is preferable for the sound file generation unit 112 to divide the sound data 104 into multiple divided sound data of substantially the same duration.
[0019] For example, if the duration of the sound data 104 is 30 seconds, the sound file generation unit 112 may cut the sound data 104 into 3-second segments and extract 10 segmented sound data A to J. In another example, the sound file generation unit 112 may cut the sound data 104 into 1-second segments and extract 30 segmented sound data.
[0020] When dividing the sound data 104 into equal intervals, it is preferable for the sound file generation unit 112 to apply a fade process to the divided sound data. Specifically, the sound file generation unit 112 applies a fade-in process to the beginning portion of the divided sound data so that the volume level at the beginning starts at a minimum value (for example, zero), and applies a fade-out process to the end portion so that the volume level at the end of the divided sound data ends at a minimum value (for example, zero). By applying a fade process to each of the divided sound data, the sound file generation unit 112 can ensure that when the divided sound data is combined and played back on the user's game device, it does not cause any discomfort to the user.
[0021] In the above example, the sound file generation unit 112 applied a fade process to each of the multiple divided sound data that were extracted with substantially the same time length. However, the sound data 104 may also be divided into multiple divided sound data by cutting it at the zero-crossing point. In this case as well, it is preferable that the time lengths of the multiple divided sound data are substantially the same. "Substantially the same time length" means that the ratio of the time length T1 of the shortest divided sound data to the time length T2 of the longest divided sound data satisfies the relationship T1 / T2 > 2 / 3. The sound file generation unit 112 may cut the sound data 104 at the zero-crossing point so that the divided sound data have substantially the same time length.
[0022] The sound file generation unit 112 generates multiple sound files containing segmented sound data and records them in the sound file recording unit 122. In this embodiment, the sound file generation unit 112 generates sound files A to J, each containing segmented sound data A to J, and records them in the sound file recording unit 122. At this time, the sound file generation unit 112 includes information indicating the type of ambient sound as metadata in the sound files. For example, if the ambient sound is the sound of a forest, the sound file generation unit 112 includes information indicating that the ambient sound is the sound of a forest as metadata in each of the sound files A to J.
[0023] In this embodiment, the similarity derivation unit 114 derives the similarity between two divided sound data (S14). The similarity derivation unit 114 derives the similarity between two divided sound data for all combinations of two divided sound data in the plurality of divided sound data A to J. The similarity derivation unit 114 may derive the cross-correlation coefficient between the two divided sound data as the similarity.
[0024] The cross-correlation function is a function used to determine the similarity (cross-correlation coefficient) between two time series data. For one time series data, the other time series data is shifted by a small amount of time, and the cross-correlation coefficient at each lag is calculated. The closer the cross-correlation coefficient is to 1, the higher the similarity; the closer it is to 0, the lower the similarity. When the cross-correlation coefficient is 1, the two time series data are identical. The similarity derivation unit 114 may extract the maximum cross-correlation coefficient among the cross-correlation coefficients calculated for the two divided sound data. The similarity derivation unit 114 records the extracted maximum cross-correlation coefficient as the similarity between the two divided sound data in the similarity recording unit 124 (S16). In another example, the similarity derivation unit 114 may derive the cross-correlation coefficient when the lag is 0, that is, the cross-correlation coefficient when there is no time shift, as the similarity between the two divided sound data and record it in the similarity recording unit 124.
[0025] Figure 5 shows an example of a similarity recording unit 124 that records the similarity between two split sound data. The similarity recording unit 124 records the similarity derived by the similarity derivation unit 114 for all combinations of two split sound data. The similarity recording unit 124 may record the similarity derived for all combinations of two split sound data as a single similarity file.
[0026] As described above, the sound file generation device 100 generates multiple sound files and similarity files. The multiple sound files and similarity files generated in this way may be incorporated into game software as part of the audio files and used by the user during gameplay.
[0027] <Audio File Playback Function> Figure 6 shows an information processing system 1 according to an embodiment. The information processing system 1 comprises an information processing device 10 operated by the user and a server device 5. The access point (hereinafter referred to as "AP") 8 has the functions of a wireless access point and a router, and the information processing device 10 connects to the AP 8 via wireless or wired connection and is connected to the server device 5 on the network 3 so as to be able to communicate. The information processing device 10 of this embodiment operates as an audio file playback device that outputs ambient sounds during the user's gameplay.
[0028] The information processing device 10 is connected wirelessly or via a wired connection to the input device 6 operated by the user, and the input device 6 outputs information operated by the user to the information processing device 10. When the information processing device 10 receives operation information from the input device 6, it reflects it in the processing of system software or game software, and outputs the processing result from the output device 4. In the information processing system 1, the information processing device 10 is a game device (game console) that runs games, and the input device 6 is a device such as a game controller that supplies user operation information to the information processing device 10. The input device 6 may also be an input interface such as a keyboard or mouse.
[0029] The auxiliary storage device 2 is a high-capacity storage device such as an HDD (hard disk drive) or an SSD (solid-state drive), and may be an internal storage device or an external storage device connected to the information processing device 10 by USB (Universal Serial Bus) or the like. The output device 4 may be a television having a display for outputting images and a speaker for outputting sound. The output device 4 may be connected to the information processing device 10 by a wired cable or by wireless connection.
[0030] The imaging device, camera 7, is located near the output device 4 and images the space around the output device 4. Figure 6 shows an example where camera 7 is mounted on top of the output device 4, but it may also be located on the side or bottom of the output device 4. In any case, it is positioned to capture images of the user located in front of the output device 4. Camera 7 may also be a stereo camera.
[0031] Server device 5 provides network services to users of information processing system 1. Server device 5 manages user identifiers (user accounts) that identify each user, and each user signs in to the network services provided by server device 5 using their user account. By signing in to the network services from information processing device 10, users can register game save data and trophies, which are virtual rewards earned during gameplay, with server device 5. Once save data and trophies are registered with server device 5, the save data and trophies can be synchronized even if the user uses an information processing device other than information processing device 10.
[0032] Figure 7 shows the hardware configuration of the information processing device 10. The information processing device 10 consists of a main power button 20, a power-on LED 21, a standby LED 22, a system controller 24, a clock 26, a device controller 30, a media drive 32, a USB module 34, a flash memory 36, a wireless communication module 38, a wired communication module 40, a subsystem 50, and a main system 60.
[0033] The main system 60 includes a main CPU (Central Processing Unit), main memory (memory and memory controller), and a GPU (Graphics Processing Unit). The GPU is primarily used for processing game programs. The main CPU has the function of starting system software and executing game software installed on the auxiliary storage device 2 within the environment provided by the system software. The subsystem 50 includes a sub-CPU, main memory (memory and memory controller), and other components, but does not have a GPU.
[0034] The main CPU has the function of executing game software installed on the auxiliary storage device 2, while the sub-CPU does not have such a function. However, the sub-CPU has the function of accessing the auxiliary storage device 2 and sending and receiving data with the server device 5. The sub-CPU is configured with only these limited processing functions and can therefore operate with less power consumption compared to the main CPU. These functions of the sub-CPU are executed when the main CPU is in standby mode.
[0035] The main power button 20 is an input unit that receives user input and is located on the front of the housing of the information processing device 10. It is operated to turn the power supply of the information processing device 10 to the main system 60 on or off. The power-on LED 21 lights up when the main power button 20 is turned on, and the standby LED 22 lights up when the main power button 20 is turned off. The system controller 24 detects when the user presses the main power button 20.
[0036] Clock 26 is a real-time clock that generates current date and time information and supplies it to the system controller 24, subsystem 50, and main system 60.
[0037] The device controller 30 is configured as a Large-Scale Integrated Circuit (LSI) that performs information transfer between devices, similar to a Southbridge. As shown in the figure, devices such as the system controller 24, media drive 32, USB module 34, flash memory 36, wireless communication module 38, wired communication module 40, subsystem 50, and main system 60 are connected to the device controller 30. The device controller 30 absorbs the differences in electrical characteristics and data transfer speeds of each device and controls the timing of data transfer.
[0038] The media drive 32 is a drive device that drives a ROM medium 44 containing application software such as games and license information, and reads programs and data from the ROM medium 44. The ROM medium 44 is a read-only recording medium such as an optical disc, magneto-optical disc, or Blu-ray disc.
[0039] The USB module 34 is a module that connects to external devices via a USB cable. The USB module 34 may also be connected to the auxiliary storage device 2 and the camera 7 via a USB cable. The flash memory 36 is an auxiliary storage device that constitutes the internal storage. The wireless communication module 38 communicates wirelessly with the input device 6 using a communication protocol such as Bluetooth® protocol or IEEE802.11 protocol. The wired communication module 40 communicates with external devices via a wired connection and connects to the network 3 via AP8.
[0040] Figure 8 shows an example of a functional block of the information processing device 10. The information processing device 10 of this embodiment comprises a processing unit 200, a communication unit 202, and a recording device 230. The processing unit 200 comprises a game execution unit 210, a game image generation unit 212, a game sound generation unit 214, and an output unit 220. The game sound generation unit 214 has a voice file playback unit 216 and a sound file playback unit 218. The recording device 230 is an auxiliary storage device 2 or flash memory 36, which records system software and game software. In the example shown in Figure 8, a sound file recording unit 232 that records sound files included in the game software and a similarity recording unit 234 that records similarity files included in the game software are shown, but the game program itself and voice files of game characters are omitted from the illustration.
[0041] The functions of the components in the information processing apparatus 10 may be realized in a circuit (circuitry) or processing circuitry including a general-purpose processor, a specific-purpose processor, an integrated circuit, ASICs (Application Specific Integrated Circuits), a CPU (a Central Processing Unit), a conventional circuit, and / or a combination thereof, which is configured or programmed to realize the functions described in this specification. The processor may be regarded as a circuitry or processing circuitry including transistors and other circuits. The processor may be a programmed processor that executes a program stored in a memory.
[0042] In this specification, a circuitry, unit, or means is hardware programmed to realize the described functions or hardware that executes them. The hardware may be any hardware disclosed in this specification or any hardware known as being programmed or executing to realize the described functions.
[0043] When the hardware is a processor regarded as a type of circuitry, the circuitry, means, or unit may be a combination of the hardware and software used to configure the hardware and / or the processor.
[0044] The communication unit 202 has the functions of the wireless communication module 38 and the wired communication module 40 combined. During the user's game play, the communication unit 202 receives information (game operation information) on the user's operation of the input device 6 and provides it to the game execution unit 210.
[0045] The game execution unit 210 executes the game according to the user's input. During gameplay by the user, the game execution unit 210 executes the game program based on game operation information and performs calculations to move the player character in the virtual space. The game image generation unit 212 includes a GPU (Graphics Processing Unit) and, upon receiving the results of calculations in the virtual space, generates game images from the viewpoint position (virtual camera) in the virtual space. The game sound generation unit 214 has the function of generating game sounds in the game space. The voice file playback unit 216 plays voice files for the player character and enemy characters, and the sound file playback unit 218 plays sound files such as ambient sounds and sound effects. The output unit 220 outputs game images from the display of the output device 4 and outputs character voices, ambient sounds, and / or sound effects from the speaker of the output device 4. The user operates the input device 6 while viewing the game images and game sounds output from the output device 4 to advance the game.
[0046] In this embodiment, the similarity recording unit 234 holds a similarity file that records the similarity between two sound data in a plurality of sound files recorded in a plurality of sound files recorded in the sound file recording unit 232. Here, the sound file recording unit 232 records the sound files A to J described above, and the similarity recording unit 234 holds a similarity file (see Figure 5) that records the similarity between two sound data in the divided sound data A to J (hereinafter also referred to as "sound data A to J").
[0047] The sound file playback unit 218 plays back a plurality of sound files held in the sound file recording unit 232 by connecting them together based on the similarity between two sound data recorded in the similarity file, and the output unit 220 outputs the played-back sound from the speaker. In the embodiment, the sound file playback unit 218 determines the order of the sound files to be played back by referring to the similarity between two sound data. A high similarity between two sound data means that the playback sounds of the two sound data are similar, and a low similarity between two sound data means that the playback sounds of the two sound data are not similar. Therefore, when two sound data with low similarity are connected and played back, the playback sound changes suddenly at the boundary between the two sound data, which will give the user a sense of discomfort. Therefore, when playing back a plurality of sound data seamlessly connected together, it is preferable to connect and play back two sound data with high similarity.
[0048] Therefore, the sound file playback unit 218 of the embodiment determines the order of the sound files to be played back so that after playing back the first sound file including the first sound data, it plays back the second sound file including the second sound data with high similarity to the first sound data. That is, the sound file playback unit 218 determines the sound file to be played back next to the first sound file including the first sound data as the second sound file including the second sound data with high similarity to the first sound data. Hereinafter, a method for determining the order of the sound files to be played back will be described. FIGS. 9(a) to 9(g) show the playback order of the sound files determined by the sound file playback unit 218.
[0049] (Step 1) The sound file playback unit 218 first randomly selects one from the sound files A to J as the sound file to be played back first. FIG. 9(a) shows a state where the sound file A including the sound data A is selected as the sound file to be played back first.
[0050] (Step 2) Figure 9(b) shows the state where sound file B is selected after sound file A. The sound file playback unit 218 determines that the sound file to be played after sound file A, which contains sound data A, is sound file B, which contains sound data B with a high similarity to sound data A. Referring to the similarity file shown in Figure 5, the similarity (cross-correlation coefficient) between sound data A and sound data B is 0.9, and with respect to sound data A, the similarity between sound data A and sound data B is the highest. For this reason, the sound file playback unit 218 may decide to play sound file B after sound file A.
[0051] (Step 3) Figure 9(c) shows the state in which sound file I is selected after sound file B. The sound file playback unit 218 determines that the sound file to be played after sound file B, which contains sound data B, is sound file I, which contains sound data I with a high similarity to sound data B. Referring to the similarity file shown in Figure 5, the similarity (cross-correlation coefficient) between sound data B and sound data I is 0.9, and with respect to sound data B, the similarity between sound data B and sound data I is the highest. For this reason, the sound file playback unit 218 may decide to play sound file I after sound file B.
[0052] The similarity between sound data B and sound data A is 0.9, which is equal to the similarity between sound data B and sound data I. However, the sound file playback unit 218 does not decide to play sound file A after sound file B. Referring to Figure 9(b), since sound file A is played before sound file B, playing sound file A after sound file B would increase the playback frequency of sound file A per unit time. Because a high playback frequency of a particular sound file may cause discomfort to the user, the sound file playback unit 218 imposes a limit on the playback frequency of a single sound file.
[0053] In this embodiment, the sound file playback unit 218 determines that the sound file to be played after the first sound file is the second sound file, provided that a predetermined time has elapsed since the previous playback end time of the second sound file. In other words, the sound file playback unit 218 does not determine that the second sound file is the sound file to be played after the first sound file if the predetermined time has not elapsed since the previous playback end time of the second sound file. This predetermined time for determining the playback frequency may be, for example, 20 seconds. In this embodiment, the duration of each sound file is 3 seconds, and when sound file B finishes playing, only 3 seconds have elapsed since the previous playback end time of sound file A. Therefore, the sound file playback unit 218 determines that sound file A is not suitable as the sound file to be played after sound file B.
[0054] (Step 4) Figure 9(d) shows the state in which sound file E is selected after sound file I. The sound file playback unit 218 determines that the sound file to be played after sound file I, which contains sound data I, is sound file E, which contains sound data E with a high similarity to sound data I. Referring to the similarity file shown in Figure 5, the similarity (cross-correlation coefficient) between sound data I and sound data E is 0.9, and with respect to sound data I, the similarity between sound data I and sound data E is the highest. The similarity between sound data I and sound data B is also 0.9, which is the same as the similarity between sound data I and sound data E, but when sound file I finishes playing, only 3 seconds have passed since the previous playback end time of sound file B. Therefore, the sound file playback unit 218 determines that it is not appropriate to play sound file B after sound file I, and decides to play sound file E after sound file I.
[0055] (Step 5) Figure 9(e) shows the state where sound file D is selected after sound file E. The sound file playback unit 218 determines that the sound file to be played after sound file E, which contains sound data E, is sound file D, which contains sound data D and has a high similarity to sound data E. Referring to the similarity file shown in Figure 5, the similarity between sound data E and sound data I is 0.9, which is the highest, but when sound file E finishes playing, only 3 seconds have passed since the previous playback end time of sound file I. Therefore, the sound file playback unit 218 determines that sound file I is not suitable as the sound file to be played after sound file E.
[0056] Next, searching for the highest similarity, we find that the similarity between sound data E and sound data D is 0.7, and similarly, the similarity between sound data E and sound data F is also 0.7. When there are multiple options, the sound file playback unit 218 may determine one of the multiple options using a random function or the like. In this example, the sound file playback unit 218 decided to play sound file D after sound file E, but it may also decide to play sound file F after sound file E.
[0057] (Step 6) Figure 9(f) shows the state in which sound file F is selected after sound file D. The sound file playback unit 218 determines that the sound file to be played after sound file D, which contains sound data D, is sound file F, which contains sound data F with a high similarity to sound data D. Referring to the similarity file shown in Figure 5, the similarity between sound data D and sound data F is 0.9, and with respect to sound data D, the similarity between sound data D and sound data F is the highest. Therefore, the sound file playback unit 218 decides to play sound file F after sound file D.
[0058] (Step 7) Figure 9(g) shows the state in which sound file G is selected after sound file F. The sound file playback unit 218 determines that the sound file to be played after sound file F, which contains sound data F, is sound file G, which contains sound data G with a high similarity to sound data F. Referring to the similarity file shown in Figure 5, the similarity between sound data F and sound data D is 0.9, which is the highest, but when sound file F finishes playing, only 3 seconds have passed since the previous playback end time of sound file D. Therefore, the sound file playback unit 218 determines that sound file D is not suitable as the sound file to be played after sound file F.
[0059] Next, searching for high similarity, the similarity between sound data F and sound data A is found to be 0.8, and similarly, the similarity between sound data F and sound data G is also 0.8. At this point, the elapsed time since the end of the previous playback of sound file A is 15 seconds, which is not 20 seconds. Therefore, the sound file playback unit 218 determines that it is not appropriate to play sound file A after sound file F, and decides to play sound file G after sound file F.
[0060] As described above, the sound file playback unit 218 determines the order in which to combine multiple sound files. By combining and playing multiple sound files, the sound file playback unit 218 can provide variations in ambient sounds. Therefore, ambient sounds can be provided to the user without the user realizing that they are hearing a loop. The sound file playback unit 218 sequentially plays the multiple sound files in the determined order, and the output unit 220 outputs the played ambient sounds from the speaker. The sound file playback unit 218 seamlessly combines the end of one sound file with the beginning of the next sound file during playback, but it may also overlap a portion of the end of one sound file with a portion of the beginning of the next sound file during playback.
[0061] In this embodiment, the sound file playback unit 218 determines that the sound file to be played after the first sound file containing the first sound data is the second sound file containing the second sound data with the highest similarity to the first sound data. In a modified example, the sound file playback unit 218 may determine that the sound file to be played after the first sound file is the second sound file containing the second sound data whose similarity to the first sound data is greater than or equal to a predetermined value. Here, the predetermined value of similarity may be set to, for example, 0.6. If there are multiple candidates (choices) for the sound file to be played next, the sound file playback unit 218 may randomly select one from the multiple candidates using, for example, a random function, a random number table, or a random number generation program. By randomly selecting the sound file to be played next from multiple candidates in this way, the sound file playback unit 218 can reduce the regularity of the ambient sounds to be played.
[0062] In this embodiment, the sound file playback unit 218 determines the next sound file to be played based on its similarity to the last sound file played. In a modified example, the sound file playback unit 218 may determine the next sound file to be played based on its similarity to a plurality of sound files that were last played. For example, as shown in Figure 9(b), when sound file A and sound file B are played consecutively, the sound file playback unit 218 determines the next sound file to be played after sound file B based on its similarity to sound file A and sound file B. In this case, the sound file playback unit 218 may treat sound file A and sound file B as a single sound file AB and determine the next sound file to be played based on its similarity to this single sound file AB. Alternatively, the sound file playback unit 218 may determine the next sound file to be played based on the average or sum of the similarity between sound file A and sound file B.
[0063] Figure 10 shows an example of two sound files to be played simultaneously. The sound file playback unit 218 determines that the sound file to be played simultaneously with the first sound file containing the first sound data is the third sound file containing third sound data with a low similarity to the first sound data. When two sound files are played simultaneously, if the similarity between the two sound data is high, resonance or a phenomenon close to resonance may occur, resulting in an unnatural sound. Therefore, when playing two sound files simultaneously, it is preferable to select two sound files with low similarity.
[0064] If the sound file playback unit 218 initially selects sound file A, it may determine that sound file H, which contains the sound data H with the lowest similarity to sound data A, will be played simultaneously. If the sound file playback unit 218 initially selects sound file H, it may determine that sound file E, which contains the sound data E with the lowest similarity to sound data H, will be played simultaneously. In this way, the sound file playback unit 218 may determine that the sound file to be played simultaneously with the first sound file containing the first sound data is the third sound file containing the third sound data with the lowest similarity to the first sound data.
[0065] The sound file playback unit 218 may also determine that the sound file to be played simultaneously with the first sound file is a second sound file containing second sound data whose similarity to the first sound data is less than or equal to a predetermined value. Here, the predetermined value of similarity may be set to, for example, 0.4. If there are multiple candidate (option) sound files to be played simultaneously, the sound file playback unit 218 may select one sound file from the multiple candidates using, for example, a random function. By randomly selecting the sound file to be played simultaneously from multiple candidates in this way, the sound file playback unit 218 can reduce the regularity of the ambient sounds to be played.
[0066] When the sound file playback unit 218 plays sound file A and sound file H simultaneously, the next sound file to be played will be a sound file containing sound data with a high similarity between the combined data (synthesized sound signal) of sound data A and sound data H. For this reason, it is preferable that the similarity recording unit 234 records the similarity between the data obtained by combining the two sound data and another sound data. For this reason, in the sound file generation device 100, it is preferable that the similarity derivation unit 114 derives the similarity between the data obtained by combining the two sound data and another sound data in advance and records it in the similarity recording unit 124.
[0067] The present disclosure has been described above based on embodiments. These embodiments are illustrative, and it will be understood by those skilled in the art that various modifications are possible in combinations of their components and processing processes, and that such modifications are also within the scope of the present disclosure. In the embodiments, the similarity derivation unit 114 divided the sound data 104 into 3-second intervals, but it may be divided into shorter intervals (for example, 1 second) or longer intervals. Also, in the embodiments, the sound data 104 was ambient sound data, but it may be other types of sound data.
[0068] In this embodiment, the similarity derivation unit 114 derives the similarity between the entirety of the two divided sound data, but in a modified example, the similarity between parts of the two divided sound data may be derived. Figure 11 is a diagram illustrating the similarity derivation method in a modified example. In the modified example, the divided sound data is given a front part and a back part. In this example, the section from the start point ts to t1 of the divided sound data is given as the front part, and the section from t2 to the end point te is given as the back part. The section from the start point ts to the end point te is 3 seconds long. The section from the start point ts to t1 and the section from t2 to the end point te are set to be equal in length, for example, 1 second long. Note that t1 and t2 may be at the same timing, and the front part and the back part may overlap.
[0069] In the modified version, the similarity derivation unit 114 derives the similarity between the latter part of the first divided sound data and the former part of the second divided sound data, and also derives the similarity between the latter part of the second divided sound data and the former part of the first divided sound data. Suppose the similarity between the latter part of the first divided sound data and the former part of the second divided sound data is high, for example, 0.8, while the similarity between the latter part of the second divided sound data and the former part of the first divided sound data is low, for example, 0.2. This means that there is no problem in playing the second divided sound data after the first divided sound data, but playing the first divided sound data after the second divided sound data may cause discomfort to the user. Thus, in the modified version, the similarity derivation unit 114 derives the similarity near the joint between the two divided sound data and records it in the similarity recording unit 124. According to the modified version, the sound file playback unit 218 in the user's information processing device 10 can appropriately determine the playback order of multiple sound files.
[0070] This disclosure may include the following embodiments: [Item 1] A sound file generation device comprising a circuit configured to, the circuit, acquires sound data, divides the acquired sound data into a plurality of divided sound data, generates a plurality of sound files containing the divided sound data, derives the similarity between two divided sound data, and records the derived similarity between the two divided sound data. [Item 2] The sound file generation device according to Item 1, wherein the circuit cuts out the sound data at zero-crossing points and divides it into a plurality of divided sound data. [Item 3] The sound file generation device according to Item 1, wherein the circuit applies a fade process to the divided sound data. [Item 4] The sound file generation device according to Item 1, wherein the circuit divides the sound data into a plurality of divided sound data of substantially the same time length. [Item 5] The sound file generation device according to Item 1, wherein the sound data is ambient sound data obtained by recording ambient sounds. [Item 6] The sound file generation device according to Item 1, wherein the circuit derives the similarity for all combinations of two divided sound data in a plurality of divided sound data. [Item 7] The sound file generation device according to Item 1, wherein the circuit derives the cross-correlation coefficient between two divided sound data. [Item 8] A method for generating a sound file, comprising: acquiring sound data; dividing the acquired sound data into a plurality of divided sound data to generate a plurality of sound files containing the divided sound data; deriving the similarity between two divided sound data; and recording the derived similarity between the two divided sound data. [Item 9] A recording medium on which a program executed on a computer is recorded, wherein the program enables the computer to perform the following functions: acquiring sound data; dividing the acquired sound data into a plurality of divided sound data to generate a plurality of sound files containing the divided sound data; deriving the similarity between two divided sound data; and recording the derived similarity between the two divided sound data.[Item 10] An audio file playback device that plays an audio file containing sound data, comprising a circuit configured to the following: the circuit records the similarity between two sound data, concatenates and plays multiple audio files based on the recorded similarity between the two sound data, and outputs the sound to be played. [Item 11] The audio file playback device according to Item 10, wherein the circuit records the similarity for all combinations of two sound data in multiple sound data contained in multiple audio files. [Item 12] The audio file playback device according to Item 10, wherein the circuit determines the audio file to be played after the first audio file containing first sound data to be the second audio file containing second sound data with a high similarity to the first audio data. [Item 13] The audio file playback device according to Item 12, wherein the circuit determines the audio file to be played after the first audio file to be the second audio file, on the condition that a predetermined time has elapsed since the previous playback end timing of the second audio file. [Item 14] The sound file playback device according to Item 12, wherein the circuit does not determine the second sound file to be played after the first sound file if a predetermined time has not elapsed since the previous end of playback of the second sound file. [Item 15] The sound file playback device according to Item 10, wherein the circuit determines the third sound file to be played simultaneously with the first sound file containing the first sound data to be a third sound file containing third sound data with a low similarity to the first sound data. [Item 16] A method for playing sound files, comprising: recording the similarity between two sound data; concatenating and playing multiple sound files based on the recorded similarity between the two sound data; and outputting the sound to be played. [Item 17] A recording medium on which a program executed in a computer is recorded, wherein the program enables the computer to: record the similarity between two sound data; concatenate and play multiple sound files based on the recorded similarity between the two sound data; and output the sound to be played.
[0071] This disclosure can be used in the technical field of generating and / or playing audio files.
[0072] 1... Information processing system, 10... Information processing device, 100... Sound file generation device, 102... Processing device, 104... Sound data, 110... Sound data acquisition unit, 112... Sound file generation unit, 114... Similarity derivation unit, 120... Recording device, 122... Sound file recording unit, 124... Similarity recording unit, 200... Processing unit, 202... Communication unit, 210... Game execution unit, 212... Game image generation unit, 214... Game sound generation unit, 216... Voice file playback unit, 218... Sound file playback unit, 220... Output unit, 230... Recording device, 232... Sound file recording unit, 234... Similarity recording unit.
Claims
1. A sound file generation device comprising: a sound data acquisition unit that acquires sound data; a sound file generation unit that divides the acquired sound data into multiple divided sound data and generates multiple sound files containing the divided sound data; a similarity derivation unit that derives the similarity between two divided sound data; and a recording unit that records the derived similarity between the two divided sound data.
2. The sound file generation device according to claim 1, characterized in that the sound file generation unit cuts out sound data at zero-crossing points and divides it into multiple divided sound data.
3. The sound file generation device according to claim 1, characterized in that the sound file generation unit applies a fade process to the divided sound data.
4. The sound file generation device according to claim 1, characterized in that the sound file generation unit divides sound data into a plurality of segmented sound data having substantially the same time length.
5. The sound file generation device according to claim 1, characterized in that the sound data is ambient sound data obtained by recording ambient sounds.
6. The sound file generation device according to claim 1, characterized in that the similarity derivation unit derives a similarity for all combinations of two divided sound data in a plurality of divided sound data.
7. The sound file generation device according to claim 1, characterized in that the similarity derivation unit derives a cross-correlation coefficient between two divided sound data.
8. A method for generating sound files, comprising the steps of: acquiring sound data; dividing the acquired sound data into multiple divided sound data to generate multiple sound files containing the divided sound data; deriving the similarity between two divided sound data; and recording the derived similarity between the two divided sound data.
9. A program to implement the following functions on a computer: acquiring sound data, dividing the acquired sound data into multiple segmented sound data and generating multiple sound files containing the segmented sound data, deriving the similarity between two segmented sound data, and recording the derived similarity between the two segmented sound data.
10. An audio file playback device for playing an audio file containing sound data, comprising: a similarity recording unit for recording the similarity between two audio data; an audio file playback unit for concatenating and playing a plurality of audio files based on the similarity between the two audio data recorded in the similarity recording unit; and an output unit for outputting the sound to be played.
11. The sound file playback device according to claim 10, characterized in that the similarity recording unit records the similarity for all combinations of two sound data in multiple sound data contained in multiple sound files.
12. The sound file playback device according to claim 10, characterized in that the sound file playback unit determines the sound file to be played after the first sound file containing the first sound data to be a second sound file containing second sound data which has a high degree of similarity to the first sound data.
13. The sound file playback device according to claim 12, characterized in that the sound file playback unit determines the second sound file to be played after the first sound file, on the condition that a predetermined time has elapsed since the end of the previous playback of the second sound file.
14. The sound file playback device according to claim 12, characterized in that the sound file playback unit does not determine the second sound file to be played after the first sound file if a predetermined time has not elapsed since the previous playback end time of the second sound file.
15. The sound file playback device according to claim 10, characterized in that the sound file playback unit determines that the sound file to be played simultaneously with the first sound file containing the first sound data is a third sound file containing third sound data with a low similarity to the first sound data.
16. A method for playing an audio file, comprising the steps of: recording the similarity between two audio data; concatenating and playing multiple audio files based on the similarity between the two recorded audio data; and outputting the sound to be played.
17. A program to implement the following functions on a computer: a function to record the similarity between two audio data; a function to concatenate and play multiple audio files based on the similarity between the two recorded audio data; and a function to output the sound to be played.