Sound variation generating device, sound variation generating method, and storage medium

The sound variation generation device uses a machine learning model to automate the creation of sound variations, addressing inefficiencies in manual sound production by generating high-quality, scenario-specific audio for games and other content.

WO2026009294A1PCT designated stage Publication Date: 2026-01-08SONY INTERACTIVE ENTERTAINMENT LLC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/023810
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-01
Publication Date
2026-01-08

AI Technical Summary

Technical Problem

Conventional methods for generating sound variations in games and other entertainment content require significant manual effort and are inefficient, particularly for creating numerous variations of sounds based on different game scenarios or situations.

Method used

A sound variation generation device utilizing a machine learning model to automatically generate sound variations based on input sounds, with a variation mode designation unit specifying the desired variation type, and a similarity evaluation unit ensuring high-quality output.

Benefits of technology

Efficiently produces a large number of high-quality sound variations tailored to specific game scenarios or attributes, reducing manual labor and enhancing the realism of game audio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024023810_08012026_PF_FP_ABST
    Figure JP2024023810_08012026_PF_FP_ABST
Patent Text Reader

Abstract

A sound variation generating device 1 includes: a sound variation generating unit 12 that uses a sound variation generating model M, which is a machine learning model, to generate, from an input sound IS, an output sound OS, which is a variation of the input sound IS; and a variation mode specifying unit 13 that specifies a mode of the variation to be generated by the sound variation generating unit 12 from the input sound IS. The input sound IS and the output sound OS are game sounds. The variation mode specifying unit 13 specifies a situation during a game in which the output sound OS is to be used.
Need to check novelty before this filing date? Find Prior Art

Description

Sound variation generation device, sound variation generation method, and storage medium

[0001] The present disclosure relates to a technique for generating variations of sounds in games and the like.

[0002] Patent Document 1 discloses a method for effectively utilizing sound to enhance the dramatic effects in racing games. Various sound data, such as dramatic sounds and sound effects, are prepared in advance, and appropriate sound data is selected according to the progress of the racing game and output from a speaker.

[0003] JP 2009-213828 A

[0004] As in Patent Document 1, in the production of content for games, movies, etc., conventionally, a wide variety of sounds or tones had to be prepared manually. In particular, for a single sound (for example, the footsteps of a specific character wearing specific shoes), it was often necessary to prepare a large number of variations depending on the scene or situation in which the sound was played in the game, movie, etc., requiring a huge amount of man-hours.

[0005] The present disclosure has been made in light of these circumstances, and aims to provide a sound variation generation device and the like that can efficiently generate sound variations.

[0006] In order to solve the above problem, a sound variation generation device according to one aspect of the present disclosure includes a sound variation generation unit that uses a sound variation generation model, which is a machine learning model, to generate an output sound that is a variation of an input sound from the input sound, and a variation mode designation unit that designates the mode of variation that the sound variation generation unit should generate from the input sound.

[0007] According to this aspect, the sound variation generation unit can use the sound variation generation model to efficiently generate, from the input sound, an output sound with an appropriate variation according to the aspect specified by the variation aspect specification unit.

[0008] Another aspect of the present disclosure is a sound variation generation method, which uses a sound variation generation model that is a machine learning model to generate, from an input sound, an output sound that is a variation of the input sound, and specifies aspects of the variation to be generated from the input sound.

[0009] Yet another aspect of the present disclosure is a storage medium storing a sound variation generation program that causes a computer to generate, from an input sound, an output sound that is a variation of the input sound, using a sound variation generation model that is a machine learning model, and to specify aspects of the variation to be generated from the input sound.

[0010] Any combination of the above components, or any conversion of these expressions into methods, devices, systems, recording media, computer programs, etc., are also encompassed within the present disclosure.

[0011] 1 is a schematic functional block diagram of a sound variation generation device according to a first embodiment;FIG. 2 is a schematic functional block diagram of a sound variation generation device according to a second embodiment;FIG.

[0012] Hereinafter, with reference to the drawings, a detailed description of embodiments of the present disclosure (hereinafter also referred to as "embodiments") will be given. In the description and / or drawings, identical or equivalent components, members, processes, etc. will be designated by the same reference numerals, and redundant description will be omitted. The scale and shape of each part shown in the drawings are set for convenience to simplify the description and should not be interpreted as limiting unless otherwise specified. The embodiments are merely examples and do not limit the scope of the present disclosure in any way. Not all features and combinations thereof presented in the embodiments are necessarily essential to the present disclosure. For convenience, the embodiments are presented broken down into components for each function and / or functional group that realizes the features. However, one component in an embodiment may actually be realized by a combination of multiple separate components, or multiple components in an embodiment may actually be realized by a single integrated component. Furthermore, although multiple embodiments and variants may be disclosed in parallel, any components of each embodiment and / or each variant may be combined in any manner as long as they do not interfere with each other's functions.

[0013] In this embodiment, the generation of sound variations in a game or computer game is exemplified in detail. However, the sound variation generation device according to the present disclosure is equally applicable to the generation of sound variations in entertainment content other than games, such as movies and music. Furthermore, the sound variation generation device according to the present disclosure is not limited to entertainment content, and is also useful in any other application in which an output sound that is a variation of an input sound is substantially automatically generated (e.g., the generation of output sound in the distribution of online advertisements, news, articles, etc.).

[0014] FIG. 1 is a schematic functional block diagram of a sound variation generation device 1 according to a first embodiment. The sound variation generation device 1 includes a sound input unit 11, a sound variation generation unit 12, a variation mode designation unit 13, a similarity evaluation unit 14, and a sound variation storage unit 15. Some of these functional blocks may be omitted as long as the sound variation generation device 1 can achieve at least some of the functions and / or effects described below. These functional blocks are realized by the cooperation of hardware resources, such as a computer's central processing unit, memory, input devices, output devices, and peripheral devices connected to the computer, and software executed using these resources. Regardless of the type and location of the computer, each of the above functional blocks may be realized by the hardware resources of a single computer or by a combination of hardware resources distributed across multiple computers. Furthermore, the game status acquisition unit 21 and sound playback unit 22, described below, may be realized by a game machine C, described below, a game server (not shown), or the like.

[0015] The functions of the components in the sound variation generation device 1 may be realized by circuitry or processing circuitry, including a general-purpose processor, an application-specific processor, an integrated circuit, an application-specific integrated circuit (ASIC), a central processing unit (CPU), a graphics processing unit (GPU), a conventional circuit, and / or a combination thereof, configured or programmed to realize the functions described herein. A processor is understood to mean a circuitry or processing circuitry including transistors and other circuit elements. A processor may also be a programmed processor that executes a program stored in a memory.

[0016] The circuits, units, and means herein are hardware that is programmed or configured to perform the described functions, which may be any hardware disclosed herein or known hardware that is programmed or configured to perform the described functions.

[0017] If the hardware is a processor that is interpreted as a type of circuitry, the circuitry, means, or unit may be a combination of hardware and software used to configure the hardware and / or processor.

[0018] 1, the "During Production" area on the left schematically shows elements related to the generation of sound variations during game production, and the "During Play" area on the right schematically shows elements related to the playback of sounds during game play. In the example of FIG. 1, the sound variation generation device 1 is primarily used during the "production" stage.

[0019] The sound input unit 11 inputs an input sound IS that serves as the basis for generating a variation (output sound OS) by the sound variation generation device 1. As shown schematically in Fig. 1, the sound input unit 11 may input the input sound IS through manual operation of a computer CP by a sound designer D, or may input any variation stored in a sound variation storage unit 15 (described below) as the input sound IS (i.e., a further variation is generated from this variation), or may input an existing game sound 3 (not shown in the figure but described below) as the input sound IS.

[0020] Since this embodiment relates to the generation of sound variations in a game, the input sound IS and its variation, the output sound OS, are game sounds. The game sounds for which variations are generated in this embodiment are any auditory stimuli, such as sounds, voices, sounds, vibrations, waves, music, etc., that can be reproduced in the game. For example, the game sounds may be sound effects that can be reproduced in the game in response to game operations by player P or other in-game situations.

[0021] Examples of sound effects include footsteps, gunfire, impact sounds, vehicle driving sounds, crowd noises, and environmental sounds. The manner in which these sound effects are reproduced needs to be changed depending on the situation in the game. For example, the sound of footsteps of a specific game character wearing specific shoes will differ depending on the character's walking speed and manner, the material of the floor or ground on which the character is walking, and the weather and acoustic characteristics of the location where the character is walking. Thus, even for essentially one sound, a huge number of variations need to be created.

[0022] Conventionally, the production of such sound variations has been highly dependent on human resources such as sound designers D, and has been inefficient. As will be described later, this embodiment aims to automate and streamline sound variation production by utilizing a sound variation generation model M. Note that this embodiment is applicable not only to generating variations of relatively short sounds such as sound effects, but also to generating variations of relatively long sounds such as music (for example, playing the same music in different tunes depending on the scene in the game).

[0023] The sound variation generation unit 12 uses a sound variation generation model M, which is a machine learning model, to generate an output sound OS, which is a variation of an input sound IS, from the input sound IS input by the sound input unit 11. The variation mode designation unit 13 designates the mode of variation that the sound variation generation unit 12 (sound variation generation model M) should generate from the input sound IS. In this way, the sound variation generation unit 12 inputs the input sound IS input by the sound input unit 11 and the variation mode designated by the variation mode designation unit 13 to the sound variation generation model M, and automatically generates an output sound OS of a variation that matches the mode from the input sound IS.

[0024] The variation mode designation unit 13 may designate a situation in the game in which the output sound OS should be used as the mode of variation to be generated from the input sound IS. In this case, the sound variation generation unit 12 generates the output sound OS from the input sound IS in accordance with the situation in the game designated by the variation mode designation unit 13. Examples of the situation in the game include, but are not limited to, the time in the game (such as the time of day or season), a location in the game, a scene or environment in the game, the progress or achievement of the game, the state or action of a game character or other in-game object, an operation by the player P, and input from the game machine C or the game server, and include any event or state that may bring about a change in the game or that may appear as a change in the game.

[0025] As an example, suppose the input sound IS input to the sound variation generation unit 12 (sound variation generation model M) is "footsteps of a specific game character walking indoors on a wooden floor." The variation mode designation unit 13 then designates, for example, "a situation in which the game character is running on muddy ground outdoors in the rain" as a game situation. In this case, the sound variation generation unit 12 (sound variation generation model M) generates, from the input sound IS, "footsteps of a specific game character running on muddy ground outdoors in the rain" as the output sound OS in accordance with the designated game situation (different from the game situation corresponding to the input sound IS).

[0026] As described above, the sound variation generation unit 12 functions as a sound adjustment unit that generates the output sound OS by adjusting or customizing the input sound IS in accordance with the game situation specified by the variation mode designation unit 13. While such adjustment processing can be performed by a sound mixer or the like that does not utilize machine learning or artificial intelligence, it is preferably performed by a sound variation generation model M that is a machine learning (ML) model or an artificial intelligence (AI) model, as in the example of FIG. 1 .

[0027] The sound variation generation model M is preferably a trained model or a supervised learning model that has undergone training using comprehensive learning data or training data so that it can appropriately generate an output sound OS corresponding to a game situation (specified by the variation mode designation unit 13) different from the input sound IS. For example, the training data may include a set of sound data corresponding to the input sound IS or the output sound OS and one or more metadata representing the game situation (e.g., one of the examples listed above) to which the sound data corresponds. The sound variation generation model M, which has trained a large amount of training data in which sound data and game situation data are paired, can recognize the trends and characteristics of sound data required for each game situation, and can therefore appropriately customize the input sound IS to an output sound OS corresponding to the game situation specified by the variation mode designation unit 13.

[0028] A large amount of learning data, in which sound data and game situation data are paired as described above, can be acquired from existing games that are different from the game for which sound variations are to be generated by the sound variation generation device 1. For this reason, in the example of Fig. 1, existing game sound 3 is schematically shown as learning data for the sound variation generation model M. In this case, the sound variation generation model M is a trained model that uses game sound variations in an existing game as learning data.

[0029] The sound variation generation model M is preferably trained using learning data from existing games of the same or similar type (or type, genre, or tone) as the game for which sound variations are to be generated. For example, the desired sound tendencies may differ significantly between different types of games, such as a fighting game with a dark atmosphere and a social game with a bright atmosphere. Therefore, by limiting the learning data to games of the same or similar type, the accuracy of sound variation generation by the sound variation generation model M can be improved. In this case, it is preferable to create a sound variation generation model M for each type of game in which the output sound OS is to be used.

[0030] In the above, the variation mode designation unit 13 has exemplified the in-game situations in which the output sound OS should be used, but other variation modes may also be designated. For example, the variation mode may be designated based on various sound quality evaluation indices (loudness, sharpness, fluctuation intensity, roughness, etc.), or the variation mode may be designated based on the duration, sound pressure (loudness), pitch (interval), timbre, frequency band, playback pattern, emotion, tone, etc. that the output sound OS should have. In this case, the sound variation generation unit 12 generates the output sound OS from the input sound IS in accordance with the sound quality evaluation indices and other sound attributes designated by the variation mode designation unit 13.

[0031] In this case, the training data for the sound variation generation model M uses a set of sound data corresponding to the input sound IS or the output sound OS and one or more metadata representing sound quality evaluation indices and other sound attributes indicated by the sound data. Note that the metadata representing the sound attributes can be said to be inherent in the sound data itself, and therefore may be extracted from the sound data to be trained when training the sound variation generation model M. The sound variation generation model M, which has trained a large amount of training data in which sound data and sound attribute data are paired in this way, can appropriately customize the input sound IS to the output sound OS in accordance with the sound attributes specified by the variation mode specification unit 13.

[0032] In this case, the sound variation generation model M is also preferably trained using learning data from an existing game of the same or similar type (or type, genre, or tone) as the game for which sound variations are to be generated. In this case, the sound variation generation model M is preferably created for each type of game in which the output sound OS is to be used.

[0033] 1, the variation mode designation unit 13 may input or designate a variation mode for generating an output sound OS to the sound variation generation model M through manual operation of a computer CP by a sound designer D. In this case, manual work by the sound designer D remains, but the actual output sound OS is not produced; instead, the sound designer D only needs to designate the variation mode, which significantly reduces the human resources required compared to conventional sound variation production. Alternatively, the variation mode designation unit 13 may autonomously select a mode (game situation or sound attribute) for which no variation has been created in the target game, and input or designate it to the sound variation generation model M.

[0034] In the sound variation generation model M that generates output sounds OS from input sounds IS as described above, the number of input sounds IS as inputs and the number of output sounds OS as outputs are both arbitrary. That is, where m and n are arbitrary, independent natural numbers, the sound variation generation model M generates n output sounds OS from m input sounds IS in accordance with substantially n aspects (e.g., n game situations or n sets of sound attributes) specified by the variation aspect specification unit 13.

[0035] When m is 2 or greater, the sound variation generation unit 12 uses the sound variation generation model M to generate, from a plurality (m) of input sounds IS that are variations of one another, n output sounds OS that are different from any of the plurality (m) of input sounds IS and that are variations that match the n aspects specified by the variation aspect designation unit 13. By using a plurality (m) of input sounds IS in this way, the sound variation generation model M can grasp sound attributes that are common to the plurality (m) of input sounds IS (in other words, attributes that are commonly required for the variations of the sound), and can appropriately customize the sound according to the aspect specified by the variation aspect designation unit 13 while maintaining these attributes.

[0036] Furthermore, when n is greater than m, the sound variation generation unit 12 uses the sound variation generation model M to generate a plurality (n) of output sounds OS that are more than the m input sounds IS and are variations that match the n aspects specified by the variation aspect specification unit 13. In this case, since a large number (n) of output sounds OS are generated from a small number (m) of input sounds IS, the efficiency of sound variation production can be significantly improved.

[0037] The similarity evaluation unit 14, which may constitute a part of the sound variation generation unit 12, may evaluate the similarity between n candidates for output sounds OS generated from m input sounds IS by the sound variation generation model M and the m input sounds IS. The similarity evaluation unit 14 or the sound variation generation unit 12 then finally generates the candidates evaluated as being similar to the m input sounds IS as the output sounds OS.

[0038] Among the n candidates to be evaluated by the similarity evaluation unit 14, n 0 pieces (n 0 ≦n) is evaluated as similar to m input sounds IS, 0 The output sounds OS are output as appropriate sound variations from the similarity evaluation unit 14 or the sound variation generation unit 12 and stored in the sound variation storage unit 15. On the other hand, among the n candidates to be evaluated by the similarity evaluation unit 14, n-n candidates that are evaluated as not similar to the m input sounds IS are 0 The candidates are discarded as not being appropriate sound variations. The similarity evaluation result by the similarity evaluation unit 14 may be used for re-training the sound variation generation model M.

[0039] The similarity evaluation unit 14 may evaluate the similarity between m input sounds IS and n output sounds OS (candidates) based on the attributes of the sounds (e.g., sound quality evaluation indices such as loudness and fluctuation intensity, duration, and pitch change pattern). As described above, when there are multiple input sounds IS (m≧2), attributes common to the multiple input sounds IS can be used as the basis for the similarity evaluation, thereby improving the evaluation accuracy. As a result, high-quality sound variations (output sounds OS) can be generated and stored in the sound variation storage unit 15.

[0040] The criteria for the similarity evaluation of sounds by the similarity evaluation unit 14 may be adjustable by the criteria adjustment unit 141. For example, the range of sounds that the player P perceives as similar may vary depending on the type of game. Therefore, the criteria adjustment unit 141 may automatically adjust the criteria for the similarity evaluation of sounds depending on the type of game. Alternatively, the criteria adjustment unit 141 may adjust the criteria through manual operation of the computer CP by the sound designer D (for example, an adjustment tool such as a slider for adjusting the criteria may be provided on the user interface of the computer CP). Furthermore, like the quality evaluation unit 16 described below, the similarity evaluation unit 14 may evaluate the sound quality and other qualities of the candidate output sound OS, as well as the suitability of the candidate output sound OS in terms of compliance, in addition to or instead of the similarity of sounds, and store only those that meet predetermined evaluation criteria in the sound variation storage unit 15.

[0041] At least some of the sound variations stored in the sound variation storage unit 15 as described above may be provided to the sound input unit 11 as input sounds IS in order to generate further different sound variations. Also, at least some of the sound variations stored in the sound variation storage unit 15 may be stored as existing game sounds 3, which are learning data for a sound variation generation model M for generating sound variations in other games.

[0042] As described above, during "production" on the left side of Fig. 1, a huge number of sound variations are generated corresponding to all game situations etc. in which the sound for which variations are to be generated can be played, and are stored in the sound variation storage unit 15. During "play" on the right side of Fig. 1, when player P plays the game, a sound variation corresponding to the game situation etc. is selected and played by the speaker SP etc.

[0043] For example, a player P plays a game through a remote controller R that can communicate in real time with a game machine C and a game server (not shown). Game images are displayed on a television set or the like, and game sounds are reproduced by speakers SP or the like. The speakers SP may be integrated into the television set.

[0044] A game status acquisition unit 21, which may be constituted by the game machine C or a game server (not shown), acquires the game status (including operation information of the remote controller R by the player P). A sound playback unit 22, which may also be constituted by the game machine C or a game server (not shown), reads out sounds of variations corresponding to the game status acquired by the game status acquisition unit 21 from the sound variation storage unit 15 and plays them on the speaker SP.

[0045] FIG. 2 is a schematic functional block diagram of a sound variation generation device 1 according to a second embodiment. Components similar to those in the first embodiment of FIG. 1 are assigned the same reference numerals, and redundant explanations will be omitted. Furthermore, the existing game sounds 3 shown in FIG. 1 are also useful as learning data for the sound variation generation model M in the second embodiment, but are not shown in FIG. 2. As in FIG. 1, elements "during production" are shown schematically on the left side of FIG. 2, and elements "during play" are shown schematically on the right side. In contrast to the example of FIG. 1, in the example of FIG. 2, the sound variation generation device 1 is primarily used "during play."

[0046] During "production" in this embodiment, the reference sound creation unit 31 creates one or more (typically a small number of) reference sounds that serve as the basis for sound variations created substantially in real time during "play," as will be described later. The reference sound creation unit 31 may create the reference sounds, for example, through manual operation of a computer CP by a sound designer D. The reference sounds created by the reference sound creation unit 31 are stored in a reference sound storage unit 32.

[0047] In this embodiment, during "play," the sound variation generation device 1 generates appropriate variations (output sound OS) according to the game situation in approximately real time based on the reference sound (input sound IS) stored in the reference sound storage unit 32.

[0048] First, a game situation acquisition unit 21 (which also functions as the variation mode designation unit 13, as described below), which may be configured by the game machine C or a game server (not shown), acquires the situation during the game (including operation information of the remote controller R by the player P). The sound variation generation device 1 reads a reference sound corresponding to the game situation acquired by the game situation acquisition unit 21 from the reference sound storage unit 32 as an input sound IS. This reference sound serves as a base or standard for the sound to be reproduced in the game situation acquired by the game situation acquisition unit 21, but is not a variation optimized for that game situation. Therefore, as in the first embodiment of FIG. 1 , the sound variation generation unit 12 (sound variation generation model M) is used to generate a variation (output sound OS) from the reference sound (input sound IS) optimized for the game situation.

[0049] While the player P is playing the game, the sound variation generation unit 12 uses the trained sound variation generation model M to generate, in approximately real time, an output sound OS, which is a variation of the input sound IS, from the input sound IS input from the reference sound storage unit 32. Specifically, the sound variation generation unit 12 and / or the sound variation generation model M generate, in approximately real time, a variation (output sound OS) from the reference sound (input sound IS) in accordance with the game situation (variation mode) acquired by the game situation acquisition unit 21 functioning as the variation mode designation unit 13.

[0050] The quality evaluation unit 16, which may constitute a part of the sound variation generation unit 12, may evaluate the quality of the candidates for the output sound OS generated by the sound variation generation model M from the input sound IS. The quality evaluation unit 16 or the sound variation generation unit 12 then ultimately generates or outputs the candidates evaluated as having good quality as the output sound OS. On the other hand, candidates not evaluated as having good quality by the quality evaluation unit 16 are discarded.

[0051] If all candidates fail to pass the "gate" of the quality evaluation unit 16 and are discarded, it becomes necessary to redo the generation of variations by the sound variation generation model M (or, although undesirable, the reference sound (input sound IS) may be reproduced as is by the sound reproduction unit 22). To avoid such a situation, it is preferable that the sound variation generation model M generates multiple (preferably more than the input sound IS) candidates for the output sound OS in parallel based on one or multiple (typically a small number) input sounds IS. If at least one of these multiple candidates "passes" or "survive" the quality evaluation unit 16, there is no need to redo the generation of variations. Such a quality evaluation result by the quality evaluation unit 16 may be used for offline relearning of the sound variation generation model M.

[0052] The quality evaluation by the quality evaluation unit 16 may be based on the similarity between the input sound IS and the candidate output sound OS, as with the similarity evaluation unit 14. Furthermore, in addition to or instead of the similarity of the sounds, the quality evaluation unit 16 may evaluate the sound quality and other qualities of the candidate output sound OS, or the suitability of the candidate output sound OS from the viewpoint of compliance, etc.

[0053] The criteria for evaluating sound quality by the quality evaluation unit 16 may be adjustable by the criteria adjustment unit 161. For example, the range of sound quality that the player P can tolerate may differ depending on the type of game. Therefore, the criteria adjustment unit 161 may automatically adjust the criteria for evaluating sound quality depending on the type of game.

[0054] The output sound OS (a variation suited to the game situation) that has been evaluated as "acceptable" by the quality evaluation unit 16 is sent as is to the sound playback unit 22 and played back by a speaker SP or the like.

[0055] As described above, the sound variation generation device 1 according to this embodiment generates appropriate variations (output sounds OS) in substantially real time according to the game situation, based on the reference sounds (input sounds IS) stored in the reference sound storage unit 32. However, since it takes time for the sound variation generation device 1 to generate variations and to evaluate or verify them, it is difficult to achieve completely real-time processing.

[0056] Therefore, it is preferable that the sound variation generation device 1 generates in advance variations that may be needed for the immediate or near future game situation that is expected in light of the current game situation acquired by the game situation acquisition unit 21.

[0057] The present disclosure has been described above based on the embodiments. Various modifications are possible to the combinations of the components and processes in the exemplary embodiments, and it will be obvious to those skilled in the art that such modifications are included within the scope of the present disclosure.

[0058] For example, while Fig. 1 shows an example in which the sound variation generation device 1 is primarily used "during production," and Fig. 2 shows an example in which the sound variation generation device 1 is primarily used "during play," the sound variation generation device according to the present disclosure may have a hybrid configuration between Fig. 1 and Fig. 2. A sound variation generation device with such a hybrid configuration may create variations in advance for some sounds (e.g., sounds with high importance, playback frequency, complexity, etc.) "during production," as shown in Fig. 1, and create variations for the remaining sounds in approximately real time "during play," as shown in Fig. 2.

[0059] The configuration, operation, and function of each device and method described in the embodiments can be realized by hardware resources, software resources, or a combination of hardware and software resources. Examples of hardware resources include processors, ROM, RAM, and various integrated circuits. Examples of software resources include operating systems, applications, and other programs.

[0060] The present disclosure may be expressed as follows:

[0061] Item 1: A sound variation generation device comprising circuitry configured to perform the following: generating, from an input sound, an output sound that is a variation of the input sound using a sound variation generation model that is a machine learning model; and specifying an aspect of the variation to be generated from the input sound. Item 2: The sound variation generation device of item 1, in which the input sound and the output sound are game sounds. Item 3: The sound variation generation device of item 2, in which specifying the aspect of the variation specifies a situation in a game in which the output sound should be used. Item 4: The sound variation generation device of item 3, in which generating the output sound generates the output sound from the input sound in accordance with a situation in the game that is specified by specifying the aspect of the variation during game play. Item 5: The sound variation generation device of item 4, in which generating the output sound includes evaluating the quality of candidates for the output sound generated from the input sound, and generating the candidate evaluated to be of good quality as the output sound. Item 6: The sound variation generation device according to any one of items 2 to 5, wherein the sound variation generation model is a trained model that uses variations of game sounds in existing games as training data. Item 7: The sound variation generation device according to any one of items 2 to 6, wherein the sound variation generation model is created for each type of game in which the output sound is to be used. Item 8: The sound variation generation device according to any one of items 1 to 7, wherein generating the output sound uses the sound variation generation model to generate a plurality of output sounds that are more than the input sound and that match the specified aspects by specifying the aspects of the variation.Item 9: The sound variation generation device according to any one of items 1 to 8, wherein generating the output sound uses the sound variation generation model to generate, from a plurality of input sounds that are variations of one another, the output sound being a variation that is different from any of the plurality of input sounds and that matches the specified aspect by specifying a mode of the variation. Item 10: The sound variation generation device according to item 8 or 9, wherein generating the output sound includes evaluating the similarity of a candidate output sound generated from the input sound to the input sound, and generating the candidate evaluated to be similar to the input sound as the output sound. Item 11: A sound variation generation method that performs the steps of: using a sound variation generation model that is a machine learning model to generate, from an input sound, an output sound that is a variation of the input sound; and specifying the aspect of the variation to be generated from the input sound. Item 12: A storage medium that stores a sound variation generation program that causes a computer to perform the steps of: using a sound variation generation model that is a machine learning model to generate, from the input sound, an output sound that is a variation of the input sound; and specifying the aspect of the variation to be generated from the input sound.

[0062] The present disclosure relates to a technique for generating variations of sounds in games and the like.

[0063] 1 sound variation generation device, 11 sound input unit, 12 sound variation generation unit, 13 variation mode designation unit, 14 similarity evaluation unit, 15 sound variation storage unit, 16 quality evaluation unit, 21 game status acquisition unit, 22 sound playback unit.

Claims

1. A sound variation generation device comprising: a sound variation generation unit that generates an output sound that is a variation of an input sound from the input sound using a sound variation generation model that is a machine learning model; and a variation mode designation unit that designates the mode of variation that the sound variation generation unit should generate from the input sound.

2. The sound variation generating device according to claim 1, wherein the input sound and the output sound are game sounds.

3. The sound variation generating device according to claim 2, wherein the variation mode designation unit designates a situation in a game in which the output sound should be used.

4. A sound variation generation device as described in claim 3, wherein the sound variation generation unit generates the output sound from the input sound in accordance with the situation in the game specified by the variation mode specification unit while the game is being played.

5. The sound variation generation device of claim 4, wherein the sound variation generation unit is provided with a quality evaluation unit that evaluates the quality of the candidate output sound generated from the input sound, and generates the candidate evaluated to be of good quality as the output sound.

6. The sound variation generation device according to claim 2, wherein the sound variation generation model is a trained model that uses game sound variations in existing games as training data.

7. The sound variation generation device according to claim 2, wherein the sound variation generation model is created for each type of game in which the output sound is to be used.

8. The sound variation generation device of claim 1, wherein the sound variation generation unit uses the sound variation generation model to generate a plurality of output sounds that are more than the input sound and that are variations that match the aspect specified by the variation aspect specification unit.

9. The sound variation generation device of claim 1, wherein the sound variation generation unit uses the sound variation generation model to generate, from a plurality of input sounds that are variations of each other, an output sound that is different from any of the plurality of input sounds and that matches the aspect specified by the variation aspect specification unit.

10. A sound variation generation device as described in claim 8 or 9, wherein the sound variation generation unit is provided with a similarity evaluation unit that evaluates the similarity of the candidate output sound generated from the input sound to the input sound, and generates the candidate evaluated to be similar to the input sound as the output sound.

11. A sound variation generation method that performs the steps of: generating an output sound that is a variation of an input sound from the input sound using a sound variation generation model that is a machine learning model; and specifying the manner in which the variation should be generated from the input sound.

12. A storage medium storing a sound variation generation program that causes a computer to perform the following steps: generate an output sound that is a variation of an input sound from the input sound using a sound variation generation model that is a machine learning model; and specify the form of the variation to be generated from the input sound.

Citation Information

Patent Citations

  • Game audio processing method and device, storage medium and electronic device

    CN117563223A

  • Adaptive Audio Mixing

    JP2024503584A

  • Neural Synthesis of Sound Effects Using Deep Generative Models

    US20230390642A1

  • Interface customized generation of gaming music

    US20240029691A1

  • Arrangement generation method, arrangement generation device, and generation program

    WO2021166745A1