Information processing device, audio data generation device, acoustic output method, and audio data generation method

The information processing device optimizes sound sources in virtual spaces by adjusting volumes and distributions based on evaluation and positional relationships, addressing the challenge of balancing audio accuracy and user experience with reduced manual effort.

WO2026009420A1PCT designated stage Publication Date: 2026-01-08SONY INTERACTIVE ENTERTAINMENT LLC
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/024430
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-05
Publication Date
2026-01-08

AI Technical Summary

Technical Problem

Existing audio processing technologies struggle to maintain a balance between audio accuracy and user experience in virtual spaces, often leading to noisy or distracting sounds that hinder the perception of important audio cues, and require significant manual effort from content creators to adjust sound settings.

Method used

An information processing device with sound source status acquisition and control units that adjust sound sources based on evaluation results, synthesizing sounds while minimizing noise by disabling or reducing volumes of unnecessary sound sources, and optimizing sound distribution based on positional relationships and areas of sound sources.

Benefits of technology

This approach maintains audio accuracy and enhances user experience by reducing noise and effort, allowing for realistic sound generation with minimal manual adjustments by content creators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024024430_08012026_PF_FP_ABST
    Figure JP2024024430_08012026_PF_FP_ABST
Patent Text Reader

Abstract

In an information processing device 10, an output state acquisition unit 70 acquires the state of output audio collected by a microphone 18. A sound source state acquisition unit 72 arranges necessary sound sources in a virtual space to be displayed. A sound source control unit 74 adjusts at least one of the volume and the number of the sound sources on the basis of at least one of the state of the output audio and the state of the sound sources. An acoustic generation unit 76 combines sounds from the adjusted sound sources. An output unit 64 outputs data of the combined sounds to a speaker 20 or a storage device.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, audio data generating device, sound output method, and audio data generating method

[0001] The present invention relates to an information processing device that processes sound included in content, an audio data generating device that generates audio data, an audio output method, and an audio data generating method.

[0002] Advances in acoustic technology have made it possible to enjoy various forms of auditory effects in electronic content such as games and moving images. For example, by simulating how the voices of various objects and environmental sounds in a virtual space sound from the user's viewpoint and processing the sound source, it is possible to add impact and realism to video content (see, for example, Patent Document 1).

[0003] Japanese Patent Application Laid-Open No. 2020-18620

[0004] On the other hand, accurately reproducing all sounds occurring in a virtual space does not necessarily lead to an improved quality of user experience. For example, if animal objects in a virtual space each make sounds, a small number of them may create a sense of realism, but when many of them are gathered together, it may be perceived as noisy to the user. In some cases, these sounds may become distracting, making it difficult to hear the sounds that should be heard, or even causing distress. For this reason, maintaining a good balance between audio accuracy and the quality of the user experience is always an important challenge.

[0005] Furthermore, to enhance the sense of realism in virtual spaces, it is necessary to generate high-quality sounds that sound similar to those occurring in the real world. In the simulations mentioned above, sound propagation is taken into account, but if the sound before propagation deviates from reality, the resulting sound will sound unnatural. This requires manual recording of various real-world sounds and adjusting them to match the images, which is a huge effort for content creators.

[0006] The present invention has been made in consideration of these problems, and its purpose is to provide an audio processing technology that can maintain a good balance between audio accuracy and the quality of the user experience. Another purpose of the present invention is to provide a technology that can generate realistic audio while minimizing the work required by content creators.

[0007] One aspect of the present invention relates to an information processing device, which includes a sound source status acquisition unit that sets a sound source for a virtual space to be displayed, a sound source control unit that evaluates the state of a sound generated by the sound source settings and adjusts the sound source settings according to the evaluation result, and a sound generation unit that synthesizes and outputs a sound emitted from the adjusted sound source.

[0008] Another aspect of the present invention relates to a sound output method, which includes the steps of setting a sound source for a virtual space to be displayed, evaluating the state of a sound generated by the setting of the sound source and adjusting the setting of the sound source according to the evaluation result, and synthesizing and outputting a sound emitted from the adjusted sound source.

[0009] Yet another aspect of the present invention relates to an information processing device including a sound source state acquisition unit that sets sound sources for a virtual space to be displayed, and a sound generation unit that synthesizes and outputs sounds emitted from the set sound sources, wherein the sound source state acquisition unit adjusts the number of set sound sources depending on the area of ​​the sound generating sources.

[0010] Yet another aspect of the present invention relates to a sound output method, the sound output method including a step of setting sound sources for a virtual space to be displayed, and a step of synthesizing and outputting sounds emitted from the set sound sources, wherein the step of setting the sound sources includes adjusting the number of the set sound sources depending on the area of ​​the sound generating sources.

[0011] Yet another aspect of the present invention relates to a voice data generating device, comprising: a processing pattern storage unit for storing combinations of parameter values ​​of a plurality of processing means to be applied to voice; a voice processing unit for processing voice data designated by a user using the combinations; and an output unit for outputting the processed voice data.

[0012] Yet another aspect of the present invention relates to a voice data generation method, which includes the steps of: reading from a memory a combination of parameter values ​​of a plurality of processing means to be applied to voice; processing voice data designated by a user using the combination; and outputting the processed voice data.

[0013] Any combination of the above components and any transformation of the present invention between methods, devices, etc. are also valid aspects of the present invention.

[0014] According to the present invention, it is possible to maintain a good balance between audio accuracy and the quality of the user experience, and to generate realistic sound while minimizing the work required by content creators.

[0015] 1 is a diagram illustrating an example of the configuration of a content processing system in the present embodiment. FIG. 2 is a diagram illustrating the internal circuit configuration of an information processing device in the present embodiment. FIG. 3 is a diagram illustrating an aspect of adjusting a sound source according to an output situation in the present embodiment. FIG. 4 is a diagram illustrating the configuration of functional blocks of an information processing device in the present embodiment. FIG. 5 is a flowchart illustrating a processing procedure in which an information processing device outputs sound while adjusting a sound source as necessary in the present embodiment. FIG. 6 is a diagram illustrating an example of the structure of audio data stored in an audio data storage unit in the present embodiment. FIG. 7 is a diagram illustrating an example of a method of determining the priority of a sound source based on a positional relationship with a sound receiving point in the present embodiment. FIG. 8 is a diagram illustrating an aspect of adjusting the distribution of a point sound source according to the area of ​​the sound source in the present embodiment. FIG. 9 is a diagram illustrating the configuration of functional blocks of an information processing device in the present embodiment. FIG. 10 is a diagram illustrating an example of the structure of audio data stored in an audio data storage unit in the present embodiment. FIG. 11 is a flowchart illustrating a processing procedure in which an information processing device outputs sound while adjusting the distribution of a point sound source according to the area of ​​the sound generating source. FIG. 12 is a diagram illustrating the configuration of functional blocks of an audio data generating device in the present embodiment. FIG. 13 is a diagram illustrating an example of the data structure of a processing pattern stored in a processing pattern storage unit in the present embodiment. FIG. 14 is a diagram illustrating an example of a user interface image for audio processing generated by an image data generating unit in the present embodiment. FIG. 15 is a flowchart illustrating a processing procedure in which an audio data generating device processes and outputs source sound in the present embodiment. 1 is a diagram showing a functional block configuration of the voice data generating device according to the present embodiment;FIG. 2 is a diagram showing an example of a waveform of voice data analyzed by a waveform analysis unit according to the present embodiment;FIG.

[0016] 1 shows an example of the configuration of a content processing system according to the present embodiment. The content processing system 2 includes an information processing device 10 that processes content, an input device 14 that accepts user operations, a display device 16 that displays images of the content, speakers 20a and 20b that output the sound of the content, and a microphone 18 that collects the output sound. The information processing device 10 may also be connected to a network such as the Internet to enable communication with external devices such as servers.

[0017] The input device 14, display device 16, speakers 20a and 20b, and microphone 18 may each be connected to the information processing device 10 by a cable, or may be connected wirelessly via Bluetooth (registered trademark) or the like. Although the speakers 20a and 20b are shown integrally with the display device 16, they may have housings separate from the display device 16, and the number and arrangement of these housings are not limited. Furthermore, the external shape of each device is not limited to that shown in the drawings.

[0018] In this embodiment, the "user" who uses the content processing system 2 may be a person who appreciates or plays games on finished electronic content, or a person who creates electronic content. Hereinafter, the stage in which the former user uses the content processing system 2 may be referred to as the "appreciation phase," and the stage in which the electronic content creator uses the content processing system 2 may be referred to as the "production phase."

[0019] The specific configuration of the devices included in the content processing system 2 may differ depending on which phase this embodiment is used in. For example, in the viewing phase, the information processing device 10 is a game device, a personal computer, or the like, which processes electronic content selected by a user and supplies the resulting moving image data to the display device 16 and the sound data to the speakers 20a and 20b. In the production phase, the information processing device 10 is a computer or the like used for production, which accepts operations on the images and sound data of the content being produced, and supplies the image and user interface data to the display device 16 and the sound data to the speakers 20a and 20b.

[0020] The input device 14 is, for example, at least one of a game controller, a remote controller, a keyboard, a mouse, a joystick, a camera, etc. In the viewing phase, the input device 14 accepts user operations for content, such as command input for an electronic game. In the production phase, the input device 14 accepts user operations related to the production of programs, images, and sounds. The input device 14 sequentially transmits the accepted content to the information processing device 10.

[0021] The display device 16 may be a flat display such as a liquid crystal display or an organic EL display, or a head-mounted display worn by the user to display an image in front of the user's eyes, and appropriately processes and displays image data supplied from the information processing device 10. The speakers 20a and 20b are general speakers installed in the space where the user is present, and appropriately process and output audio data supplied from the information processing device 10. Note that the audio output device is not limited to speakers, and may also be earphones, headphones, or the like. The speakers 20a and 20b may be mounted on the information processing device 10 or the input device 14. Hereinafter, the speakers 20a and 20b will be collectively referred to as speakers 20.

[0022] The microphone 18 picks up sounds emitted by the speaker 20. The microphone 18 may also function as part of an input device that accepts voice commands from the user. The microphone 18 converts the sounds into electrical signals and transmits them to the information processing device 10. The microphone 18 may be mounted on the information processing device 10, the input device 14, or the display device 16.

[0023] 2 shows the internal circuit configuration of the information processing device 10. The information processing device 10 includes a CPU (Central Processing Unit) 23, a GPU (Graphics Processing Unit) 24, and a main memory 26. These components are interconnected via a bus 28. An input / output interface 30 is also connected to the bus 28. Connected to the input / output interface 30 are a communication unit 32 including a peripheral device interface such as a USB or a network interface for a wired or wireless LAN, a storage unit 34 such as a hard disk drive or nonvolatile memory, an output unit 36 ​​that outputs data to the display device 16, an input unit 38 that receives data from the input device 14 or the microphone 18, a recording medium drive unit 40 that drives a removable recording medium such as a magnetic disk, an optical disk, or a semiconductor memory, and an audio processing unit 42 that processes signals to the speaker 20.

[0024] The CPU 23 controls the entire information processing device 10 by executing an operating system stored in the storage unit 34. The CPU 23 also executes various programs read from removable recording media and loaded into the main memory 26, or downloaded via the communication unit 32. The GPU 24 performs various image processing operations in response to requests from the CPU 23. The main memory 26 is composed of RAM (Random Access Memory) and stores programs and data required for processing. The sound processing unit 42 generates electrical signals representing sound in response to requests from the CPU 23 and outputs them to the speaker 20.

[0025] In the present embodiment, with the configuration described above, audio output that enhances the sense of realism of the displayed world and the quality of the user experience is realized with minimal effort on the part of the content creator. To that extent, the type of electronic content and the display content are not particularly limited, but as a typical example, the following description will be given assuming content that represents a virtual world composed of three-dimensional objects. Hereinafter, several means for achieving this will be described, but each may be achieved independently, or two or more may be combined to achieve the same.

[0026] 3A and 3B are diagrams for explaining how a sound source is adjusted depending on the output status in this embodiment. Fig. 3A illustrates an example of a frame of an image displayed on the display device 16. This example shows a main character 102 that can be controlled by a user who is a player of the game, and other characters (e.g., other characters 104) that are defined by the program or that can be controlled by other players, in a virtual space that includes a river 106.

[0027] 1(b) shows a bird's-eye view of the virtual space. The information processing device 10 constructs a three-dimensional virtual space 108 consisting of a main character 102, other characters (e.g., other characters 104, 112), a river 106, and the like. The information processing device 10 sequentially changes the position, posture, shape, and other states of each object based on user operations and programs. The information processing device 10 also places a virtual camera 110 in the virtual space to generate a display image such as that shown in 1(a). For example, as shown in the figure, the virtual camera 110 is set behind the main character 102, and the position and posture of the virtual camera 110 are changed in accordance with the movement of the main character 102.

[0028] By rendering the virtual space as seen from the virtual camera 110 at a predetermined rate, a moving image representing the virtual space can be displayed, realizing a so-called third-person shooter game. However, in this embodiment, the position of the virtual camera 110 is not particularly limited. The information processing device 10 also sets a position corresponding to the user's viewpoint, such as the position of the virtual camera 110, as a sound receiving point, and generates and outputs sounds of the virtual space 108 that can be heard there. To this end, the information processing device 10 extracts objects that generate sounds in the virtual space 108 and sets sound sources at those positions. The information processing device 10 then processes and synthesizes each sound, taking into account the propagation of sound from each sound source to the sound receiving point. This allows the user to hear realistic sounds that match the situation in the virtual space.

[0029] However, accurately reproducing all sounds occurring in the virtual space 108 may actually degrade the quality of the user experience. For example, if many other characters each speak, or if the sound of the flowing water in the river 106 becomes louder, the synthesized sound may be perceived as noisy, or it may be difficult to hear necessary sounds, such as the sound of the main character 102's movements. Whether or not such problems occur can be determined by the results of synthesizing various sounds, so content creators have had to go to the trouble of listening to the actual sounds for each scene and adjusting the volume settings in the program, for example.

[0030] In this embodiment, the information processing device 10 evaluates the "loudness" based on at least one of the state of the sound output from the speaker 20 and the state of the sound source in the virtual space 108, and determines whether sound source adjustment is necessary. If adjustment is necessary, the information processing device 10 disables some sound sources or reduces their volume. At this time, the objects to be adjusted are determined based on priorities according to the type of sound source and its positional relationship with the sound receiving point. For example, in the example shown in the figure, the priority of other characters 112 outside the field of view of the virtual camera 110 is lowered, and the sound source is disabled or its volume is reduced. By setting such adaptive rules, "loudness" can be easily eliminated with the minimum necessary adjustments.

[0031] Fig. 4 shows the configuration of functional blocks of the information processing device 10 in this embodiment. Each functional block shown in Fig. 4 and Figs. 9, 12, and 16 described below can be realized in hardware terms by the configuration of the CPU, GPU, various memories, data bus, etc. shown in Fig. 2, and in software terms by a program that performs various functions such as data input function, data retention function, calculation function, image processing function, and sound processing function, loaded into memory from a recording medium, etc. Therefore, it will be understood by those skilled in the art that these functional blocks can be realized in various forms using only hardware, only software, or a combination thereof, and are not limited to any one of them.

[0032] The information processing device 10 includes an input information acquisition unit 50 that acquires input information from the input device 14 and the microphone 18, a virtual space control unit 52 that controls the virtual space, an image data generation unit 54 that generates data for images to be displayed, an audio data generation unit 58 that generates data for audio to be output, and an output unit 64 that outputs audio and image data.

[0033] The input information acquisition unit 50 sequentially acquires the content of user operations from the input device 14. The input information acquisition unit 50 also acquires audio data picked up by the microphone 18. In the viewing phase, the virtual space control unit 52 processes the electronic content in response to user operations and controls objects present in the display space. At this time, the virtual space control unit 52 processes the electronic content using well-known techniques based on programs and various data stored internally.

[0034] In the production phase, the virtual space control unit 52 also controls objects present in the space to be displayed by calling up data of the scene to be produced in response to user operations and processing it appropriately. In this case, the virtual space control unit 52 also processes the electronic content using well-known techniques based on the program and various data being produced.

[0035] The image data generation unit 54 generates data for images to be displayed as a result of processing by the virtual space control unit 52. That is, the image data generation unit 54 draws an image representing the field of view of the virtual camera in the virtual space controlled by the virtual space control unit 52. Alternatively, the image data generation unit 54 may combine an image of the virtual space with an image of the real world captured by a camera serving as the input device 14. The image data generation unit 54 generates the images at a predetermined rate and supplies them sequentially to the output unit 64.

[0036] The sound data generation unit 58 generates sound to be output in accordance with the displayed image. Specifically, the sound data generation unit 58 includes an output state acquisition unit 70, a sound source state acquisition unit 72, a sound source control unit 74, a sound generation unit 76, and an audio data storage unit 78. The output state acquisition unit 70 acquires the state of the sound output from the speaker 20, which is picked up by the microphone 18. Here, the "state of the output sound" refers to, for example, the volume, but this is not intended to limit the present embodiment to this, and may also refer to a frequency band or the like.

[0037] The sound source state acquisition unit 72 acquires the sound source state based on the state of the virtual space controlled by the virtual space control unit 52. Here, the "sound source state" refers to the location of a sound source, if any, currently being generated, the need for a new sound to be generated, and the location of a new sound to be generated, if any. For example, the sound source state acquisition unit 72 determines whether or not speech or cries from people, animals, robots, etc. present in the virtual space are required, as well as whether or not footsteps of these objects, sounds of weapons, vehicles, etc., and environmental sounds, such as rivers, wind, and rain, are required. The sound source state acquisition unit 72 may also determine whether or not additional sounds without a clear source location, such as background music or narration, are required. If the location of the sound source can be defined, the sound source state acquisition unit 72 sets the location of the sound source according to the state of the virtual space.

[0038] The sound source control unit 74 controls the sound source based on the state of the output sound acquired by the output state acquisition unit 70 and the sound source state acquired by the sound source state acquisition unit 72. That is, the sound source control unit 74 determines whether or not at least one of the state of the output sound and the sound source state satisfies the condition for adjusting the sound source, and when adjustment is required, determines to disable or reduce the volume of the sound source in order of lowest priority. Here, the "condition for adjusting the sound source" is, qualitatively, a condition indicating that the sound exceeds a threshold at which a person perceives it as noisy.

[0039] For example, when the volume from the speaker 20 exceeds a threshold, the sound source control unit 74 determines to disable some sound sources or reduce the volume. In this case, the sound source control unit 74 may increase the number of sound sources to disable or increase the extent to which the volume is reduced as the difference between the actual volume and the threshold increases. The condition for adjustment may be determined taking into account the volume and frequency band as the state of the output sound. Alternatively, the sound source control unit 74 may determine to disable some sound sources or reduce the volume as the number of sound sources exceeds a threshold. In this case, the sound source control unit 74 may also increase the number of sound sources to disable or increase the extent to which the volume is reduced as the difference between the set number of sound sources and the threshold increases.

[0040] The sound source control unit 74 may determine that sound source adjustment is necessary when the sum of the volumes emitted from each sound source is greater than a threshold value. The sound source control unit 74 determines target sound sources to be disabled or have their volumes reduced according to a priority preset for each sound to be generated. The sound source control unit 74 may adaptively determine the priority based on the positional relationship of the sound source with respect to the sound receiving point. Alternatively, the sound source control unit 74 may determine a final priority by combining a preset priority with a priority based on the positional relationship with the sound receiving point.

[0041] In either case, the sound source control unit 74 disables sound sources in ascending order of priority, or reduces the volume of sound sources in descending order of priority. Alternatively, the sound source control unit 74 may disable a certain number of sound sources in descending order of priority, and then reduce the volume of a certain number of sound sources in descending order of priority. The amount of volume reduction does not need to be fixed, and the amount of reduction may be greater the lower the priority. As a modified example, the sound source control unit 74 may apply filtering to the sound of low-priority sound sources to make them less noticeable.

[0042] The sound generator 76 generates sound data to be output by synthesizing sounds emitted from valid sound sources in accordance with the determination of the sound source controller 74. At this time, the sound generator 76 may synthesize the sound after processing it in consideration of sound propagation based on the positional relationship between the sound source, sound receiving point, surrounding obstructions, etc. Any practically usable method may be used for such processing.

[0043] The audio data storage unit 78 stores data related to sounds to be generated in the virtual space. Specifically, the audio data storage unit 78 stores data associating the object that generates the sound with the original data of the sound and its priority. Here, "original data" refers to sound data at the sound source before it propagates through space, and includes data created or recorded in advance, as well as data such as synthesized voice generated on the spot. Hereinafter, such sounds may be referred to as "source sounds." The audio data storage unit 78 may also store data associating the sound data with priority for sounds not associated with specific objects, such as background music and narration.

[0044] The output unit 64 outputs the display image data generated by the image data generation unit 54 to the display device 16. The output unit 64 also outputs the sound data generated by the sound data generation unit 58 to the speaker 20. The sound output from the speaker 20 is picked up by the microphone 18 and used to acquire the output state by an output state acquisition unit 70 of the sound data generation unit 58. In the production phase, the output unit 64 further stores the sound data, in which the sound source has been adjusted as necessary, in a storage device in response to user operation. The sound data may be associated with a scene of the content being produced and may become part of the electronic content.

[0045] Next, the operation of the information processing device 10 that can be realized by the configuration of this embodiment will be described. Figure 5 is a flowchart showing the processing procedure in which the information processing device 10 outputs sound while adjusting the sound source as necessary. This flowchart is executed in parallel with the content being processed by user operation and the image of the virtual space being displayed. First, the output state acquisition unit 70 and the sound source state acquisition unit 72 of the sound data generation unit 58 respectively acquire the output sound state and sound source state at that time (S10).

[0046] That is, the output state acquisition unit 70 acquires the state of sound from the speaker 20 that is being picked up by the microphone 18. The sound source state acquisition unit 72 determines the need to generate sound based on the state of objects in the virtual space controlled by the virtual space control unit 52, and if sound generation is required, sets a sound source at the corresponding position. The sound source state acquisition unit 72 also determines the need to generate background music, narration, etc., and if sound generation is required, sets the sound source at an indefinite position.

[0047] Next, the sound source control unit 74 determines whether or not the sound source should be adjusted based on at least one of the state of the output sound and the state of the sound source, and on the conditions described above (S12). If the conditions for adjusting the sound source are met (Y in S12), the sound source control unit 74 adjusts the sound source by the means described above (S14). At this time, the sound source state acquisition unit 72 adjusts the sound source in order of relative priority, in accordance with at least one of the priority of each sound stored in the audio data storage unit 78 and the positional relationship between the sound source and the sound receiving point.

[0048] If the conditions for adjusting the sound source are not met, the sound source control unit 74 skips the sound source adjustment process (N at S12). Next, the sound generation unit 76 reads the source sound data of the valid sound source from the audio data storage unit 78, plays it, and synthesizes it after processing it taking into account propagation to the sound receiving point (S16). The output unit 64 outputs the synthesized sound data to the speaker 20 or a storage device (not shown) (S18). The processes of S10 to S18 are repeated until the content to be played or created is completed (N at S20), and all processes are terminated when the content is completed or the user ends the creation (Y at S20).

[0049] 6 illustrates the structure of the audio data stored in the audio data storage unit 78. In this example, the audio data 120 associates an object ID 122a, audio data 122b, and a priority 122c. The object ID 122a is identification information for an object existing in the virtual space, and the audio data 122b is data on the source sound emitted by each object. For example, the cry of an animal with an object ID of "01" is stored as data on the source sound "AAA." However, the actual audio data may be stored in a separate storage area.

[0050] For simplicity, in the figure, one source sound data is associated with one object, but the number of associated sound data is not limited. For example, multiple action sound data may be associated with one object. In this case, multiple actions and sound source positions on the object are associated with the object ID 122a shown in the figure, and action sound data is associated with each of them. Priority 122c is a parameter that represents the priority of the sound that should be maintained even when adjusting the sound source. In the example shown in the figure, the priorities are shown as "1," "2," and "3," with "BBB," "AAA," and "CCC" indicating the highest priority in this order. However, the priority parameter is not particularly limited, and the same priority may be assigned to multiple sound data.

[0051] 7 is a diagram illustrating an example of a technique for determining the priority of a sound source based on its positional relationship with a sound receiving point. The diagram shows a bird's-eye view of a sound receiving point 130 and the space around it in a virtual space. In this example, a reference direction 132 is defined as the direction indicated by the outline arrow with respect to the sound receiving point 130. When the sound receiving point 130 is set at the position of a virtual camera, the reference direction 132 is, for example, the optical axis direction of the virtual camera. When the sound receiving point 130 is set at the position of a main character, the reference direction 132 is, for example, the front direction of the main character.

[0052] The priority boundary is then defined by thresholds given to the distance D from the sound receiving point and the angle θ with the reference direction 132 as the central axis. Note that the distance D and angle θ may be defined relative to a reference plane such as the ground, or may be defined three-dimensionally including the height direction. The threshold value for the angle θ may be less than or greater than 180°. In the simplest case, one pair of thresholds for the distance D and the angle θ is set, and the priority of sound sources within a range 134 below the threshold is set to high, and the priority of other sound sources is set to low. In this case, the information processing device 10 does not adjust sound sources within the range 134, but disables or lowers the volume of sound sources outside the range.

[0053] On the other hand, two or more pairs of threshold values ​​for distance D and angle θ may be set, and three or more priorities may be set. Qualitatively, the smaller the distance D and angle θ, the higher the priority. This allows the sound of an object that is likely to attract the user's attention to remain as is as much as possible. The threshold value may be fixed or may be changed according to the characteristics of the scene. Furthermore, even for sound sources within the same priority range, the priority may be differentiated according to the priority 122c previously set in the audio data 120 shown in FIG. 6.

[0054] According to the above-described sound source adjustment, the information processing device 10 determines the need to reduce the number of sound sources or the volume of some sound sources based on the actual output sound state or the state of the sound sources that suggests that state. This allows for necessary and sufficient adjustments to be made based on the user's sensory perception of a "noisy" situation. Furthermore, by assigning priorities flexibly based on the characteristics of the sound and its positional relationship with the sound receiving point, the sound estimated to be necessary can be accurately maintained. This saves content creators the trouble of adjusting sound sources while checking the output sound for each scene.

[0055] <Adjusting the distribution of point sound sources according to the area of ​​the sound source> Generally, sound sources set in a virtual space are point sound sources, and each sound source is assigned a single position coordinate. On the other hand, since the area of ​​sound sources in the real world varies, if the change in area can be reflected in the sound, a more realistic acoustic expression can be achieved. For example, when an object talks or laughs, the sound spread and volume change depending on the size of the mouth, even if it is the same speaking or laughing voice. To reproduce this, in this embodiment, the number of point sound sources is increased or decreased and their distribution is adjusted according to the area of ​​the sound source, such as the mouth.

[0056] FIG. 8 is a diagram illustrating an embodiment in which the distribution of point sound sources is adjusted according to the area of ​​the sound source. The diagram shows variations of the object's face, with the degree of mouth opening varying horizontally and the size of the face itself varying vertically. When setting the sound source for the voice emitted by the object, face 140a has a small mouth opening, so one point sound source 142a is set in the center of the mouth. In contrast, the number of point sound sources increases as the mouth opening increases, as in faces 140b and 140c. In the example shown in the diagram, three point sound sources (e.g., point sound source 142b) are set for face 140b, and six point sound sources (e.g., point sound source 142c) are set for face 140c.

[0057] Furthermore, even if the degree of mouth opening is the same, the number of point sound sources is increased if the area of ​​the mouth increases as a result of the object itself becoming larger. In the example shown in the figure, six point sound sources (e.g., point sound source 142d) are set for face 140d. Here, the size of the object may refer to the size of the object itself in three-dimensional space, or the apparent size from a virtual camera. Qualitatively, the more the area of ​​the mouth increases, the more point sound sources are added, and the distribution range of the sound sources is adjusted according to the shape of the surface, so that the sound sources are arranged to cover the surface. In other words, the number and distribution of point sound sources are adjusted so that the density of point sound sources is approximately constant relative to the area of ​​the sound source.

[0058] Note that sound sources that require adjustment of the number and distribution of sound sources are not limited to mouths. For example, they may be contact surfaces that generate collision sounds with other objects, such as hands clapping or the size of tools, or large objects that themselves emit sound, such as weapons, musical instruments, or rivers. The number and distribution of sound sources do not need to be changed over time in accordance with changes in the size of the sound source as shown in the figure, but may also be set to a fixed value for objects that do not change in size. Furthermore, the adjustments that are made depending on changes in the size of the sound source are not limited to the number and distribution of sound sources, and the processing applied to the source sound may also be changed. For example, filtering may be applied so that the frequency band of the sound emitted by the larger the object, becomes lower.

[0059] 9 shows the functional block configuration of the information processing device 10a in this embodiment. Functional blocks similar to those of the information processing device 10 shown in FIG. 4 are denoted by the same reference numerals, and descriptions thereof will be omitted where appropriate. The information processing device 10 includes an input information acquisition unit 50a that acquires input information from the input device 14, a virtual space control unit 52 that controls the virtual space, an image data generation unit 54 that generates image data to be displayed, an audio data generation unit 58a that generates audio data to be output, and an output unit 64 that outputs audio and image data.

[0060] The input information acquisition unit 50a sequentially acquires the contents of user operations from the input device 14. The virtual space control unit 52, the image data generation unit 54, and the output unit 64 have functions similar to those of the corresponding functional blocks in the information processing device 10 shown in FIG. 4. The sound data generation unit 58a generates sound to be output in accordance with the displayed image. In detail, the sound data generation unit 58a includes a sound source state acquisition unit 72a, a sound generation unit 76, and an audio data storage unit 78a.

[0061] 4, the sound source state acquisition unit 72a acquires the state of the sound to be generated as the sound source state based on the state of the virtual space controlled by the virtual space control unit 52. In particular, in this aspect, when the sound source position of the sound to be generated can be defined, the number and distribution of point sound sources are adjusted according to the area of ​​the sound source. In detail, the sound source state acquisition unit 72a includes an area acquisition unit 82 and a sound source setting unit 84.

[0062] The area acquisition unit 82 acquires the area of ​​the sound source in the virtual space. As exemplified above, the area of ​​the sound source may be the area of ​​the part of the object that emits the sound, or the contact area when sound is generated by collision with another object. The sound source setting unit 84 determines the number and distribution of sound sources according to a predetermined rule based on the acquired area. This determines the position coordinates of one or more point sound sources for one source sound. The sound generation unit 76 generates sound data to be output by synthesizing the sounds emitted from each point sound source in accordance with the determination by the sound source state acquisition unit 72a. At this time, the sound generation unit 76 may perform processing that takes into account sound propagation based on the positional relationship between the sound source, sound receiving point, surrounding obstructions, etc., before synthesizing.

[0063] The sound data storage unit 78a stores data related to sounds to be generated in the virtual space. Specifically, the sound data storage unit 78 stores data relating to objects that generate sounds and source sounds, as well as data associating the data with rules for determining the number and distribution of point sound sources. The area acquisition unit 82 and the sound source setting unit 84 determine the targets for which the area is to be acquired and adjust the number and distribution of point sound sources according to the data stored in the sound data storage unit 78a.

[0064] FIG. 10 illustrates an example of the structure of audio data stored in the audio data storage unit 78a. In this example, audio data 150 associates an object ID 152a, audio data 152b, and a point sound source setting rule 152c. The object ID 152a and audio data 152b may be the same as the object ID 122a and audio data 122b shown in FIG. 6. The sound source setting rule 152c represents a setting rule for the number and distribution of sound sources according to the area. In the example shown in the figure, the names of rules "rule a" and "rule b" are shown, but in reality, the number of point sound sources may be expressed as a function of the area, or rules for deriving the position coordinates of each point sound source, such as the spacing between point sound sources or arrangement rules. The setting rule may also define a rule for determining the range of the surface to be acquired by the area acquisition unit 82.

[0065] Depending on the type of sound, no rules may be set, i.e., the number of point sound sources may be fixed regardless of the area. In the figure, such settings are indicated as "N / A." In this way, by being able to set whether or not to add point sound sources for each object or sound, and if so, the number and distribution generation rules, point sound sources can be set in an optimal state depending on the characteristics of the object or sound.

[0066] Next, the operation of the information processing device 10, which can be realized by the configuration of this embodiment, will be described. Figure 11 is a flowchart showing the processing procedure in which the information processing device 10 outputs sound while adjusting the distribution of point sound sources according to the area of ​​the sound source. This flowchart is executed in parallel with the processing of content through user operation and the display of images in the virtual space. First, the sound source status acquisition unit 72a of the sound data generation unit 58a determines whether a new sound source has appeared or the area of ​​the source has changed based on the processing results of the virtual space control unit 52 (S30). If no such change occurs, the sound source status acquisition unit 72a waits (N in S30).

[0067] If a new sound source appears or the area of ​​a sound source changes (Y in S30), the sound source state acquisition unit 72a refers to the audio data stored in the audio data storage unit 78a and checks whether there is a corresponding point sound source setting rule (S32). If there is a setting rule (Y in S32), the sound source state acquisition unit 72a adjusts the point sound source in accordance with the setting rule (S34). That is, the area acquisition unit 82 identifies the area of ​​the part of the object that will become the sound source, and the sound source setting unit 84 adjusts the number and distribution of point sound sources according to the identified area. The sound source setting unit 84 then sets the point sound source in accordance with the adjustment result (S36).

[0068] If there is no rule for setting a point sound source (N in S32), the sound source setting unit 84 simply sets one point sound source at the location of the object that will be the sound source (S36). However, even in this case, the sound source setting unit 84 may set multiple predetermined point sound sources at the object that will be the sound source. Next, the sound generation unit 76 reads the source sound data of the set sound source from the audio data storage unit 78a, plays it, and synthesizes it after processing it to take into account propagation to the sound receiving point (S38). The output unit 64 outputs the synthesized sound data to the speaker 20 or a storage device (not shown) (S40). The processes of S30 to S40 are repeated until the content to be played or created is completed (N in S42). When the content is completed or the user ends the creation, all processes are terminated (Y in S42).

[0069] According to the above-described aspect of adjusting the distribution of point sound sources, the information processing device 10 identifies the size of the surface from which sound is generated, and then adjusts and distributes the number of point sound sources accordingly. Therefore, the point sound source setting rules are optimized in advance according to the characteristics of the objects and sounds. This allows content creators to simulate surface sound sources where necessary, realizing a more realistic sound representation, without having to take the time and effort of individually considering the distribution of point sound sources for each scene.

[0070] <Automatic Generation of Diverse Source Sounds> The aspects described above are based on the premise that source sounds appropriate for the sound source are prepared. In other words, if the source sound itself lacks realism, the sense of realism of the virtual space will be lost. For example, in the case of footsteps, the sound changes depending on various factors such as the material and size of the shoes, the length of the feet, and the condition of the ground, so it has been time-consuming to re-record the footsteps in an environment that matches the settings of the virtual space. Therefore, in this embodiment, a device is realized that can perform various processing on recorded sounds and present a large number of options, making it possible to easily obtain the desired source sound.

[0071] Fig. 12 shows the functional block configuration of the audio data generating device in this embodiment. The audio data generating device 160 may be included in the content processing system 2 shown in Fig. 1 in place of or in combination with the information processing device 10. The audio data generating device 160 is used by the content creator to generate source sound in the production phase. The generated source sound data is ultimately stored in the audio data storage units 78, 78a of the information processing device 10.

[0072] The audio data generating device 160 includes an input information acquiring unit 162 that acquires input information from the input device 14, an image data generating unit 164 that generates data for images to be displayed, an audio data storage unit 166 that stores recorded audio data, an audio processing unit 168 that processes audio, and an output unit 174 that outputs audio and image data. The input information acquiring unit 50 sequentially acquires the content of user operations from the input device 14. Here, the input information includes selection of audio data to be processed, selection of processing content, and instructions to output data to a storage device. The image data generating unit 164 generates image data for a user interface for audio processing.

[0073] The audio data storage unit 166 stores pre-recorded audio data. The audio processing unit 168 processes audio data selected by the user from the audio data stored in the audio data storage unit 166. The audio processing unit 168 internally holds a processing pattern storage unit 172. The processing pattern storage unit 172 stores data related to processing variations. Here, processing applied to audio data refers to changing one or a combination of the values ​​of well-known adjustment parameters such as volume, pitch, ADSR (attack, decay, sustain, release), cutoff, reverb, and echo. When multiple audio signals are synthesized, the processing may also include the synthesis time, synthesis ratio, etc.

[0074] The processing pattern storage unit 172 stores a plurality of sets of combinations of these parameter values. The audio processing unit 168 reads out audio data selected by the user from the audio data storage unit 166 and processes the data using the parameter values ​​prepared in the processing pattern storage unit 172 in accordance with the user's operation. The output unit 174 outputs the processed audio data to the speaker 20 or a storage device. The output unit 174 also outputs the user interface image data generated by the image data generation unit 164 to the display device 16.

[0075] 13 illustrates an example of the data structure of a processing pattern stored in the processing pattern storage unit 172. The processing pattern data 180 associates a pattern name 182a with a combination of multiple parameter values ​​used for processing. In this example, values ​​are set for each of parameters such as volume 182b, pitch 182c, attack 182d, decay 182e, etc., and a pattern name is assigned to each set of these combinations. The specific setting values ​​may vary depending on the type of parameter. For example, the amount or rate of change from the original source sound may be set depending on the parameter, or an absolute value may be set.

[0076] In the illustrated example, the pattern names 182a are "PA," "PB," "PC," etc., but they may also be words that represent the processing effects obtained by the combinations of parameter values. For example, by giving names such as "dry sound," "wet sound," "dark sound," and "bright sound," the user can intuitively grasp the desired processing content and make selections easier. The processing pattern data 180 can be created, for example, by listening to test sounds processed with various combinations of parameter values ​​and extracting useful combinations. In this case, the pattern names can be determined by verbalizing the processing effects perceived by the creator when listening to them.

[0077] 14 shows an example of a user interface image for audio processing generated by the image data generation unit 164. In this example, the user interface image 190 includes a target audio data designation field 192, a playback operation field 194, a corresponding content image field 196, processing pattern selection buttons 198a, 198b, and 198c, and a parameter adjustment field 200. The user inputs the name of the audio data or the like in the target audio data designation field 192 to designate the audio data to be processed. The playback operation field 194 is a GUI (Graphical User Interface) that accepts operations such as playing, stopping, and pausing the audio before and after processing.

[0078] The corresponding content image field 196 shows the scene of the content for which the processed audio should be output. By viewing the scene of the content and listening to the processed audio, the user can determine whether the processing is appropriate. The processing pattern selection buttons 196a, 196b, and 196c are GUI buttons that allow the user to select a pattern name for the processing content. In the illustrated example, the thick frame indicates that the button 198b labeled "PB" has been selected. However, as mentioned above, the pattern name may also be a word that represents the effect of the processing. When a playback operation is performed in this state using the playback operation field 194, the audio processing unit 168 processes the target audio data using the combination of parameter values ​​for "PB" described in the processing pattern data 180 and outputs the processed audio from the speaker 20 via the output unit 174.

[0079] The parameter adjustment field 200 is a field for accepting fine adjustments of various parameters used in the processing. In this example, a slider GUI is provided for changing the value of each parameter, such as "pitch." By allowing further adjustments to the processing results based on the combination of default values ​​selected using the processing pattern selection button, the variety of processing content can be increased without limit. Once the optimal processing pattern is obtained in this way, the user can assign a pattern name to the combination of parameter values ​​and store it as a new processing pattern in the audio data storage unit 166. This allows the processing pattern to be reused at any time in the future.

[0080] Note that the illustrated screen configuration is merely an example and is not intended to limit the present embodiment. For example, while the configuration shown in the figure allows only one target voice data to be designated, it is also possible to accept the designation of multiple voice data and display the option of synthesizing them. In this way, for example, by designating a large number of data representing the speech of various people, it is possible to synthesize them to generate voice data representing the hustle and bustle of a crowd.

[0081] Next, the operation of the audio data generating device 160 that can be realized by the configuration of this embodiment will be described. Fig. 15 is a flowchart showing the processing procedure by which the audio data generating device 160 processes and outputs source sound. This flowchart starts in a state where the image data generating unit 164 displays a user interface image on the display device 16 and the user has designated target audio data. First, the audio processing unit 168 reads the designated audio data from the audio data storage unit 166 (S50).

[0082] When the user specifies the name of a processing pattern (S52), the voice processing unit 168 reads out the corresponding combination of parameter values ​​from the processing pattern storage unit 172, processes the target voice using the combination, and outputs the processed voice to the speaker 20 (S54, S56). The processes from S52 to S56 are repeated until the user inputs an instruction to store the processed voice data in the storage device (N in S58). Note that in the process of S52, fine adjustment of the parameter values ​​may be accepted using the parameter adjustment field 200 shown in FIG. 14.

[0083] When an instruction to store the audio data in the storage device is input (Y in S58), the audio processing unit 168 stores the audio data resulting from the latest processing in the storage device via the output unit 174 (S60). At this time, the audio processing unit 168 may also store the processing pattern used in the storage device so that it can be reused. The audio data stored in the storage device in this way can be used as source sound when the information processing device 10 processes content by associating it with objects and various actions that exist in the virtual space.

[0084] According to the above-described sound source adjustment, the sound data generating device 160 processes recorded sound data using various combinations of parameter values ​​to create a wide variety of source sound options. For this reason, it prepares combinations of parameter values ​​that are considered useful, and uses pattern names that intuitively express their effects. This makes it easy to obtain sound data that matches the scene, eliminating the need to re-record similar sounds to account for minute differences in the virtual space.

[0085] <Extracting Audio Breakpoints> When outputting sound in accordance with a virtual space, it is possible to create sound of any duration by repeatedly playing back recorded audio data of a finite duration. For example, by recording footsteps of a few seconds and repeatedly playing them back, it is possible to make the footsteps sound continuous so that they match the time of an object walking in the virtual space. In this case, if the beginning and ending sounds of the footsteps as the source sound are not appropriate, the sound may be interrupted at the joints of the repetition, or may not sound like periodic footsteps.

[0086] Furthermore, for sounds that are not inherently periodic, such as flowing river water or chirping birds, it is necessary to smoothly connect the end and start points of longer audio data so that the repeated playback is not noticeable. Therefore, in this embodiment, the audio data generation device extracts segments from the audio data that are suitable for the start and end points of playback. This function is not limited to cases where audio data is repeatedly played by connecting the end and start points, but can also be used in cases where music is played from the middle in a specific scene. For example, if a virtual space is set up as a coffee shop with background music playing, one of the extracted segments is randomly selected and played each time the main character enters the coffee shop. This prevents the unnatural feeling of the same song starting from the beginning each time the main character enters the coffee shop. Furthermore, playing the audio from a suitable starting point eliminates the abrupt feeling, encouraging a smooth entrance.

[0087] Fig. 16 shows the functional block configuration of the audio data generation device in this embodiment. Functional blocks similar to those of the audio data generation device 160 shown in Fig. 12 are assigned the same reference numerals, and descriptions thereof will be omitted where appropriate. The audio data generation device 160a includes an input information acquisition unit 162 that acquires input information from the input device 14, an image data generation unit 164 that generates data for images to be displayed, an audio data storage unit 166 that stores recorded audio data, an audio segment extraction unit 210 that extracts segments of the audio data, and an output unit 174 that outputs audio and image data.

[0088] The input information acquisition unit 50 sequentially acquires the contents of user operations from the input device 14. Here, the input information includes the selection of audio data to be processed, the selection of extraction rules, and instructions to output data to the storage device. The image data generation unit 164a generates image data of a user interface for segment extraction. The audio data storage unit 166 stores pre-recorded audio data. The audio segment extraction unit 210 extracts segments suitable for the start and end points of playback from the audio data selected by the user among the audio data stored in the audio data storage unit 166.

[0089] In detail, the audio segment extraction unit 210 includes a waveform analysis unit 211, a segment setting unit 212, and an extraction rule storage unit 214. The waveform analysis unit 211 analyzes the waveform of the audio data to be processed and extracts one or more locations suitable as segmentations according to preset rules. The locations are extracted as time points (timings) on the time axis of the audio data. The segment setting unit 212 processes the audio data to be processed so as to clearly indicate the extracted segmentations.

[0090] For example, the delimiter setting unit 212 generates new audio data by extracting the audio data from a start point to an end point. Alternatively, the delimiter setting unit 212 marks a time point in the audio data that marks a delimiter. For example, an audio signal with a frequency outside the audible range is embedded at the point to be marked in the audio data. This allows the information processing device 10, when processing content, to start playback of the audio data from the point where the signal is embedded, or to connect the end point and start point represented by two points and repeatedly play the audio data. Points suitable as playback start points and end points can be distinguished by the wavelength of the embedded signal.

[0091] The extraction rule storage unit 214 stores pre-generated extraction rules for segmentation. For example, multiple extraction rules may be set depending on the purpose of segment extraction, such as repeated playback or playback from the middle to a desired timing. The extraction rule storage unit 214 may also store options for clearly indicating segmentation, such as extracting new audio data or embedding audio signals. Such options are presented to the user by the image data generation unit 154a so that the user can select them. The waveform analysis unit 211 and segment setting unit 212 extract and indicate segmentation according to the rules selected by the user.

[0092] The output unit 174 outputs the user interface image generated by the image data generation unit 164 to the display device 16. The output unit 174 also repeatedly plays back the extracted audio data and outputs it to the speaker 20, and outputs the audio data or data in which an audio signal representing a break is embedded to a storage device.

[0093] 17 shows an example of a waveform of audio data analyzed by the waveform analysis unit 211. In this waveform, the horizontal axis represents time and the vertical axis represents sound amplitude. In the example shown, there are silent periods at the beginning and end of the audio data, so it is desirable to trim these periods. Also, in order to smoothly connect the end and start points of the audio during repeated playback, it is desirable to use quiet sections with small amplitude as separators. For these reasons, the waveform analysis unit 211 extracts points where the amplitude remains below a predetermined value for a predetermined duration, as indicated by the arrows, as separator candidates.

[0094] Alternatively, the waveform analysis unit 211 may estimate the content and characteristics of the audio by performing pattern matching on the waveform, and identify suitable segments based on the results. However, this is not intended to limit the waveform analysis method. The image data generation unit 164a may present candidate segments to the user along with the waveform as shown in the figure, and the user may determine the final segments and the distinction between the start and end points by listening to and confirming the actual sound. The segment setting unit 212 generates new audio data extracted at the determined segments, or embeds an audio signal representing the segment, and stores the data in a storage device. The audio data stored in the storage device in this way can be used as source sound when the information processing device 10 processes content by associating it with objects and various actions existing in a virtual space.

[0095] According to the above-described audio segment extraction method, the audio data generation device 160a analyzes the waveform of the recorded audio to extract segments suitable for use as end and start points for repeated playback or as start points for playback that has been interrupted. The audio data generation device 160a then extracts audio data at the extracted segments to generate new audio data, or embeds audio signals with wavelengths outside the audible band, and stores the data in a storage device. This allows for efficient generation of audio that matches the scene, without the need to repeatedly listen to all sound sources required to represent the virtual space and manually set segments or extract audio data.

[0096] The present invention has been described above based on the embodiments. The above embodiments are merely examples, and it will be understood by those skilled in the art that various modifications are possible in the combination of the respective components and treatment processes, and that such modifications are also within the scope of the present invention.

[0097] 10 Information processing device, 14 Input device, 16 Display device, 18 Microphone, 20 Speaker, 23 CPU, 24 GPU, 26 Main memory, 50 Input information acquisition unit, 52 Virtual space control unit, 54 Image data generation unit, 58 Sound data generation unit, 64 Output unit, 70 Output state acquisition unit, 72 Sound source state acquisition unit, 74 Sound source control unit, 76 Sound generation unit, 78 Audio data storage unit, 82 Area acquisition unit, 84 Sound source setting unit, 160 Audio data generation device, 162 Input information acquisition unit, 164 Image data generation unit, 166 Audio data storage unit, 168 Audio processing unit, 172 Processing pattern storage unit, 174 Output unit, 210 Audio segment extraction unit, 211 Waveform analysis unit, 212 Segment setting unit, 214 Extraction rule storage unit.

[0098] As described above, the present invention can be used in various devices such as information processing devices, game machines, content processing devices, sound processing devices, and audio data generating devices, as well as systems including any of these devices.

[0099] The present disclosure may include the following aspects. [Item 1] An information processing device including a circuit configured to: set a sound source for a virtual space to be displayed; evaluate a state of sound resulting from the setting of the sound source, adjust the setting of the sound source according to the result; and synthesize and output sounds emitted from the adjusted sound source. [Item 2] The information processing device according to item 1, wherein the circuit acquires an output volume as the state of the sound, and reduces at least one of the number of sound sources and the volume of some of the sound sources when the volume exceeds a threshold. [Item 3] The information processing device according to item 1, wherein the circuit acquires the number of sound sources set by a sound source state acquisition unit as the state of the sound, and reduces at least one of the number of sound sources and the volume of some of the sound sources when the number of sound sources exceeds a threshold. [Item 4] The information processing device according to Item 1, wherein the circuitry obtains the priority of each sound source set when setting a sound source for the virtual space, and performs adjustments to disable sound sources in order of lowest priority. [Item 5] The information processing device according to Item 4, wherein the circuitry determines the priority based at least on a positional relationship with a sound receiving point set in the virtual space. [Item 6] An audio output method comprising: setting a sound source for a virtual space to be displayed, evaluating a state of sound produced by the sound source setting, adjusting the sound source setting according to the result, and synthesizing and outputting sound emitted from the adjusted sound source. [Item 7] A recording medium having recorded thereon a program for causing a computer to implement the following functions: setting a sound source for a virtual space to be displayed, evaluating a state of sound produced by the sound source setting, adjusting the sound source setting according to the result, and synthesizing and outputting sound emitted from the adjusted sound source.[Item 8] An information processing device comprising a circuit configured to: set sound sources for a virtual space to be displayed; synthesize and output sounds emitted from the set sound sources; and, when setting the sound sources, adjust the number of sound sources to be set according to the area of ​​the sound sources. [Item 9] The information processing device according to item 8, wherein the circuit adjusts the distribution range of the set sound sources according to the shape of the surface of the sound sources. [Item 10] The information processing device according to item 8, wherein the circuit changes the number of sound sources according to a change in the area over time. [Item 11] The information processing device according to item 8, wherein the circuit determines the number of sound sources according to a rule set for each sound source. [Item 12] An audio output method, comprising: set sound sources for a virtual space to be displayed; synthesize and output sounds emitted from the set sound sources; and, when setting the sound sources, adjust the number of sound sources to be set according to the area of ​​the sound sources. [Item 13] A recording medium having recorded thereon a program that causes a computer to realize a function of setting sound sources for a virtual space to be displayed, and a function of synthesizing and outputting sounds emitted from the set sound sources, wherein the function of setting the sound sources adjusts the number of sound sources to be set depending on the area of ​​the sound sources. [Item 14] A voice data generation device comprising a circuit configured to: store combinations of parameter values ​​for a plurality of processing means to be applied to voice, process voice data specified by a user using the combinations, and output the processed voice data. [Item 15] The voice data generation device according to Item 14, wherein the circuit stores data that associates a plurality of the combinations with words that express the processing effects produced by each combination, and processes the voice data using one combination selected by the user based on the words.[Item 16] The voice data generation device according to Item 15, wherein the circuitry displays a user interface that accepts selection of the word and adjustment of parameter values ​​associated with the word, processes the voice data after changing the parameter values ​​associated with the selected word in accordance with the adjustment operation, and further stores the adjusted combination of parameter values. [Item 17] A voice data generation method that reads from a memory a combination of parameter values ​​for a plurality of processing means to be applied to voice, processes voice data specified by a user using the combination, and outputs the processed voice data. [Item 18] A recording medium having recorded thereon a program for causing a computer to implement the functions of reading from a memory a combination of parameter values ​​for a plurality of processing means to be applied to voice, processing voice data specified by a user using the combination, and outputting the processed voice data.

Claims

1. An information processing device comprising: a sound source state acquisition unit that sets a sound source for a virtual space to be displayed; a sound source control unit that evaluates the state of the sound generated by the sound source setting and adjusts the sound source setting according to the evaluation result; and a sound generation unit that synthesizes and outputs the sound emitted from the adjusted sound source.

2. The information processing device of claim 1, further comprising an output state acquisition unit that acquires the output volume as the sound state, and wherein the sound source control unit reduces at least one of the number of sound sources and the volume from some of the sound sources when the volume exceeds a threshold value.

3. The information processing device of claim 1, further comprising a sound source state acquisition unit that acquires the number of sound sources set by the sound source state acquisition unit as the sound state, and wherein the sound source control unit reduces at least one of the number of sound sources and the volume from some of the sound sources when the number of sound sources exceeds a threshold value.

4. An information processing device according to any one of claims 1 to 3, characterized in that the sound source control unit acquires the priority of each sound source set by the sound source status acquisition unit, and performs adjustments to disable sound sources in order of lowest priority.

5. The information processing apparatus according to claim 4, wherein the sound source control unit determines the priority based on at least a positional relationship with a sound receiving point set in a virtual space.

6. A sound output method comprising the steps of: setting a sound source for a virtual space to be displayed; evaluating the state of the sound generated by the setting of the sound source and adjusting the setting of the sound source according to the evaluation result; and synthesizing and outputting the sound emitted from the adjusted sound source.

7. A computer program that causes a computer to perform the following functions: setting a sound source for a virtual space to be displayed; evaluating the state of the sound generated by the setting of said sound source and adjusting the setting of the sound source according to the evaluation result; and synthesizing and outputting the sound emitted from the adjusted sound source.

8. An information processing device comprising: a sound source state acquisition unit that sets sound sources for a virtual space to be displayed; and a sound generation unit that synthesizes and outputs sounds emitted from the set sound sources, wherein the sound source state acquisition unit adjusts the number of sound sources to be set depending on the area of ​​the sound generation sources.

9. The information processing device according to claim 8, wherein the sound source state acquisition unit further adjusts the distribution range of the sound source to be set in accordance with the shape of the surface of the sound source.

10. The information processing device according to claim 8 or 9, wherein the sound source state acquisition unit changes the number of the sound sources in accordance with a change in the area over time.

11. The information processing device according to claim 8 or 9, wherein the sound source state acquisition unit determines the number of sound sources in accordance with rules set for each sound source.

12. A sound output method comprising the steps of: setting sound sources for a virtual space to be displayed; and synthesizing and outputting sounds emitted from the set sound sources, wherein the step of setting the sound sources adjusts the number of sound sources to be set according to the area of ​​the sound generating sources.

13. A computer program that causes a computer to implement the functions of setting sound sources for a virtual space to be displayed and synthesizing and outputting sounds emitted from the set sound sources, wherein the function of setting the sound sources adjusts the number of sound sources to be set depending on the area of ​​the sound generating sources.

14. A voice data generating device comprising: a processing pattern memory unit that stores combinations of parameter values ​​for multiple processing means to be applied to voice; a voice processing unit that processes voice data specified by a user using said combinations; and an output unit that outputs the processed voice data.

15. The audio data generating device of claim 14, wherein the processing pattern memory unit stores data associating a plurality of the combinations with words that represent the processing effect resulting from each combination, and the audio processing unit processes the audio data using one combination selected by the user based on the words.

16. The voice data generating device of claim 15, further comprising an image data generating unit that displays a user interface that accepts the selection of the word and an adjustment operation of the parameter value associated with the word, wherein the processing pattern storage unit processes the voice data by changing the parameter value associated with the selected word in accordance with the adjustment operation, and wherein the processing pattern storage unit further stores the combination of the adjusted parameter values.

17. A method for generating voice data, comprising the steps of: reading from memory a combination of parameter values ​​for a plurality of processing means to be applied to voice; processing voice data designated by a user using said combination; and outputting the processed voice data.

18. A computer program that causes a computer to perform the following functions: reading from memory a combination of parameter values ​​for multiple processing means to be applied to voice; processing voice data specified by a user using said combination; and outputting the processed voice data.

Citation Information

Patent Citations

  • Method and apparatus for correction of voice characteristic of synthesized voice

    JP1997127970A

  • Game apparatus

    JP1998211358A

  • Sounding processing apparatus, sounding processing method and sounding processing program

    JP2011123374A

  • Sound controller, program and control method

    JP2013007921A

  • Acoustic characteristic setting support system and acoustic characteristic setting device

    JP2014026146A