Acoustic processing device, acoustic processing method, and acoustic processing program

The acoustic processing device addresses the challenge of creating immersive acoustic experiences in virtual spaces by dynamically correcting audio parameters based on character movement, enhancing immersion and reducing processing burdens.

WO2025120910A1PCT designated stage expired Publication Date: 2025-06-12SONY GROUP CORP +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/027010
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-05
Filing Date
2024-07-29
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Existing technologies face challenges in creating a realistic and immersive acoustic experience in virtual spaces, as faithfully reproducing sounds according to physical laws can lead to significant processing burdens and hinder immersion.

Method used

An acoustic processing device that acquires information on the movement of characters in virtual spaces and uses this information to correct parameters for generating audio signals, thereby enhancing the sense of immersion by simulating real-world acoustic phenomena.

Benefits of technology

The proposed solution effectively enhances the sense of immersion in virtual spaces by dynamically adjusting audio parameters based on character movement, reducing processing burdens while maintaining realistic acoustic simulations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024027010_12062025_PF_FP_ABST
    Figure JP2024027010_12062025_PF_FP_ABST
Patent Text Reader

Abstract

An acoustic processing device according to one embodiment of the present disclosure comprises: an acquisition unit that acquires information relating to movement of a character serving as a sound reception point of a sound source in a virtual space; a correction unit that, on the basis of the information relating to the movement of the character, corrects a parameter used for sound signal generation at the sound reception point; and a generation unit that generates a sound signal at the sound reception point on the basis of the parameter corrected by the correction unit.
Need to check novelty before this filing date? Find Prior Art

Description

Sound processing device, sound processing method, and sound processing program

[0001] The present disclosure relates to a sound processing device, a sound processing method, and a sound processing program.

[0002] Technologies for simulating sounds in virtual spaces are being actively developed. Known sound simulation technologies include wave acoustic simulation and geometric acoustic simulation. It is also known that sound information for reproducing the location of an object in a virtual space is not simply information about the sound, but is also influenced by visual or psychological factors of the user.

[0003] JP 2000-267675 A JP 2013-34107 A

[0004] "Sound Effects in Video Games: Development of Interactive Reverb," ​​Kishi Tomoya, Kojima Kenji, Nakahara Masataka, Hanyu Toshiki, Hoshi Kazuma, Journal of the Acoustical Society of Japan, 2012, Vol. 68, No. 7, pp. 362-368; "Application of Multidimensional Scaling to the Analysis of Sound Perception Phenomena," Ogushi Kengo, Journal of the Acoustical Society of Japan, 2011, Vol. 67, No. 11, pp. 557-558; "Physical and Psychological Quantities of Sound," Nikaido Seiya, Measurement and Control, 1980, Vol. 19, No. 3; "Handbook of Industrial Educational Equipment Systems, edited by the Educational Equipment Editorial Committee," JUSE Press, 1972

[0005] Conventional technology can realize a realistic acoustic space in a virtual space. However, if sounds are faithfully reproduced in accordance with real-world physical laws in order to enhance the user's immersive experience in virtual space content such as games, the amount of calculation required by the information processing device becomes enormous, which may cause problems in processing.

[0006] In this regard, as mentioned above, since the reproduction of sound is influenced by the user's visual and psychological factors, it may be possible to provide content that is more immersive for the user by not only reproducing sound in accordance with the laws of physics but also taking other factors into account.

[0007] Therefore, the present disclosure proposes a sound processing device, a sound processing method, and a sound processing program that can realize a sound space that is more immersive for the user in content in a virtual space.

[0008] It should be noted that the above problem or object is merely one of multiple problems or objects that can be solved or achieved by multiple embodiments disclosed in this specification.

[0009] In order to solve the above problems, one form of sound processing device according to the present disclosure includes an acquisition unit that acquires information regarding the movement of a character that serves as a sound receiving point of a sound source in a virtual space, a correction unit that corrects parameters used to generate an audio signal at the sound receiving point based on the information regarding the movement of the character, and a generation unit that generates an audio signal at the sound receiving point based on the parameters corrected by the correction unit.

[0010] 1 is a diagram illustrating an overview of information processing according to an embodiment. FIG. 1 is a flowchart illustrating audio signal generation processing. FIG. 2 is a diagram illustrating audio simulation in a virtual space. FIG. 3 is a diagram illustrating audio simulation in a virtual space. FIG. 4 is a diagram illustrating reflection in an audio simulation. FIG. 5 is a diagram illustrating diffraction in an audio simulation. FIG. 6 is a diagram illustrating transmission in an audio simulation. FIG. 7 is a diagram illustrating an example of a synthesized and output audio signal. FIG. 8 is a diagram illustrating a camera field of view in an audio simulation. FIG. 9 is a diagram illustrating an example configuration of an audio processing device according to an embodiment. FIG. 10 is a flowchart illustrating the procedure of audio processing according to an embodiment. FIG. 11 is a diagram illustrating an overview of sound receiving point correction processing. FIG. 12 is a flowchart illustrating the procedure of sound receiving point correction processing. FIG. 13 is a diagram illustrating a localization difference accompanying update of sound receiving point coordinates. FIG. 14 is a diagram illustrating calculation processing of a movement amount limit coefficient. FIG. 15 is a diagram illustrating calculation processing of a movement amount limit coefficient. FIG. 16 is a diagram illustrating a movement limit range of a sound receiving point. FIG. 17 is a diagram illustrating a modified example of sound receiving point correction processing. FIG. 18 is a diagram illustrating a modified example of sound receiving point correction processing. FIG. 2 is a diagram (2) for explaining parameter emphasis processing in space. FIG. 3 is a diagram showing a specific example of parameter adjustment processing. FIG. 4 is a diagram for explaining near-wall representation. FIG. 5 is a diagram for explaining frequency changes according to head rotation. FIG. 6 is a diagram showing an overview of acoustic representation according to a modified example. FIG. 7 is a diagram showing a flow of acoustic processing according to a modified example. FIG. 8 is a diagram showing an example of setting a parent listener. FIG. 9 is a diagram showing an example of setting a child listener. FIG. 10 is a diagram showing an example of the configuration of acoustic processing according to a modified example. FIG. 11 is a diagram showing an example of listener generation processing. FIG. 12 is a diagram showing an example of listener generation processing. FIG. 13 is a diagram showing a hardware configuration showing an example of a computer that realizes the functions of a sound processing device.

[0011] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In the following embodiments, the same components are designated by the same reference numerals, and redundant description will be omitted.

[0012] One or more embodiments (including examples and modified examples) described below can be implemented independently. However, at least a portion of the embodiments described below may be implemented in appropriate combination with at least a portion of another embodiment. These embodiments may include novel features that are different from one another. Therefore, these embodiments may contribute to solving different purposes or problems and may produce different effects.

[0013] The present disclosure will be described in the following order: 1. Overview 1-1. Overview of Acoustic Signal Generation Processing 1-2. Audio Signal Generation Processing by Acoustic Simulation 2. Acoustic Processing According to Embodiment 2-1. Overview of Acoustic Processing According to Embodiment 2-2. Configuration of Acoustic Processing Device 2-3. Correction Processing Related to Sound Receiving Points 2-4. Details of Correction Processing Related to Sound Receiving Points 2-5. Modified Example of Correction Processing Related to Sound Receiving Points 2-6. Direct Effect Parameter Control 2-7. Representation of the Vicinity of a Wall 3. Modified Examples 3-1. Use of Movement Information 3-2. Accessibility 3-3. Use of Artificial Intelligence, Machine Learning, Cloud, etc. 3-4. Parameters 3-5. Acquisition of Object Data from a Library 3-6. Speech Synthesis Processing 3-7. Application Example of the Technology of the Present Disclosure 3-8. Another Example of Acoustic Expression 3-8-1. Overview 3-8-2. Flow of Acoustic Processing 3-8-3. Other configurations of acoustic processing 3-8-4. Creation of listeners 3-8-5. Examples of use 3-8-6. Processing variations 4. Other modifications 5. Effects of the acoustic processing method according to the present disclosure 6. Hardware configuration examples

[0014] (1. Overview) (1-1. Overview of Sound Signal Generation Processing) Information processing according to the present disclosure is executed by a sound processing device 100 shown in Fig. 1. Prior to describing the sound processing according to the embodiment, an overview of sound signal generation processing in a virtual space will first be described with reference to Figs. 1 to 5.

[0015] 1 is a diagram illustrating an overview of information processing according to an embodiment. A sound processing device 100 is an example of a sound processing device according to the present disclosure. The sound processing device 100 is an information processing device such as a personal computer (PC), a server device, a game console, or a tablet terminal. The sound processing device 100 is used by a user U1 who uses or creates content related to virtual space (e.g., a game or a metaverse).

[0016] The following description focuses on processing audio information in a virtual three-dimensional space in game content, which is interactive content operated by a user. Note that the sound processing device 100 of the present disclosure can also be used to create audio for video content, create content that includes audio information that reproduces a sound field in a virtual space, or present audio information.

[0017] The sound processing device 100 has output units such as a display and a speaker, and outputs various information to the user U1. For example, the sound processing device 100 displays a virtual three-dimensional space (hereinafter simply referred to as a "virtual space") of a game or the like on the display, and outputs an audio signal (Audio Signal) generated by the information processing of the embodiment from the speaker. Alternatively, the sound processing device 100 displays a user interface of software related to sound production on the display, and outputs an audio signal generated in accordance with an operation of the user U1 from the speaker.

[0018] In the embodiment, the sound processing device 100 calculates how a sound output from a sound source object, which is a sound emission point, will be reproduced at a listening point (sound receiving point) in a virtual space, generates an audio signal based on the calculation result, and reproduces the generated sound. For example, the sound processing device 100 performs an acoustic simulation in the virtual space, and performs processing to make the sounds emitted in the virtual space closer to those in the real world or to reproduce the reverberation and tone color desired by the sound producer (creator).

[0019] FIG. 1 shows a virtual space V1 in a game. For example, the virtual space V1 is displayed on a display provided in the sound processing device 100. In the virtual space V1, a listening point (sound receiving point) is set along with the position (coordinates) of a sound source object (sound generating point). The listening point is the position where a user U1 virtually hears sound in the virtual space V11. In real spaces, various physical phenomena cause differences between sounds observed near the sound source and sounds observed at the listening point. Therefore, the sound processing device 100 virtually reproduces (simulates) real physical phenomena in the virtual space V1 and generates audio signals appropriate for the space to enhance the realism of the sounds experienced by the user U1 in the virtual space V1.

[0020] Here, the audio signal generated by the sound processing device 100 will be described. Graph G1 schematically shows the sound intensity when a pulsed sound emitted from a sound source object is observed at a listening point. At the listening point, the direct sound is observed first, followed by diffracted sounds of the direct sound, etc. Then, at the listening point, the first-order reflected sound reflected at the boundary of the virtual space and the transmitted sound that has passed through the object are observed. Reflected sounds are observed every time the sound reflects at a boundary, and for example, first- to third-order reflected sounds, which are considered to be early reflected sounds, are observed. Then, at the listening point, higher-order reflected sounds, which are considered to be late reverberation sounds, are observed. Because the sound emitted from the sound source decays over time, graph G1 depicts an envelope (decay curve) that asymptotically approaches 0, with the direct sound at its peak.

[0021] (1-2. Audio signal generation processing by acoustic simulation) Next, the audio signal generation processing executed by the audio processing device 100 will be described with reference to the drawings. First, a general generation processing of an audio signal in a virtual space will be described with reference to Figs. 2 to 6B. Then, the audio processing according to the embodiment will be described with reference to Figs. 7 and subsequent figures.

[0022] 2 is a flowchart showing the audio signal generation process, which is initiated by an event occurring in the virtual three-dimensional space (for example, an event occurring as a result of an operation by the user U1 or the progress of the game).

[0023] In the following description, the sound processing device 100 is assumed to execute the audio signal generation process, but the audio signal generation process may be executed by a single computer or by multiple computers working together. For example, one computer (e.g., a server device) among multiple computers (e.g., a user terminal and a server device) connected via a network may execute some steps of the audio signal generation process, and another computer (e.g., a user terminal) may execute the remaining steps.

[0024] In the following description, a scene in a game will be assumed as an example. Specifically, a virtual space V01 shown in FIGS. 3A and 3B will be assumed. FIGS. 3A and 3B are diagrams for explaining acoustic simulation in a virtual space. The virtual space V01 is a virtual three-dimensional space for a game or the like. FIG. 3A is a perspective view of the virtual space V01, and FIG. 3B is a plan view of the virtual space V01. Note that, in the following description, an XYZ coordinate system may be used for ease of understanding. In the drawings, the X-axis and Y-axis directions are horizontal, and the Z-axis direction is vertical.

[0025] In the virtual space V01, a sound emission point, which is a position that serves as a sound source S01, and a listening point, which is a sound receiving point T01, are set. The sound source S01 is, for example, an object that emits any sound toward the sound receiving point T01 (i.e., an object that can serve as a sound source). For example, the sound emission point is the coordinates at which the sound source object is located. Furthermore, the sound receiving point T01 is a position at which the sound output from the sound source S01 is observed. For example, the sound receiving point T01 is the position at which a game character operated by the user U1 is located (more precisely, the coordinates of a position corresponding to the head of the game character).

[0026] Furthermore, a plurality of three-dimensional objects are placed in the virtual space V01. For example, an object J2, which is a furniture-type object, and an object J3, which is a human-type object, are placed in the virtual space V01. Furthermore, a wall J1 and the like, which serve as a boundary that forms the virtual space V, are also placed in the virtual space V. Since acoustic characteristics, as will be described later, are set for boundaries such as the wall J1, in the embodiment, the boundary such as the wall J1 is also treated as one of the objects placed in the virtual space V.

[0027] 3A and 3B, sound is generated at the position of a sound source S01. The sound processing device 100 generates an audio signal at a sound receiving point T01 while taking into account the propagation characteristics of sound in a virtual space V. The audio signal generation process of the embodiment will be described below with reference to the flowchart of FIG.

[0028] First, the sound processing device 100 acquires audio data of a sound emitted from the sound source S01 (step S101). When a vocal event is initiated in the sound source S01, the sound processing device 100 acquires pre-recorded audio data of the sound to be produced. The audio data is, for example, audio data recorded in an in-game library. Considering that the audio data will undergo subsequent signal processing, data that does not include reverberation (dry source) is preferable. Alternatively, the audio data may be a simplified sound presentation that omits some of the subsequent signal processing, or a wet source that adds reverberation to express a distinctive tone. Furthermore, the sound processing device 100 may not only reference audio data from a library, but also acquire audio data generated on demand, such as by a synthesizer. Alternatively, the sound processing device 100 may acquire audio data generated on demand by simulating audio production from a structural physical simulation using FEM (Finite Element Method) or the like.

[0029] Next, the sound processing device 100 acquires spatial data of the space in which the sound source S01 and the sound receiving point T01 exist (step S102). For example, the sound processing device 100 acquires three-dimensional data of the virtual space V01 and metadata (e.g., material information) attached to the three-dimensional data. The three-dimensional data includes, for example, coordinates indicating the boundary shape of the walls of the virtual space V01 and coordinates of objects placed within the virtual space V01. If the sound source S01 and the sound receiving point T01 exist in a closed space, the sound processing device 100 acquires three-dimensional data of the entire closed space, or three-dimensional data of the space on the path from the sound source S01 to the sound receiving point T01 and the entire space connected to that space. Note that if the three-dimensional data includes detailed textures for video display, the data volume will increase. Therefore, the sound processing device 100 may acquire data that simplifies color information, surface fine shapes, etc., in a form that maintains the acoustic parameters required for subsequent physical simulations. The three-dimensional data used in the route search described below is preferably polygon information that does not include minute surface irregularities (for example, polygon information without texture applied).

[0030] Next, the sound processing device 100 calculates the sound path from the sound source S01 to the sound receiving point T01 (step S103). For example, the sound processing device 100 calculates the sound propagation path from the sound source S01 to the sound receiving point T01 using the acquired three-dimensional data. At this time, the path of the direct sound is a line-of-sight path from the sound source S01 to the sound receiving point T01. Note that if a transmission phenomenon such as a wall is included, the path including the obstacle is calculated even if there is no line-of-sight path. In the examples of Figures 3A and 3B, the propagation path of the direct sound from the sound source S01 to the sound receiving point T01 is shown by a solid line.

[0031] Furthermore, the sound processing device 100 calculates a path that includes reflection boundaries and the like along which the early reflected sound reaches the sound receiving point T01. For example, the sound processing device 100 determines the propagation path of the early reflected sound using a ray tracing method, which is a geometric physical simulation method. Note that in the embodiment, the early reflected sound includes not only sound that reaches the sound receiving point T01 after reflecting once on a boundary surface, but also sound that reaches the sound receiving point T01 after reflecting a small number of times, such as two or three times.

[0032] FIG. 4A shows an example of propagation paths calculated by the sound processing device 100. FIG. 4A is a diagram schematically illustrating reflections in a sound simulation. In the example shown in FIG. 4A, path R1 is the propagation path of a direct sound from a sound source S01 to a sound receiving point T01. Path R2 is the propagation path of a primary reflected sound that reflects once at a boundary B1 and reaches the sound receiving point T01. Path R3 is the propagation path of a secondary reflected sound that reflects twice at boundaries B2 and B3 and reaches the sound receiving point T01. In FIG. 4A, paths R1 to R3 are illustrated as propagation paths. However, actual calculations also include reflections from floors, ceilings, other wall surfaces, object surfaces, and the like.

[0033] Diffracted sound reaching the sound receiving point T01 from the sound source S01 can be obtained not only by the shortest path from the sound source S01 to the sound receiving point T01, but also by a path that detours around the edge of an obstacle located between the sound source S01 and the sound receiving point T01. In the examples of Figures 3A and 3B, object J3 is an example of an obstacle located between the sound source S01 and the sound receiving point T01. In this case, the sound processing device 100 can, for example, obtain a detour route that detours around the obstacle in the shape of an arc, or can obtain the detour route by a method that connects the sound source S01 and the edge of the object J3, and the edge of the object J3 and the sound receiving point T01, with a straight line.

[0034] When directivity is set for the sound source, the sound processing device 100 may omit sound rays in directions where there is no sound radiation according to the directivity, or may change the density of sound rays according to the strength of radiation. By performing this processing, it is possible to reduce the impact on auditory sensation while reducing subsequent processing performed for each path.

[0035] Based on the obtained information, the sound processing device 100 generates an audio signal to be heard at the sound receiving point T01. For example, the sound processing device 100 performs generation processes such as direct sound generation, early reflected sound generation, diffracted sound generation, and late reverberation sound generation. These generation processes will be described below.

[0036] First, the sound processing device 100 generates a sound signal based on the direct sound among the sound signals heard at the sound receiving point T01 (step S104). Specifically, the sound processing device 100 receives the acquired sound data as an input and generates a sound signal using parameters such as the spatial size of the sound source, directivity, attenuation according to the distance from the sound source to the listening point, and a deviation in observation time due to propagation time. For example, if the sound source S01 is an omnidirectional point sound source and the distance between the sound source S01 and the sound receiving point T01 is x, the attenuation is 20log 10 x. The propagation time can be calculated by dividing the distance x by the speed of sound c. If a content creator wishes to emphasize the sense of distance between the sound source S01 and the sound receiving point T01, the sound processing device 100 can change the amount of attenuation without being bound by physical phenomena in the real space.

[0037] Next, the sound processing device 100 generates a sound signal based on the early reflection sound from the sound signal heard at the sound receiving point T01 (step S105). Specifically, the sound processing device 100 receives the acquired sound data as input, and calculates the attenuation and time lag according to the propagation path length in the medium based on the propagation path calculated in advance, as in the case of direct sound. Furthermore, the sound processing device 100 performs signal processing that simulates reflection at a boundary surface based on information about the boundary surface that reflects the input signal. For example, the sound processing device 100 may provide a predetermined attenuation for each frequency depending on the sound absorption coefficient of the boundary surface that reflects the input signal.

[0038] For example, for the reflected sound corresponding to path R2 shown in FIG. 4A , the section from the sound source S01 to the boundary B1 and the section from the boundary B1 to the sound receiving point T01 can be considered as the propagation path of the medium. The sound processing device 100 calculates the attenuation and time lag from the distances of each of these propagation paths. Note that the sound processing device 100 may calculate attenuation by storing sound absorption coefficient data for each frequency for boundary B1 in a library in advance and appropriately reading out that data. The sound processing device 100 may generate an audio signal of the reflected sound corresponding to path R2 by applying attenuation according to the sound absorption coefficient to the input signal in addition to the attenuation on the previous propagation path. In this case, if the effects of plate vibration and resonance at boundary B1 and nonlinear phenomena are not taken into consideration, the sound processing device 100 may calculate attenuation, etc. by combining the section from the sound source S01 to the boundary B1 and the propagation path from boundary B1 to the sound receiving point T01.

[0039] Similarly, the sound processing device 100 performs calculations for reflected sound corresponding to path R3 shown in FIG. 4A . That is, the sound processing device 100 performs calculations according to the propagation paths from the sound source S01 to boundary B2, from boundary B2 to boundary B3, and from boundary B3 to sound receiving point T01. When calculating sound reflected at multiple boundary surfaces such as boundary B2 and boundary B3, the sound processing device 100 may perform calculations of attenuation according to the sound absorption coefficient all at once if nonlinear behavior according to the sound pressure at the boundary surfaces is not taken into consideration. A digital filter such as an IIR (Infinite Impulse Response) or FIR (Finite Impulse Response) filter may be used to calculate the attenuation for each frequency according to the sound absorption coefficient. In the process of generating early reflected sound, the sound processing device 100 may appropriately use a technique such as ray tracing to simulate the reflection from each wall surface.

[0040] Next, the sound processing device 100 generates an audio signal related to the diffracted sound (step S106). FIG. 4B is a diagram schematically illustrating the state of diffraction in an acoustic simulation. As shown in FIG. 4B, diffraction is a phenomenon in which sound reaches the sound receiving point T01 by circling around the diffraction point even in a situation where there is no line of sight between the sound source S01 and the sound receiving point T01. As with early reflection sounds, the sound processing device 100 generates an audio signal related to the diffracted sound using a geometric method that calculates attenuation and time lag based on the propagation path. It is known that the attenuation of diffracted sound increases as the frequency increases in the sound diffraction phenomenon. Therefore, the sound processing device 100 may generate diffracted sound by applying a filter with different attenuation amounts depending on the frequency. It is noted that if a more accurate method based on the physical laws of wave phenomena is used, the sound processing device 100 can use not only the geometric simulation method described above, but also a wave sound simulation method from the sound source S01 to the sound receiving point T01. In this case, the amount of calculation may be greater than in geometric simulation, so user U1 may use equipment with sufficient computing power as the sound processing device 100, or may prepare a library in which the characteristics of typical propagation path shapes have been calculated in advance.

[0041] Next, the sound processing device 100 generates an audio signal related to the transmitted sound (step S107). FIG. 4C is a diagram schematically illustrating the state of transmission in the sound simulation. As shown in FIG. 4C , transmission is a phenomenon in which sound passes through the object J5 and reaches the sound receiving point T01 even in a scene in which the sound source S01 and the sound receiving point T01 are not visible due to the presence of an object J5 or the like between the sound source S01 and the sound receiving point T01. When sound transmission is set for the object J5, the sound processing device 100 calculates sound synthesis using a path including the transmitted sound. For example, the sound processing device 100 calculates attenuation and delay based on the propagation path, similar to the processing of direct sound. Specifically, the sound processing device 100 performs filter processing simulating the transmission phenomenon on the object J5 for which transmission is set, to generate transmitted sound.

[0042] Next, the sound processing device 100 generates late reverberation sounds (step S108). Late reverberation sounds refer to the sound emitted from the sound source S01, which is repeatedly reflected and diffracted in space and reaches the sound receiving point T01, excluding the early reflected sound. This reflection and diffraction continues until the sound is completely attenuated (for example, until the quantization error on the computer can be considered to be zero). Here, the completely attenuated sound may be a state that the listener perceives as a limit at which the reproduced sound including the late reverberation sounds generated by the sound processing device 100 can be discriminated.

[0043] In a simulation, it is also possible to perform calculations until the sound is completely attenuated. However, since the amount of calculation increases with each reflection, the sound processing device 100 may calculate the early reflection sound by regarding up to a predetermined number of reflections as early reflections. The sound processing device 100 may then calculate subsequent sounds (late reverberation sounds) using a different method and combine them with the early reflection sound, etc. Here, the predetermined number of reflections included in the early reflection sound may be set by the creator, or may be dynamically changed or set within the sound processing device 100 depending on the type of content or virtual space to be reproduced, the size of the space, the sound absorption coefficient of boundary surfaces within the space, object surfaces, etc.

[0044] In the embodiment, the sound processing device 100 generates a late reverberation sound using, for example, a short-time impulse response, thereby generating a realistic late reverberation sound while reducing the calculation load.

[0045] In addition, methods for calculating late reverberation sounds are also known in the field of architectural acoustics, such as a statistical method for calculating reverberation time based on the sound absorption coefficients of the space's size, boundary surfaces within the space, and object surfaces. When the sound processing device 100 does not perform the processing described above, it may generate late reverberation sounds using these known techniques. For example, the sound processing device 100 may use a comb filter that uses acquired audio data as input, multiplies the input by a predetermined gain, and feeds back the input signal, causing the input signal to decay and repeat at a constant cycle, or a method that convolves the input signal with impulse response data stored in a library or the like. As described above, late reverberation sounds have a significant impact on a user's auditory perception of a space and are therefore important in content creation. Many existing modular reverberation generators allow adjustment of late reverberation settings (parameters), such as the late reverberation level, late reverberation delay time, decay time, frequency-specific decay ratio, echo density, and modal density.

[0046] The sound processing device 100 generates direct sound, early reflection sound, diffracted sound, transmitted sound, and late reverberation sound, and then synthesizes these signals (step S109). For example, the sound processing device 100 converts the direct sound, early reflection sound, diffracted sound, and transmitted sound, which require consideration of the sound's arrival direction, as well as the late reverberation sound generated using some processing methods, into signals tailored to the actual playback environment via HOA (Higher Order Ambisonics) or the like. For example, in a 7.1-channel speaker environment, the signals tailored to the actual playback environment are channel-based signals with a total of eight channels. Alternatively, in a headphone environment, the signals tailored to the actual playback environment are binaural signals that take into account HRTF (Head Related Transfer Function). Furthermore, for sounds for which directionality is not considered, the sound processing device 100 directly adjusts the volume balance and the processing and timing of sounds that require consideration of directionality, and synthesizes them into signals tailored to the actual playback environment.

[0047] The sound processing device 100 outputs the synthesized sound signal to a speaker or the like as a sound observed at the sound receiving point T01 (step S110). Fig. 5 is a diagram showing an example of the synthesized and output sound signal. When the output is completed, the sound processing device 100 ends the sound signal generation process.

[0048] The generation processes from step S104 to step S107 above may be switched around because there is no dependency on the order of the steps. When considering the case where diffraction occurs after reflection in the propagation path, the signal generated by first calculating the reflection attenuation at the boundary from the sound source may be used as input for calculating the diffracted sound.

[0049] 3A and the like, the sound processing device 100 may add processing using a porting simulation technique that simulates sound transmission through a boundary surface or the characteristics of a small sound passing through a space such as a window or door of a building, thereby enabling the sound processing device 100 to realize sound representation that is closer to phenomena occurring in real space.

[0050] Furthermore, if there are multiple reflected sounds, diffracted sounds, and transmitted sounds on the path, the sound processing device 100 may perform these calculations for each location where each phenomenon occurs, and after all calculations have been completed, synthesize and output the audio signal.

[0051] Furthermore, when there are multiple sound sources S01, the sound processing device 100 may perform all of the above-described processing in parallel for the sounds output from the multiple sound sources S01. Alternatively, the sound processing device 100 may perform sequential processing by delaying the time until output until all processing is completed, and further combine and output each combined signal in which the time series of the sounds arriving at the sound receiving point T01 are aligned.

[0052] (2. Acoustic Processing According to Embodiment) (2-1. Overview of Acoustic Processing According to Embodiment) Next, acoustic processing for realizing a more immersive acoustic space for the user in content in a virtual space, which is a feature of the embodiment, will be described.

[0053] As described above, a user may recognize changes in the environment in a virtual space through changes in sound. Therefore, if changes in sound due to changes in the environment can be appropriately emphasized, the user's spatial awareness can be further enhanced. Specifically, when a character controlled by the user moves through a space during a game, the user can sense from the changes in sound whether the character is currently in a large space, or whether the space to which the character has moved has become large. The sound processing device 100 emphasizes the changes in sound by appropriately correcting the sound generation process so that the user can easily recognize these changes in sound.

[0054] For example, in content such as games, whether the user recognizes the object that is the sound source, i.e., whether the sound source is included in the user's field of view, etc., affects the user's perception of sound. For this reason, the sound processing device 100 may acquire information regarding the user's field of view, etc., before performing sound synthesis processing.

[0055] This point will be explained with reference to Fig. 6A, which is a diagram schematically showing the field of view of a camera in an acoustic simulation.

[0056] 6A shows the field of view displayed by the virtual camera when the character operated by the user is object J6 in the virtual space V01 shown in FIG. 3A etc. Specifically, the screen display 50 shows that the object J6 is viewing an object serving as sound source S01, located in the upper left corner of a desk 51. Note that the screen display 60 shows the viewpoint of the screen display 50 as a plan view.

[0057] The field of view displayed by the virtual camera is the first-person perspective of the object J6. That is, in this disclosure, the virtual space V01 is assumed to be a so-called FPV (First Person View) game content. However, the processing according to the embodiment is not limited to the first-person perspective, and can also be applied to content such as a so-called TPV (Third Person View).

[0058] In FPV, the virtual camera is set, for example, at the center between the eyes of the character controlled by the user. Therefore, when the user changes the character's orientation, posture, position, etc., the camera's orientation, position, etc. also change to track the character's movement. Note that while the screen display 60 in FIG. 6A schematically shows the range of the field of view of the object J6, the character's field of view angle is not limited to this. For example, the virtual camera can be set arbitrarily, such as to have different horizontal and elevation field of view angles depending on the settings of the content creator. As will be described in detail later, the sound processing device 100 uses information about the camera in the sound receiving point coordinate correction process, and therefore, for example, acquires information about the camera prior to the process of step S104 shown in FIG. 2. The information about the camera refers, for example, to the orientation of the camera, which means the direction of the center of the field of view displayed on the screen.

[0059] The sound processing device 100 also acquires information about the movement of the character corresponding to the field of view of the virtual camera. The information about the movement of the character is, for example, a movement vector when the character is moving. The movement vector includes, for example, information about the direction and speed of the movement of the character.

[0060] As described above, in the embodiment, the sound processing device 100 corrects sound generation by adjusting various parameters for synthesizing sound in accordance with the movement of a character. For example, the sound processing device 100 executes a process for correcting parameters that directly affect sound generation (hereinafter referred to as "direct effect parameter control") and a process for correcting parameters that indirectly affect sound generation (hereinafter referred to as "indirect effect parameter control").

[0061] Direct effect parameter control refers to correcting parameters that are directly used in generating sound (such as volume (gain) and reverberation time). Indirect effect parameter control refers to changing the sound perceived by the user by correcting the coordinates of the sound receiving point, for example.

[0062] In the procedure shown in FIG. 2, for example, the correction process is executed in a process prior to the calculation of the propagation of direct sound and the like (step S104).

[0063] In addition to the correction process, the sound processing device 100 may also perform a process that changes the sound depending on the location of the character in the virtual space. Specifically, when a character is near an object that reflects sound (for example, when the character is near a wall), the sound processing device 100 outputs sound that emphasizes the effect of such reflection. This allows the user to perceive a sound that more closely resembles a real-world sound environment. Hereinafter, the above correction process may be referred to as "wall proximity representation." In the procedure shown in FIG. 2, this process is performed, for example, after the synthesis process (step S109).

[0064] By constructing a sound environment in a virtual space through the various correction processes described above, the sound processing device 100 can provide the user with an immersive sound environment that allows the user to more accurately perceive the environment through changes in sound.

[0065] (2-2. Configuration of Sound Processing Device) Next, a configuration of the sound processing device 100 for realizing sound processing according to the embodiment will be described. Fig. 6B is a diagram showing an example of the configuration of the sound processing device 100 according to the embodiment.

[0066] In addition, the term "acoustic processing" that appears in the following description can be replaced with "acoustic signal processing," "audio processing," "audio signal processing," "sound processing," "sound signal processing," "speech / voice processing," or "speech / voice signal processing."

[0067] In the following description, the term "audio signal" may be replaced with "acoustic signal," "sound signal," or "speech / voice signal." Similarly, the term "audio data" may be replaced with "acoustic data," "sound data," or "speech / voice data."

[0068] The sound processing device 100 is an information processing device (computer) that performs processing related to content relating to virtual space (for example, games, the Metaverse, etc.) For example, the sound processing device 100 is a user terminal used for games, the Metaverse, etc.

[0069] The sound processing device 100 is typically a game console (e.g., a dedicated game console), but is not limited to a game console. Any type of computer can be used for the sound processing device 100. For example, the sound processing device 100 may be a mobile terminal such as a mobile phone, a smart device (smartphone or tablet), a PDA (Personal Digital Assistant), or a notebook PC. The sound processing device 100 may also be an imaging device or a car navigation device. The sound processing device 100 may also be an M2M (Machine to Machine) device or an IoT (Internet of Things) device. The sound processing device 100 may also be a wearable device such as a smart watch. As long as a user can play a game on the device, these devices can also be considered game consoles (general-purpose game consoles).

[0070] The sound processing device 100 may also be an xR device such as an AR (Augmented Reality) device, a VR (Virtual Reality) device, or an MR (Mixed Reality) device. In this case, the xR device may be a glasses-type device such as AR glasses or MR glasses, or a head-mounted device such as a VR head-mounted display. The xR device may be an optical see-through type, a video see-through type, or any other type. When the sound processing device 100 is an xR device, the sound processing device 100 may be a standalone device consisting only of a user-worn part (e.g., glasses). The sound processing device 100 may also be a terminal-linked device consisting of a user-worn part (e.g., glasses) and a terminal part (e.g., a smart device) linked to the user-worn part. In this case, the sound processing device 100 may also be an information processing device (e.g., a game console) connected to a user-worn part (e.g., a head-mounted display) via a wired or wireless connection. If a user can play a game on the device, these devices can also be considered game consoles (general-purpose game consoles).

[0071] The sound processing device 100 may also be an information processing device connected to a user terminal and transmitting the results of processing related to the virtual space to the user terminal. For example, the sound processing device 100 may be a server device connected to a user terminal (e.g., a game console) via a network. In this case, the sound processing device 100 may be an application server or a web server. The sound processing device 100 may also be a PC server, a mid-range server, or a mainframe server. The sound processing device 100 may also be an information processing device that performs data processing (edge ​​processing) near the user terminal. For example, the sound processing device 100 may be an information processing device attached to or built into a base station. Of course, the sound processing device 100 may also be an information processing device that performs cloud computing.

[0072] As shown in Fig. 6B , the sound processing device 100 includes a communication unit 110, a storage unit 120, a control unit 130, an input unit 140, and an output unit 150. Note that the configuration shown in Fig. 6B is a functional configuration, and the hardware configuration may be different from this. Furthermore, the functions of the sound processing device 100 may be statically or dynamically distributed and implemented in multiple physically separated configurations. For example, the sound processing device 100 may be configured by multiple information processing devices (computers) connected by communication.

[0073] The communication unit 110 is a communication interface for communicating with other devices. The communication unit 110 may be a network interface or a device connection interface. For example, the communication unit 110 may be a LAN (Local Area Network) interface such as a NIC (Network Interface Card), or a USB (Universal Serial Bus) interface configured by a USB host controller, a USB port, etc. The communication unit 110 may be a wired interface or a wireless interface. The communication unit 110 exchanges information with other information devices, etc. via the network N.

[0074] Here, the network N is a communication network such as a LAN, a WAN (Wide Area Network), a cellular network, a fixed telephone network, a regional IP (Internet Protocol) network, or the Internet. The network N may include a wired network or a wireless network. The network N may also include a core network. The core network is, for example, an EPC (Evolved Packet Core) or a 5GC (5G Core network). Of course, the network N may also be a data network connected to the core network. The data network may be a service network of a telecommunications carrier, for example, an IMS (IP Multimedia Subsystem) network. The data network may also be a private network such as an in-house network or a home network.

[0075] The storage unit 120 is a storage device that stores various types of information. For example, the storage unit 120 is a data readable / writable storage device such as a dynamic random access memory (DRAM), a static random access memory (SRAM), a semiconductor memory (e.g., a flash memory), or a hard disk. The storage unit 120 may be an optical drive such as a Blu-ray (registered trademark) drive, a DVD drive, or a CD drive. The storage unit 120 stores information for processing related to virtual space (hereinafter referred to as "virtual space information"). For example, the storage unit 120 stores sound source information, object information, sound receiving point information, and setting information as virtual space information.

[0076] The sound source information is, for example, position information of the sound emission point S, data of the sound output from the sound emission point S (for example, waveform data), and information on the directivity of the sound output from the sound emission point S. The sound source information may also include information on the type of sound. If the sound is dialogue, the sound source information may include data on the character emitting the sound, and metadata such as angry voices or laughter. Furthermore, the object information is, for example, information on the position, shape, material, sound absorption coefficient, and acoustic impedance of the object.

[0077] The sound receiving point information is, for example, information such as the position of the sound receiving point T01, directional characteristics, and HRTF (Head Related Transfer Function) that expresses the characteristics from the sound source S01 to both ears of the listener. The positions of the sound generating point S and the sound receiving point T01 may be expressed using a Cartesian coordinate system, a cylindrical coordinate system, or a spherical coordinate system.

[0078] The setting information includes information about a player that plays content related to the virtual space, information about a platform for creating the content, and specific scene information within the content. The virtual space information may also include, for example, CAD data of a virtual space V01 arbitrarily created by a creator, as shown in FIG. 1 , or CAD data of a virtual space V01 in a scene from a game in which a user U1 controls a character in the game. In addition to CAD data, the virtual space information may also include voxel data, mesh data, point cloud data, etc. The sound source information and object information may also include, as metadata, information indicating the instrument type of the sound source (vocals, guitar, bass, drums, etc.), priority information set for each object, and information indicating the spatial extent of each sound source.

[0079] Note that the information stored in the storage unit 120 is not limited to these pieces of information. The virtual space information stored in the storage unit 120 may also include setting information related to sound settings (for example, setting information related to diffracted sound, reflected sound, transmitted sound, and late reverberation sound). The virtual space information may be input to the storage unit 120 via the communication unit 110 or the input unit 140.

[0080] The control unit 130 is a controller that controls each unit of the sound processing device 100. The control unit 130 is realized by a processor such as a CPU (Central Processing Unit), MPU (Micro Processing Unit), GPU (Graphics Processing Unit), APU (Accelerated Processing Unit), or DSP (Digital Signal Processor). For example, the control unit 130 is realized by executing various programs (e.g., sound processing programs according to the present disclosure) stored in a storage device inside the sound processing device 100 using RAM (Random Access Memory) or the like as a working area. The CPU, MPU, GPU, APU, ASIC, and FPGA can all be considered as controllers.

[0081] As shown in FIG. 6B , the control unit 130 includes an acquisition unit 131, a correction unit 132, a generation unit 133, and an output control unit 134. Each block constituting the control unit 130 (from the acquisition unit 131 to the output control unit 134) is a functional block that represents a function of the control unit 130. These functional blocks may be software blocks or hardware blocks. For example, each of the above-described functional blocks may be a software module implemented by software (including a microprogram) or a circuit block on a semiconductor chip (die). Of course, each functional block may be a processor or an integrated circuit. The control unit 130 may be configured with functional units different from the above-described functional blocks. The functional blocks may be configured in any manner. In other words, in the present disclosure, the processing performed by the sound processing device 100 may be interpreted as the processing performed by each block. The operations of the acquisition unit 131, correction unit 132, and generation unit 133, among the blocks constituting the control unit 130, will be described later. The output control unit 134 outputs the audio signal synthesized by the operations of the acquisition unit 131, the correction unit 132, and the generation unit 133 to an output unit such as a speaker.

[0082] The input unit 140 is an input device that accepts various inputs from the outside. For example, the input unit 140 is a data input interface for inputting various data related to a network. In this case, the data input interface may be a device connection interface such as a USB interface. For example, the input unit 140 may be an operation device for a user to perform various operations, such as a keyboard, a pointing device (e.g., a mouse), or operation keys. The input unit 140 may be a microphone for a user to input sound. The input unit 140 may also be a camera for a user to input a line of sight or a gesture. If a touch panel is adopted as an input / output device for the sound processing device 100, the touch panel is also included in the input unit 140. In this case, the user performs various operations by touching the screen with a finger or a stylus.

[0083] The output unit 150 is a device that outputs various types of information to the outside, such as sound, light, vibration, and images. For example, the output unit 150 may be an acoustic device such as a speaker, or a display device such as a display. The display device may be, for example, a liquid crystal display or an organic electroluminescence display (OLED). The output unit 150 may be a touch panel display device. In this case, the input unit 140 and the output unit 150 may be considered to be an integrated configuration. The output unit 150 may also be an output unit of an xR device.

[0084] The output unit 150 performs various outputs to the user under the control of the control unit 130. For example, the output unit 150 outputs synthesized sound from headphones, earphones, hearing aids, sound collectors, speakers, sound bars, or a head mounted display (HMD) under the control of the output control unit 134. Alternatively, the output unit 150 displays a virtual space or a user interface on a display under the control of the output control unit 134.

[0085] (2-3. Correction Process for Sound Receiving Point) The acoustic process according to the embodiment is realized by adding a correction process for emphasizing sound changes to the voice synthesis process procedure shown in Fig. 2. The correction process will be described below in order.

[0086] First, the correction process for the sound receiving point will be described. This process is executed between steps S102 and S104 shown in Fig. 2. This procedure is shown in Fig. 7. Fig. 7 is a flowchart showing the procedure of the acoustic processing according to the embodiment.

[0087] 6A and other figures, the sound processing device 100 acquires spatial data (step S102), and then acquires a camera direction corresponding to the character's viewpoint and field of view (step S1021).The sound processing device 100 also acquires a character's movement vector as information related to the character's movement (step S1022).

[0088] As shown in FIG. 2, the sound processing device 100 calculates the propagation path from the sound source S01 to the sound receiving point T01 based on the acquired information (step S103).

[0089] Here, the sound processing device 100 determines whether the character is moving (step S1031). If the character is not moving (step S1031; No), for example, if the character's movement speed is 0, the sound processing device 100 skips the correction process for the sound receiving point shown in step S1032 and step S1033.

[0090] On the other hand, if the character is moving (step S1031; Yes), the sound processing device 100 performs a correction process for the sound receiving point, including updating the sound receiving point coordinates (step S1032). Specifically, the sound processing device 100 updates the coordinates by moving the coordinates of the sound receiving point from the original sound receiving point (i.e., the character's location) in the moving direction of the character according to the moving speed, based on the character's movement vector.

[0091] This processing will be explained with reference to Fig. 8. Fig. 8 is a diagram showing an outline of the sound receiving point correction processing.

[0092] The virtual space 74 in the right diagram of Figure 8 is a space that simulates a narrow corridor and a wide interior space that continues from the corridor in the game content. The right diagram of Figure 8 shows a character 76 moving from the narrow corridor into a wide space.

[0093] In the example shown on the right side of Fig. 8, a character 76 is moving from a narrow corridor toward an indoor space. In this case, the sound processing device 100 corrects the sound receiving point 72, which originally coincides with the location of the character 76, by shifting it from the character 76 so that it precedes the movement.

[0094] The screen display 70 on the left side of Figure 8 shows the field of view of a character 76 in a virtual space 74, looking at an indoor space from a narrow corridor. For the sake of explanation, a pseudo sound receiving point 72 is also depicted on the screen display 70. As shown in the screen display 70, the virtual sound receiving point 72 is located in the indoor space prior to the narrow corridor. Therefore, even when the character 76 is located in the narrow corridor, the user can hear the sound synthesized at the sound receiving point 72 placed in the indoor space, and can quickly obtain sound information about the indoor space.

[0095] In other words, by correcting the sound receiving point, the user can hear changes in sounds heard in the direction in which the character is moving earlier than usual. This allows the user to feel a stronger sense of spatial transition due to movement. Furthermore, by quickly obtaining sound information in the direction in which the character is moving as the game progresses, the user can quickly decide on the next action in the game.

[0096] When a correction is made to move the sound receiving point from the original position of the character, the sound processing device 100 calculates a new path from the sound source S01 to the new sound receiving point T01 and updates the path (step S1033).

[0097] Thereafter, the sound processing device 100 also performs a process of correcting parameters that directly affect the generation of sound (direct effect parameter control) (step S1034), the details of which will be described later.

[0098] (2-4. Details of Correction Processing for Sound Receiving Points) The above-mentioned correction processing for sound receiving points will be described in detail with reference to Fig. 9 and subsequent figures. Fig. 9 is a flowchart showing the procedure for the correction processing for sound receiving points.

[0099] First, the sound processing device 100 acquires spatial data related to the correction process in the virtual space (step S201). The spatial data related to the correction process is, for example, data related to the space where the character is located or the space to which the character is about to move. Note that the spatial data also includes the coordinates of the sound source reproducing the sound, the coordinates of the character, etc.

[0100] Next, the sound processing device 100 acquires a movement amount adjustment coefficient that has been set in advance by the content creator or the user (step S202). The movement amount adjustment coefficient is a coefficient used to determine the range (distance) by which the sound receiving point is shifted. For example, the adjustment coefficient is multiplied by the character movement vector to calculate the movement vector of the sound receiving point and to update the coordinates of the sound receiving point.

[0101] Thereafter, the sound processing device 100 acquires a movement vector of the character (step S203), and acquires path data of the character based on the movement vector (step S204).

[0102] Furthermore, the sound processing device 100 acquires a setting value for the allowable angle of localization difference, which is a setting value for determining the range of allowable movement of the sound receiving point (step S205). The localization difference here refers to the angular difference between the direction of movement of the character and the angle formed when the positions of the sound source and the character are connected. This point will be explained using FIG. 10. FIG. 10 is a diagram for explaining the localization difference accompanying the update of the sound receiving point coordinates.

[0103] 10 shows the positional relationship between a sound source 10 and a character 80. In content such as a game, the character 80 does not necessarily simply move toward or away from the direction of the sound source 10, but may move at some angle relative to the angle between the sound source 10 and the character 80. In such a case, an angle difference A occurs between the direction of the sound source 10 as seen from the original position of the character 80 and the position 82 of the character 80 after movement, which is calculated based on the movement vector of the sound receiving point.

[0104] Since the direction of sound localization is important information for the user to understand the space, if the angle difference A exceeds a certain amount due to correction, it will affect the user's understanding of the space. The setting value of the allowable angle for localization difference is a setting value for limiting the localization difference to a certain amount to prevent such an effect from occurring.

[0105] Next, the sound processing device 100 calculates a movement amount limit coefficient to limit the range of movement of the sound receiving point based on the acquired localization misalignment allowable angle setting value (step S206).

[0106] This process will be described with reference to Figures 11 and 12. Figures 11 and 12 are diagrams for explaining the process of calculating the movement amount restriction coefficient.

[0107] First, the sound processing device 100 calculates the movement vector of the sound receiving point, which is the basic advance amount, by multiplying the character movement vector by the movement amount adjustment coefficient. Furthermore, the sound processing device 100 calculates updated coordinates of the temporary sound receiving point using the current character coordinates and the movement vector of the sound receiving point. A temporary sound receiving point 86 shown in FIG. 12 indicates the coordinates of the calculated temporary sound receiving point.

[0108] Next, the sound processing device 100 calculates the localization difference (corresponding to angle A1 shown in FIG. 12 ) using the coordinates of the sound source 10, the coordinates of the character 80, and the coordinates of the virtual sound receiving point 86. In this case, the sound source 10 may include the final reflection point, diffraction point, or transmission point on a path that includes initial reflection, diffraction, or transmission, treated as a virtual sound source. If there are multiple sound sources as the sound source 10, the sound processing device 100 calculates the angle A1 corresponding to the localization difference for all the sound sources, and then uses the largest value among them in subsequent processing.

[0109] The sound processing device 100 uses, for example, a function as shown in graph 90 in Fig. 11 as a function for limiting angle A1. Specifically, the function shown in Fig. 11 is a function in which when the input is small, the output changes little from the input, and as the input increases, the amount of change increases so that the output becomes smaller than the input, gradually approaching a specified value. Note that the function is not limited to the example shown in Fig. 11 and may be any function as long as it acts to limit the input value (here, angle A1) to a smaller value.

[0110] The sound processing device 100 limits the angle A1 to a specified value or less (for example, angle A2 shown in FIG. 12 ) by the action of a function such as that shown in FIG. 11 . Then, the sound processing device 100 calculates a movement vector 84 indicating the limited movement based on the value of the limited angle A2. Furthermore, the sound processing device 100 calculates a sound receiving point movement amount limit coefficient by dividing the length of the movement vector 84 as the numerator by the length of the movement vector to the temporary sound receiving point 86 as the denominator.

[0111] Furthermore, as shown in FIG. 9, the sound processing device 100 calculates a coefficient related to the restriction of movement and acquires the camera direction related to the movement (step S207).

[0112] Then, the sound processing device 100 calculates a limit coefficient for the amount of movement of the sound receiving point coordinates according to the acquired camera direction (step S208).

[0113] This process is performed to prevent discrepancies between the visual perception of the user and the perception of sound at the corrected sound receiving point. In other words, in the sound receiving point correction process, the sound receiving point is moved in the character's moving direction in accordance with the character's movement, but if the sound receiving point moves out of the user's field of view, it may be difficult for the user to imagine the sound at the sound receiving point from the visual information, and the user may feel uncomfortable.

[0114] Therefore, the sound processing device 100 limits the length of the movement vector of the sound receiving point based on the information about the camera direction and the movement vector of the character. This processing will be explained using Fig. 13. Fig. 13 is a diagram for explaining the movement restriction range of the sound receiving point.

[0115] As shown in FIG. 13 , the sound processing device 100 sets an ellipse with its major axis in the movement direction of the character 80 as the movement range of the sound receiving point. The outer periphery of this ellipse is farthest from the character coordinates in the movement direction of the character 80, and this distance is set to "1." The distance between the outer periphery of the ellipse in the direction opposite to the character movement direction and the character coordinates can be thought of as the "rearward distance from the character position." That is, the major axis of the ellipse can be expressed as "1" + "rearward distance from the character position." Specifically, the sound processing device 100 sets a setting value (first setting value) in the forward / backward direction of the character as a combination of a fixed value, where the forward distance from the character position is 1, and an arbitrarily settable "rearward distance from the character position." When the movement limit range of the sound receiving point is set to an ellipse as shown in FIG. 13 , the first setting value corresponds to the major axis of the ellipse. Note that, as an example, the fixed value of the forward distance from the character position is set to 1, but this is not limiting. The fixed value may be dynamically changed by the sound processing device 100 or set by the user.

[0116] Furthermore, the sound processing device 100 sets the movement restriction range of the sound receiving point on the side of the character as a value (second setting value) of (1 + "distance behind from the character position") x (coefficient). When the movement restriction range of the sound receiving point is set to an elliptical shape as shown in FIG. 13, the second setting value corresponds to the minor axis of the ellipse. In this way, the "distance behind from the character position" and the "coefficient" are adjustment parameters for determining the shape of the ellipse. By adjusting these parameters, the sound processing device 100 can determine how much to restrict the movement of the sound receiving point in directions other than the camera direction. For example, the "coefficient" can take any value between "0" and "1."

[0117] If the "rear distance from the character position" is "1" and the "coefficient" is "1", there is no restriction on the amount of movement of the sound receiving point coordinates according to the camera direction. The creator of the game content or the user can arbitrarily set the movement restriction range of the sound receiving point by setting the above first and second setting values.

[0118] 13 indicates the length from the character coordinates to the intersection of the straight line extending from the character coordinates toward the camera and the ellipse. The distance 92 is a limiting coefficient for the amount of movement of the sound receiving point coordinates according to the camera direction.

[0119] In this example, the range for limiting the movement amount of the sound receiving point is shown as an ellipse, but the movement range of the sound receiving point can also take on a shape other than an ellipse, such as a rectangle or a cardioid, depending on the settings of the first and second setting values.

[0120] Returning to FIG. 9, the sound processing device 100 calculates the amount of movement of the sound receiving point based on the various coefficients and setting values ​​obtained by the above processing (step S209), and determines the coordinates of the corrected sound receiving point (step S210).

[0121] For example, in steps S206 and S208, the sound processing device 100 calculates the movement amount limit coefficient as a method for reducing the sense of discomfort felt by the user due to the movement of the sound receiving point. The sound processing device 100 uses the smaller of these values, multiplies it by the movement amount adjustment coefficient, and further multiplies it by the movement amount of the character to calculate the movement vector of the sound receiving point. Then, based on the movement vector of the sound receiving point, the sound processing device 100 determines the position to be corrected from the original position of the sound receiving point (i.e., the character coordinates).

[0122] Note that, if the sound processing device 100 wishes to further reduce the sense of discomfort felt by the user, the value by which the adjustment coefficient for the amount of movement is multiplied may be the value obtained by multiplying the adjustment coefficient for the amount of movement by the limit coefficient calculated in steps S206 and S208. This allows the sound processing device 100 to arbitrarily adjust the degree of discomfort felt by the user in accordance with the content of the game content and the user's sensitivity. Furthermore, in the above processing, an example has been shown in which the distance by which the sound receiving point is shifted from the character is adjusted in accordance with the camera direction, the moving speed, and the like. However, the sound processing device 100 may also adjust the direction by which the sound receiving point is shifted. For example, in the processing from steps S203 to S206 shown in FIG. 9 , the sound processing device 100 may reduce the localization difference by shifting not only the distance but also the direction of the sound receiving point.

[0123] 9, if the sound receiving point coordinates are updated to a position where the character cannot move, such as inside an object such as a wall, sounds that cannot occur in the virtual space may be synthesized, or signal processing may not be successful. Therefore, the sound processing device 100 can make predetermined adjustments.

[0124] This point will be described with reference to Figures 14 and 15. Figures 14 and 15 are diagrams for explaining a modified example of the correction process relating to the sound receiving point.

[0125] To perform an appropriate voice synthesis process at the updated sound receiving point, the sound processing device 100 acquires spatial data of the updated point and determines whether the coordinates are within the range of a character's movement. If the updated coordinates are beyond the range of a character's movement, the sound processing device 100 adjusts the length of the sound receiving point movement vector and re-updates the sound receiving point coordinates.

[0126] The example in Fig. 14 shows how the sound receiving point 200 is being corrected to the position of the sound receiving point 202. The sound receiving point 202 is a coordinate within a wall where the character cannot move. In this case, the sound processing device 100 adjusts the length and direction of the sound receiving point movement vector 204 as shown in Fig. 15 to shift the sound receiving point 202 to a position where it does not overlap the wall. For example, the sound processing device 100 can gradually adjust the length of the sound receiving point movement vector 204 until the coordinates are within the range where the character can move, thereby adjusting the sound receiving point 202 to be positioned at an appropriate position.

[0127] Furthermore, the sound processing device 100 may impose certain restrictions on the movement of the sound receiving point that accompanies a special movement method of the character.

[0128] For example, in some games, when a character jumps, a sudden change in height may occur. In this case, if the sound processing device 100 performs a normal correction process, the sound receiving point may move suddenly and unnaturally, which may cause a sense of discomfort to the user.

[0129] Therefore, when the virtual space is three-dimensional, the sound processing device 100 may adjust the movement amount adjustment coefficient so as to limit movement in the vertical direction. For example, when the content is such that a character is expected to perform a temporary action such as jumping, the sound processing device 100 may adjust the sound receiving point only in the horizontal direction, or may adjust the sound receiving point so as to reduce the degree of change in the vertical direction compared to the horizontal direction. In other words, when the virtual space is three-dimensional, the sound processing device 100 may make uniform corrections in the horizontal and vertical directions, or may adjust the sound receiving point so that it moves naturally by imposing different restrictions in certain directions.

[0130] 2 and 9, if the sound receiving point update process is included, the sound processing device 100 will perform sound path calculation twice. The path calculation is a process that requires a relatively high calculation cost in such a procedure. Therefore, in order to reduce calculation costs, when continuous character movement occurs, the sound processing device 100 may perform the path calculation in step S103 using the sound receiving point information calculated in the previous frame, and omit the path update in step S1033.

[0131] Furthermore, the sound processing device 100 may perform processing based on the human sound localization perception resolution regarding the restriction on the movement of the sound receiving point to reduce the localization difference shown in steps S203 to S206. That is, it is generally known that the human sound localization perception resolution differs depending on the direction. For example, the localization resolution of sounds from the front direction is fine, while the resolution from the sides and rear is coarser. Taking this characteristic into consideration, the sound processing device 100 may change the allowable angle for the localization difference depending on the angle from the front direction of the camera and the center of the camera to the direction of the sound source.

[0132] Furthermore, with regard to the movement amount limit of the sound receiving point, the sound processing device 100 may perform processing to change parameters used for adjusting the movement amount limit in a specific area in the virtual space based on the character coordinates. For example, with regard to the movement amount limit in the ellipse shape shown in Fig. 13, if the virtual space is an area on a long and narrow passage, the sound processing device 100 may adjust the movement amount by arranging the major axis of the ellipse along the length of the passage regardless of the movement direction.

[0133] 9, the sound processing device 100 may first perform the processes of steps S207 and S208, and then perform the processes of steps S203 to S206. In this case, the sound processing device 100 adjusts the movement amount depending on the camera direction in advance, and then calculates the localization difference using the adjustment result as a movement vector to the virtual sound receiving point 86 in step S206.

[0134] (2-6. Direct Effect Parameter Control) Next, the parameter emphasis process shown in step S1034 of Fig. 7 will be described. This process is executed to make it easier for the user to perceive changes in the surrounding space that accompany the movement of a character by emphasizing changes in the sounds heard while the character is moving. This process involves adjusting various parameters used in steps S104 to S108 shown in Fig. 2, which follow step S1034. The sound processing device 100 adjusts the parameters shown in Table 1 below, for example.

[0135]

[0136] Specifically, the sound processing device 100 adjusts the various parameters shown in Table 1 according to the procedure shown in Fig. 16. Fig. 16 is a flowchart showing the procedure of the parameter emphasis process. Note that the parameter emphasis process is basically executed separately for the parameters for each process in the subsequent steps S104 to S108. Furthermore, the parameter emphasis process may also be executed for parameters related to near-wall processing, which will be described later.

[0137] First, the sound processing device 100 reads out the spatial data and the route data acquired in steps S102 and S1033 (step S301).

[0138] Next, the sound processing device 100 determines whether or not the previously calculated result (processing result of the previous frame) is stored in the speech synthesis process (step S302). If the processing result of the previous frame is stored (step S302; Yes), the sound processing device 100 acquires the data (step S303).

[0139] In this case, the sound processing device 100 calculates the parameters at the current coordinates corresponding to each of the subsequent sound processes (e.g., steps S104 to S108) in order to synthesize the sound of the current frame to be processed (step S304).

[0140] For example, the sound processing device 100 adjusts parameters related to distance, such as attenuation and delay, as parameters common to the direct sound and other physical expressions excluding the near-wall expression, as shown in Table 1. Although these parameters can be calculated in subsequent processing using distance as an input, the sound processing device 100 here pre-calculates the characteristics of attenuation, delay, and frequency filters as an example of increasing the audible change as a result of the adjustment processing.

[0141] Regarding the representation of early reflections, if the sound absorption coefficient and filter coefficients are directly adjusted in subsequent processing, this may lead to an increase in sound pressure in the early reflection representation processing and the reproduction of sounds that sound unnatural to the ear, so the sound processing device 100 does not perform processing to adjust the sound absorption coefficient.The sound processing device 100 calculates only the angles between the sound source, the reflection point, and the sound receiving point, which are input to the reflecting surface simulation filter.

[0142] With regard to the diffraction representation, for the same reasons as for the early reflection sound representation, the sound processing device 100 calculates the distance from the sound source to the diffraction point, the distance from the diffraction point to the sound receiving point, the angle between the line segment connecting the sound source and the diffraction point and the surface on the sound source side as seen from the diffraction point, and the angle between the line segment connecting the diffraction point and the surface on the sound receiving point side as seen from the diffraction point.

[0143] With regard to the transmitted sound representation, the sound processing device 100 here performs only the above-mentioned processing related to distance.

[0144] Regarding the representation of late reverberation sounds, when a conventional feedback delay network (FDN) or the like is used, the parameters shown in Table 1 are further processed by a module for late reverberation representation, and converted into the delay length of each delay line, the number of delay lines, the filters included after the delay lines, and the input / output gains. For this reason, the sound processing device 100 performs calculations on the parameters listed in Table 1, which are input to the module. Conventionally, these parameters have been set for each space by a content creator or the like. Therefore, when using the conventional method as is, the sound processing device 100 reads these setting values ​​of the space through which the sound passes according to the propagation path. On the other hand, in the conventional method, near the boundary between spaces, the parameters of adjacent spaces are read, processed respectively, and then the processed results are further mixed and used at a mix ratio or the like according to the character coordinates and the distance between the spaces.

[0145] The sound processing device 100 can use such a method, but in an embodiment, adjustment may be made based on the distance between adjacent spaces. This point will be described with reference to Figures 17A and 17B. Figures 17A and 17B are diagrams for explaining parameter emphasis processing in a space.

[0146] 17A shows a virtual space 214 including a space 210 which is a large room and a space 212 which is a long and narrow corridor. A space 216 shown in FIG. 17A shows a predetermined range adjacent to the space 210 and the space 212.

[0147] When such adjacent spaces exist and a character moves between the spaces, the sound processing device 100 can change the parameters of each space (such as reverberation time) gradually in accordance with the distance from the boundary of the spaces, rather than abruptly. In the example shown in Fig. 17B, the parameters of the space from which the character moves (referred to as "Area A parameters" in Fig. 17B) change gradually from the parameters of the space from which the character moves (referred to as "Area B parameters" in Fig. 17B) in accordance with the position to which the character moves.

[0148] In this way, the sound processing device 100 can calculate the parameters of adjacent spaces, which change depending on the distance from the boundary between the adjacent spaces, using a function that smoothly connects them. The smoothly changing portion shown in FIG. 17B represents, for example, the state in FIG. 17A where a character is moving through the space 216. This allows the sound processing device 100 to transition each parameter more finely than the mix ratio that changes near the boundary in the prior art, thereby allowing the intention of the content creator to be more accurately reflected. Note that any method for processing late reverberation that can use similar parameters is not limited to the above-mentioned FDN, and any method can be used.

[0149] Regarding the near-wall representation described below, such processing also requires filter parameters, so the sound processing device 100 calculates the distance from the wall to the sound receiving point and the angle of the character relative to the wall surface to determine the filter parameters, as with the early reflection sound representation and diffracted sound representation.

[0150] After each of the above processes, the sound processing device 100 performs an adjustment process for each parameter (step S305).

[0151] Specifically, the sound processing device 100 updates the parameters for the current frame in step S303 so that the difference between the parameters for the previous frame and the parameters calculated in step S304 becomes larger, thereby emphasizing the change in sound accompanied by spatial expression that is output in subsequent processing.

[0152] As an example, the sound processing device 100 uses the time n of the current frame, the sound receiving point coordinates Pc(n) at time n, the parameter function f(Pc(n)) based on the coordinates shown in step S304, and the emphasis coefficient α to calculate the adjusted value A(n) of the parameter at time n using the following equation (1).

[0153]

[0154] Note that the above formula (1) is just one example, and the sound processing device 100 may use a different formula that includes an additional adjustment term in the above formula (1) so that the parameters do not diverge or become unstable when this process is performed continuously.

[0155] In the above formula (1), when the emphasis coefficient α is "1", the parameter is not emphasized, and when it is greater than "1" and less than "2", the parameter can be stably emphasized. Therefore, the sound processing device 100 can strengthen the emphasis of parameter changes by increasing the value within this range. Note that the amount of such emphasis can be arbitrarily adjusted by, for example, a content creator or a user via the emphasis coefficient.

[0156] On the other hand, when the emphasis coefficient α is greater than "0" and less than "1", the above formula (1) acts in the opposite direction, suppressing the change in the parameter.

[0157] When a parameter is emphasized, in the spatial representation of sound, changes in sound due to movement are more strongly perceived by the user than usual, thereby assisting the user's perception of spatial changes. On the other hand, behavior that suppresses changes can soften changes in sound due to movement, making it difficult for the user to notice spatial changes. In other words, the sound processing device 100 can adjust the degree of user perception by adjusting parameters. For example, by suppressing changes, the sound processing device 100 can provide game content with effects such as spatial transitions that are not noticeable to the user. The emphasis and suppression of these parameter changes can be selectively used by content creators, for example.

[0158] An example of parameter adjustment in a specific scene in a game is shown in Fig. 18. Fig. 18 is a diagram showing a specific example of parameter adjustment processing.

[0159] In the example of Figure 18, a scene is imagined in which, from a certain time to a certain time, a character moves in virtual space 214 shown in Figure 17A from a narrow passage (space 212) with a short reverberation time set to a wide space (space 210) with a long reverberation time set to a certain time, the character temporarily stops in the wide space, and then returns to the narrow space.

[0160] Even before parameter emphasis processing is performed, the parameters of adjacent spaces are adjusted to smoothly connect them, as described above, so the parameters (reverberation time in the example of Fig. 18) are set to smoothly connect the spaces. In other words, while the character is moving from a narrow passage to a wider passage, the reverberation time of the pre-adjustment parameters gradually increases in accordance with the above function.

[0161] If the above-described adjustment process is performed and the sound processing device 100 emphasizes the difference from the previous frame, the reverberation time will be longer during movement into a larger space than before the adjustment process. In other words, if the emphasis process is performed in such a situation, the sound processing device 100 will emphasize the parameters set for the larger space while the character is moving into the larger space (between frames 5 and 15 in FIG. 18 ), and calculate a longer reverberation time. As a result, while the character is moving from a narrow passage to a larger space, the user can perceive an expression that is closer to the sound of the larger space to which the character has moved.

[0162] Next, after moving from a narrow passage to a wide space, when the character temporarily stops (between frames 15 and 20), the coordinate-based parameter calculation always outputs the same value while the character is still, so the adjusted value gradually approaches the parameter obtained by the coordinate-based calculation. In other words, while the character is still, sound based on spatial information without any emphasis processing is output.

[0163] Next, when the character returns from the wide space to the narrow passage (between frames 20 and 30), the parameters obtained by calculation based on the coordinates change gradually to approach the parameters of the narrow passage, as in the previous movement. Furthermore, the sound processing device 100 outputs parameters adjusted by emphasis processing so that the reverberation time is closer to the parameters of the narrow passage.

[0164] In this way, when a character is moving between the previous frame and the current frame, the sound processing device 100 adjusts parameters for emphasizing the change in space, thereby making the change more noticeable to the user. Note that if the information on the previous frame is not stored (step S302 in FIG. 16 ; No), the sound processing device 100 performs calculation of the current frame as usual, and therefore does not perform emphasis processing.

[0165] The above describes an adjustment process that emphasizes the difference in parameters between the previous frame and the current frame. Here, near the boundary of a space, information about sounds that are heard from the destination and source spaces and whose localization should be taken into consideration is important for user recognition. Therefore, in addition to the above process, the sound processing device 100 may also perform a process that directly increases the gains (volumes) of direct sound, early reflection sound, diffracted sound, and transmitted sound from the values ​​calculated by the above process when the coordinates of a character are within a predetermined area set near the boundary of the space.

[0166] As a result, the sound processing device 100 can emphasize and present to the user sounds with a large amount of information compared to sounds such as late reverberation that do not give a sense of directionality or music that does not involve spatial expression. Note that the sound processing device 100 may adjust the amount of change in gain in proportion to the character's movement speed. As a result, the sound processing device 100 can make it less likely that the user will feel uncomfortable, even when the user moves between spaces while slowly checking the state of the adjacent space near the boundary.

[0167] After performing the above-described enhancement processing, the sound processing device 100 stores the processing results (step S307 in FIG. 16 ). Note that even if the sound processing device 100 performs gain enhancement processing on the direct sound, early reflected sound, diffracted sound, and transmitted sound near the boundary in step S305, the sound processing device 100 may not store such results so as not to affect the next frame.

[0168] (2-7. Wall Proximity Representation) Next, emphasis processing for representing a situation in which a character is located near a reflective material such as a wall will be described. This processing is executed, for example, between the voice synthesis processing (step S109) and the voice output (step S110) shown in Fig. 2. Fig. 19 is a diagram for explaining wall proximity representation.

[0169] 19, the character is located near a wall 264. As a real physical phenomenon, when a person is located near a wall, the sound that reaches both ears of that person changes depending on the distance from the wall.

[0170] For example, for low-frequency sounds with long wavelengths, both direct sound from the sound source and sound reflected from a wall reach the human ear, and the sound is emphasized depending on the distance from the wall. Therefore, as shown in the example of Figure 19, if right ear 260 is close to wall 264 and left ear 262 is far from wall 264, the low-frequency sound that reaches right ear 260 will sound louder than the sound that reaches left ear 262.

[0171] Furthermore, for high-frequency sounds with short wavelengths, the direct sound is partially blocked by the head, and the volume of the sound that reaches the ear due to diffraction is not large, so less of the sound reaches the right ear 260, which is closer to the wall. In other words, the high-frequency sounds that reach the right ear 260 sound quieter than those that reach the left ear 262.

[0172] Furthermore, when the character 250 rotates around the center of the head of the character 250, i.e., the sound receiving point, the distance to each ear and the degree to which the head blocks sound change. Figure 20 is a diagram for explaining the change in frequency in response to head rotation.

[0173] As shown in graph 270 in FIG. 20, when the character 250 rotates and the angle of the right ear 260 as seen from the wall changes from "0°" (the state shown in FIG. 19) to "90°" (the state where the character 250 is facing away from the wall 264), the low-frequency sound components decrease. On the other hand, the high-frequency sound components increase slightly. Also, as shown in graph 272, when the character 250 rotates and the angle of the left ear 262 as seen from the wall changes from "0°" (the state shown in FIG. 19) to "90°" (the state where the character 250 is facing away from the wall 264), the low-frequency sound components increase. On the other hand, the high-frequency sound components decrease slightly.

[0174] The sound processing device 100 can simulate these phenomena by, for example, filtering, thereby allowing the user to clearly recognize that the character 250 is located near a wall. For example, the sound processing device 100 performs filtering such that inputs are the distance from the wall to the sound receiving point, a preset head size or distance between the two ears, the angle of the character 250 relative to the wall, etc., and outputs converted frequencies. Specifically, the sound processing device 100 calculates coefficients for a filter that adjusts low-frequency sound pressure and a filter that adjusts high-frequency sound pressure for each of the left and right ears, and performs filtering for each of the left and right ears.

[0175] Through this processing, the sound processing device 100 can accurately reproduce in a virtual space sound phenomena that may occur near a real wall. This allows the sound processing device 100 to provide the user with a variety of virtual expressions, such as expressing a sense of narrowness through the reproduced sounds. In this processing, the sound processing device 100 may adjust only the overall gain uniformly across all bands, or may apply the same filter to both the left and right signals, in order to reduce the calculation load.

[0176] Furthermore, the physical phenomenon to be simulated by such processing is affected by the sound absorption coefficient of the wall, the sound absorption coefficient of the character's clothing, etc. Therefore, the sound processing device 100 may use this information to adjust the low and high frequency bands and sound pressure of the sound to be processed.

[0177] (3. Modifications) The above-described embodiment is merely an example, and various modifications and applications are possible. Modifications of the above-described embodiment will be described below.

[0178] (3-1. Use of Movement Information) In the above-described emphasis processing, the sound processing device 100 mainly uses information related to the movement of the character operated by the user, such as the movement speed, direction, etc. Here, the information related to the movement does not necessarily have to involve the movement of the character (change in coordinates).

[0179] For example, in games such as FPV, even if the character does not move, the user may direct their attention in a certain direction (referred to as "gazing") by directing their gaze (changing the camera direction), aiming at a specific target, etc. Alternatively, in games, even if the character itself does not move, the sound heard by the character may change if an object that serves as a sound source moves or if a character that is not the target of operation (NPC: Non-player character) moves.

[0180] As described above, even if the character itself does not move, i.e., even if the character coordinates do not change, the sound processing device 100 may acquire, as "information related to movement," a change in the direction that the character is paying attention to, or a change in surrounding objects, NPCs, etc. In other words, the "information related to movement" may include a change or setting of the gaze area, such as when the user focuses on a certain object or turns their line of sight or camera. Then, the sound processing device 100 may perform the above-described emphasis (correction) processing in accordance with such a change in the gaze area.

[0181] For example, the sound processing device 100 may apply the emphasis processing shown in Figures 16 to 18, etc., to the movement of a sound source approaching a character, rather than the movement of the character itself. Alternatively, with regard to updating the sound receiving point as shown in Figure 8, the sound processing device 100 may perform processing such as shifting the sound receiving point from the original character coordinates toward the sound source when the character does not move but the sound source is approaching. Furthermore, the sound processing device 100 may perform processing such as shifting the coordinates of a sound source approaching the character toward the character prior to the character's movement. In this case, the sound processing device 100 may control the amount of preceding movement of the sound source in consideration of the localization difference, as in the processing shown in Figure 9. That is, the sound processing device 100 may execute the processing shown in Figure 9 by replacing the coordinates of the sound receiving point with sound source coordinates and the sound source coordinates with sound receiving point coordinates.

[0182] (3-2. Accessibility) The processing according to the embodiment can be applied not only to the production of content using a virtual space but also to various purposes or aspects. For example, the sound processing device 100 may apply the processing according to the embodiment to a person with a disability whose detection ability, such as hearing, is inferior to that of a healthy person.

[0183] As described above, the sound processing according to the embodiment can emphasize changes in sound that accompany spatial changes caused by character movement. This makes it easier for able-bodied people to sense spatial changes, such as spatial transitions, than before. Furthermore, people with disabilities can also hear sounds that emphasize environmental changes more than before, so they can sense spatial changes more prominently and detect the changes more easily.

[0184] (3-3. Use of Artificial Intelligence, Machine Learning, Cloud, etc.) As described above, the emphasis processing according to the embodiment can control the effect level using parameters such as a movement amount adjustment coefficient and an emphasis coefficient. These parameters can be freely set by a content creator or a user within a predetermined value range, but the sound processing device 100 may provide appropriate parameters prepared in advance by the system for a content creator or a user who is using this processing for the first time. For example, the sound processing device 100 may set a plurality of parameters in advance and provide the parameters in a manner that allows the content creator or the user to select from among them.

[0185] Furthermore, the sound processing device 100 may acquire, via the cloud, the parameters set by a content creator or a user in a specific scene and the circumstances in which they are applied (spatial data, whether the scene is an exploration scene or a battle scene, the number of sound sources, the type of sound source, the characteristics of the playable character, etc.). The sound processing device 100 then performs machine learning on the parameters and the applicable scenes. As a result, the next time a scene similar to these appears, the sound processing device 100 can output appropriate parameters through machine learning and automatically set the output parameters.

[0186] For such machine learning, the sound processing device 100 may perform learning only through the work of a single content creator during production, or learning data may be collected on the cloud or the like to create a learning environment based on information from multiple content creators. For example, when the sound processing device 100 learns only through information from a single content creator, parameter information, which can be considered production know-how, can be kept confidential. On the other hand, when the sound processing device 100 learns based on information from multiple content creators, it can learn from more scenes and more easily adapt to various scenes. Even when learning is targeted at users rather than content creators, the sound processing device 100 can similarly utilize the respective advantages of each. Furthermore, the sound processing device 100 may upload learning data, which compiles parameters and scene information created by content creators or users, and internal data of machine learning learned based on the learning data, to the cloud, so that other creators or users can use it.

[0187] (3-4. Parameters) In the above embodiment, an example was shown in which the sound processing device 100 adjusted the gain, reverberation time, etc. However, the sound processing device 100 may also perform emphasis processing on the parameters listed in Table 1 and various other parameters set in each content.

[0188] For example, in the processing shown in Figure 16, the sound processing device 100 may emphasize the delay of sound or adjust a specific frequency band (equalizer processing) depending on the speed and direction of the character's movement between the previous frame and the current frame.

[0189] (3-5. Acquisition of Object Data from Library) In the embodiment, the sound processing device 100 acquires, as data for calculation, spatial data such as data of objects with settings such as reflection. For example, the sound processing device 100 acquires, as data for calculation, shape data of the object, texture data, and data on the material and sound absorption coefficient assigned to the object. Here, if no settings or designations have been made regarding data of an object in the virtual space to be processed, the sound processing device 100 may be configured to allow a creator to read and use any object data stored in a library. For example, the sound processing device 100 may automatically select object data from the library.

[0190] Here, the object data stored in the library may be linked not only with shape data and material information, but also with parameters related to reflection (for example, frequency characteristics of reflected sound, waveforms such as the behavior of scattered waves, and spectra).

[0191] The sound processing device 100 may be configured so that a creator can edit data (e.g., object data or waveform data) stored in the library. For example, the sound processing device 100 may be provided with a UI (User Interface) that allows a creator to arbitrarily change object data (e.g., material parameters of an object). This allows a creator to change the frequency characteristics of reflected sound to desired characteristics. The sound processing device 100 may also be provided with a UI that allows a creator to parametrically change the size of the object's unevenness, the scattering repetition frequency, etc., in order to adjust the scattering behavior.

[0192] (3-6. Voice Synthesis Processing) With regard to the synthesis processing of direct sound etc. shown in FIG. 2, the sound processing device 100 can arbitrarily use various processes described in patent applications already disclosed by the same applicant as the present disclosure.

[0193] (3-7. Application Examples of the Technology of the Present Disclosure) This disclosure has described a sound-receiving point correction process that emphasizes the sense of transition of characters in order to enhance the sense of immersion in a virtual space. This technology may be applied not only to audio processing but also to other processing.

[0194] For example, the technology of the present disclosure may be applied to video processing. When applied to video processing, the sound receiving point coordinates can be interpreted as camera coordinates for video presentation. That is, the sound processing device 100 may shift the camera coordinates, which normally coincide with the viewpoint of a character in an FPV or the like, so that the camera coordinates are slightly ahead of the character's movement in accordance with the character's movement.

[0195] This process allows the user to quickly obtain information about the direction of movement not only through audio information but also through visual information, which makes it easier for the user to take the next action in the game based on the obtained information and to feel the dynamism that accompanies movement, thereby increasing the sense of immersion in the virtual content.

[0196] (3-8. Other Examples of Acoustic Expression) (3-8-1. Overview) In the above embodiment, a method for expressing changes in the sound environment that occur as a character moves in a virtual space has been described. However, the technology according to the present disclosure is not limited to the above embodiment. As another example of acoustic expression, it is also possible to generate various sound effects by automatically setting the position of a sound receiving point or operating it on the system.

[0197] For example, the technology disclosed herein not only realizes effective audio expression for users, but can also be used for audio adjustment during game content creation. Specifically, the technology disclosed herein is incorporated into the game engine (core software that executes the main processes commonly used in game content) used by creators when creating game content. This allows creators to effectively and efficiently create images and audio in virtual spaces. This is because, when creating audio for content in virtual spaces such as games, it may be possible to provide users with a more immersive experience by not only faithfully reproducing sound according to the physical laws of reality, but also reproducing sound by taking other factors into account.

[0198] Taking a game as an example, sound adjustment during game content creation can be divided into (1) the expression process of sound-generating bodies (referred to as "emitters" or the like) and sound space, and (2) the adjustment and creation of the positions of sound-receiving points (referred to as "listeners" or the like). For example, creators can freely design sound-receiving points and sound-emitting points in a game engine by directly writing programs, utilizing the game engine's functions, or using sound software (referred to as "sound middleware" or the like). Specifically, creators set each of the game components (referred to as "objects" or the like), such as a virtual camera that determines the field of view displayed on the display during gameplay, a character controlled by the user, and an NPC (Non-Player Character), as a sound-receiving point and a sound-emitting point. This allows creators to create a sound space within the content and thereby design the sound that users will experience.

[0199] However, such sound design requires creators to individually design the sound receiving and emitting points for each individual game or project, which can be inefficient. Furthermore, even within the same content, there are scattered scenes where the sound is not properly designed, resulting in inconsistent quality. Similarly, in VR apps that use user-generated content (UGC) or when users themselves create spaces or sound-generating objects in games, the sound design may be insufficient compared to professional creators. Furthermore, in online multiplayer games, unexpected sounds may be generated within the content when users interact with each other in ways that the creators did not anticipate.

[0200] That is, considering the limitations of the game engine's supported functions and the amount of effort that creators and users can expend, it is difficult to create detailed sound throughout the content.

[0201] Therefore, the sound processing described below reduces the sound design burden on creators and users and realizes effective sound production. Specifically, the following describes an example of sound adjustment processing, such as how the user experiences sound (corresponding to (2)) by adjusting the position of the sound receiving point for sounds created using various techniques as described in the embodiments (corresponding to (1) of "Sound adjustment during game content creation" above).

[0202] Such an example will be described with reference to Fig. 21 and subsequent figures. Fig. 21 is a diagram showing an outline of the acoustic expression according to the modified example.

[0203] 21(a) shows a situation in which an NPC (Non-Player Character) 302, a virtual character in the game, is speaking to a character 300 operated by a user. This situation can be seen by the user operating the character 300 as a scene in a virtual space photographed by a virtual camera 310. Note that the virtual camera 310, which determines the field of view displayed on the display during game play, the character 300, the NPC 302, etc. are components in the game (referred to as "objects", etc.), and are defined in advance in the game engine, for example.

[0204] 21A, the sound receiving point coincides with the character 300. The sound producing body is the NPC 302. In this case, the user perceives the sound by hearing the sound emitted by the NPC 302 at the position of the character 300.

[0205] 21B, the sound processing device 100 sets a new sound receiving point by sound processing according to the modified example (step S400). Specifically, the sound processing device 100 sets a parent listener 320, which is an example of a sound receiving point, on a line connecting the character 300 and the virtual camera 310. The sound processing device 100 also sets a child listener 330, which is an example of a sound receiving point, on a line connecting the character 300 and the NPC 302 and in the vicinity of the NPC 302.

[0206] The parent listener 320 is a coordinate for determining the sound to be output from the content. The sound processing device 100 calculates what kind of audio signal is to be synthesized at that coordinate in the virtual space and outputs the calculation result as an audio signal. The child listener 330 is a coordinate that serves as a sound receiving point that assists the parent listener 320. If the child listener 330 is not generated, the sound processing device 100 outputs only the sound observed by the parent listener 320. Hereinafter, when there is no need to distinguish between the parent listener 320 and the child listener 330, they will simply be referred to as "listeners." Note that a control method for automatic listener generation will be described in detail later.

[0207] That is, the user will hear the sound emitted by NPC 302 as a mixture of the volumes observed at the positions of parent listener 320 and child listener 330. For example, the user will hear a sound that is the result of adding together the audio signal of parent listener 320, which is the normal sound receiving point, and an audio signal obtained by multiplying the sound observed at child listener 330 by a predetermined coefficient. Therefore, the user can experience audio that sounds more like NPC 302 is whispering closer to character 300, or audio that sounds more like NPC 302 is speaking to character 300 from a slightly overhead position, than if the user were to hear it from the position of character 300.

[0208] In this way, the sound processing device 100 not only matches the sound-receiving point with the character 300, but also adjusts the impression of how the sound sounds by setting a plurality of sound-receiving points in the virtual space, such as the parent listener 320 and the child listener 330. This enables the sound processing device 100 to produce a wider variety of sound expressions and realize sound expressions that attract the user's attention.

[0209] (3-8-2. Flow of Acoustic Processing) Next, a specific configuration for realizing the acoustic processing shown in Fig. 21(b) and the flow of the acoustic processing are shown in Fig. 22. Fig. 22 is a diagram showing the flow of acoustic processing according to a modified example.

[0210] 22, the sound processing device 100 according to the modified example has processing target data 400. The processing target data 400 refers to various data related to a virtual space that is the target of sound processing according to the modified example, for example, listener position control processing.

[0211] The sound processing device 100 acquires, as processing target data 400, for example, existing listener information 402, sound generating body (emitter) information 404, camera information 406, extended attributes 408 including object attributes and audio meta tag information, and avatar information 410.

[0212] The existing listener information 402 includes information such as the position and volume setting that have already been set as sound receiving points before the child listeners, etc. are generated. The sound emitter information 404 includes the position of objects that can emit sound, and the volume and tone of the sound that is emitted. The camera information 406 is setting information for a virtual camera that determines the viewing angle of the virtual space, and includes information such as whether it is FPV (First Person View), TPV (Third Person View), whether it follows a character, or whether it is a fixed camera. Note that if no explicit information is set for a camera object used for sound processing, meta information about the virtual camera may be included in information about other objects, etc.

[0213] The extended attributes 408, including object attributes and audio meta tag information, include attribute information for each object set by a game engine or the like. The attribute information includes, for example, the type and attributes set for the object (user-controlled player or NPC, whether it has an emitter, whether it is an enemy or ally, whether it has hit detection, whether it has transparency, etc.), as well as various settings such as the size and weight of each object. The audio meta tag information also includes various settings such as the role of the audio function of the object (listener or emitter) and the attributes of the emitter (sound effect or audio, etc.). The avatar information 410 includes setting information for a character (avatar) in a virtual space and various information associated with the user operating the character.

[0214] In the acoustic processing according to the modified example, the sound processing device 100 first obtains information about sound generating bodies in a space to grasp the sound pressure distribution of the sounds emitted from the sound generating bodies (step S410). The sound processing device 100 then refers to listener position control method data 420, which defines a method for automatically generating and adjusting the listener position, and selects a listener position control method (step S412).

[0215] In this process, the sound processing device 100 may acquire sound production concept setting information according to the selection of the creator or user (step S414). The sound production concept setting information is preset information for sound processing prepared in the system, and is a setting that succinctly indicates the sound expression that the creator wants to set for the scene, such as a sense of realism, fear, or tension. The sound processing device 100 selects a control method application target in the space data based on the listener position control method data 420 and the sound production concept setting information (step S416).

[0216] Next, the sound processing device 100 generates child listeners that are sound-receiving points different from the parent listener in accordance with the listener position control method data 420 and the audio production concept setting information (step S418). A specific example of the child listener generation process will be described later.

[0217] Furthermore, the sound processing device 100 sets the positions at which the parent listener and the child listeners are to be placed based on the current distribution of sound producing bodies, the distribution of objects, etc. (Step S420). The sound processing device 100 generally sets the parent listener at a position superimposed on the character operated by the user, but depending on the production concept, the parent listener can also be set at a different position.

[0218] A specific example of the position where the sound processing device 100 sets the listener will be shown with reference to Fig. 23. Fig. 23 is a diagram showing an example of setting the listener.

[0219] 23A shows an example in which the parent listener 320 is set at approximately the same position as the character 300 within the viewing angle of the virtual camera 310. This example is a common acoustic process in game content, etc., in which the position of the character 300 controlled by the user coincides with the sound receiving point. In this case, the user can hear the sounds experienced by the character 300, such as the user's own footsteps and the sound of clothes rustling, as well as other subtle sounds in the surrounding area.

[0220] 23(b) shows an example in which the parent listener 320 is positioned on a line connecting the virtual camera 310 and the character 300. In this example, the user hears sounds observed slightly closer to the virtual camera 310 than the position of the character 300. In this case, compared to the example in FIG. 23(a), the user can perceive ambient sounds being emitted off-screen.

[0221] FIG. 23( c ) shows an example in which the parent listener 320 is placed near the virtual camera 310 on a line connecting the virtual camera 310 and the character 300. In this example, the user hears sounds observed near the virtual camera 310. In this case, compared to the examples of FIGS. 23( a ) and 23 ( b ), the user can hear sounds that match the appearance of the character 300 viewed from above (i.e., the viewing angle of the virtual camera 310). This allows creators to create effects that take advantage of the position of the sound (for example, when wanting the user to perceive the sound emitted by the character 300's enemy more strongly). Note that, depending on the genre of the game, the sound processing device 100 may be flexibly configured, such as by placing the parent listener 320 at a position slightly offset from the line connecting the character 300 and the virtual camera 310.

[0222] Returning to Fig. 22 , the explanation continues. Next, the sound processing device 100 sets a mix coefficient for determining in what manner the sound observed in the child listener will be output (step S422). For example, when the sound observed in the parent listener is output at a volume of "1", the sound processing device 100 sets a coefficient for determining at what volume ratio the sound observed in the child listener will be output. Note that the mix coefficient may not only be a numerical value for determining the volume ratio, but also a setting value for determining the tone or frequency to be output. The sound processing device 100 may automatically determine the mix coefficient, or may acquire a mix coefficient setting from a creator or user (step S424).

[0223] The mix coefficients of the child listeners set by the creator may be determined based on the effects of the game, etc. This point will be described with reference to Fig. 24. Fig. 24 is a diagram showing an example of setting child listeners.

[0224] First, the sound processing device 100 determines whether the set sound is meaningful (important) in terms of the game, such as whether the scene contains a sound that the user wants to concentrate on listening to (step S450). If it is determined that the set sound is meaningful in terms of the game (step S450; Yes), the sound processing device 100 adds a predetermined effect to the sound of the child listener based on information input by the creator (step S452). For example, the sound processing device 100 changes the distance at which the sound of the child listener can be heard, for example, by setting a higher mix coefficient so that the sound of the child listener is heard more strongly. As an example, when a trigger occurs, such as an NPC nearby a character trying to talk to the character, the sound processing device 100 sets information about the child listener such as a coefficient that will result in an effect that emphasizes the voice of the NPC and the position of the child listener.

[0225] On the other hand, if it is determined that the setting is meaningless from a game perspective (step S450; No), the sound processing device 100 processes the sounds of the child listeners using a physics-based common set, which is sound processing that imitates real physical phenomena (step S454). Through this processing, the sound processing device 100 can set information about the child listeners in accordance with the performance that the creator wants to achieve.

[0226] 22 , the explanation will be continued. Once the mix coefficients for the child listeners have been set, the sound processing device 100 performs post-processing such as a compressor on the sounds observed by the child listeners to adjust the sound pressure and the like (step S426), and then performs mixing processing with the parent listener (step S428).

[0227] Finally, the sound processing device 100 outputs an audio signal in which the sounds of the parent listener and the child listener are mixed (step 430). Through this processing, the sound processing device 100 can generate child listeners and realize sound expression processing using the child listeners.

[0228] In the above process, the sound processing device 100 may provide a creator or a user with a GUI (Graphical User Interface) and accept input information on the GUI. For example, input may be made using sensor values ​​acquired by the creator or the user or various input devices (such as the rotation or angle of a controller or a dial provided on a device, or information detected by a pressure sensor).

[0229] The sound processing device 100 may acquire some or all of information about objects and emitters in the space to be processed, and audio meta tag information set therefor, and control the listener position. Furthermore, when generating the listener position, the sound processing device 100 may calculate in advance an audio signal to be observed at a certain position, and use this information in the generation process.

[0230] (3-8-3. Other Configurations of Sound Processing) The above-described modified example may be implemented by having the sound processing device 100 (for example, a device that executes content) use information set in a game engine. Such an example will be described with reference to Fig. 25. Fig. 25 is a diagram showing an example of the configuration of sound processing according to the modified example.

[0231] 25, the game engine 500 has setting information commonly used for various game contents, and the sound processing device 100 uses this setting information. For example, the game engine 500 has sound production concept setting information 502 that indicates the production concept for a scene for which a creator or the like sets the sound. The sound production concept setting information 502 corresponds to the same items shown in FIG. 22, for example.

[0232] The game engine 500 also has child listener coordinate information 504 that indicates the coordinates of child listeners generated in the content. The game engine 500 also has emitter coordinates / emitter extended attributes 506 that indicate extended information including the coordinates of sound-producing bodies in virtual space and audio meta tag information. The game engine 500 also has parent listener coordinate information 508 that indicates the coordinates of parent listeners set in the content.

[0233] The sound processing unit 600 (corresponding to, for example, the control unit 130 shown in FIG. 6B ) included in the sound processing device 100 acquires various information from the game engine 500 and performs sound output processing. Specifically, as sound processing for the child listener, the sound processing unit 600 calculates the relative position from the coordinates of the child listener and the emitter, and performs sound processing for the sound source at the child listener coordinates (step S460). For example, the sound processing unit 600 calculates the sound pressure output by the child listener based on the volume and relative position of the sound output from the emitter, applies various filters, and performs convolution processing for synthesis.

[0234] Similarly, the sound processing unit 600 calculates the relative position from the coordinates of the parent listener and the emitter as sound processing for the parent listener, and executes sound processing for the sound source at the parent listener coordinates (step S462).

[0235] Thereafter, the sound processing unit 600 performs spatialization processing (rendering) of the sound observed at the child listener coordinates (step S464). Similarly, the sound processing unit 600 performs spatialization processing (rendering) of the sound observed at the parent listener coordinates (step S466). Spatialization processing is processing for converting or adjusting audio to a mode suitable for the playback environment. For example, when multiple speakers (output devices) are assumed, the spatialization processing may include decoding into a multi-directional speaker layout. The spatialization processing may also include decoding into a binaural layout (in this case, user-specific or general-purpose HRTF may be used during decoding). The spatialization processing may also include gain panning processing such as VBAP (Vector Based Amplitude Panning) for the multi-directional speaker layout.

[0236] Next, the sound processing unit 600 acquires the sound production concept setting information 502 and mixes the signals observed by the child listener and the parent listener, taking into account the type of production to be added to the sound that will ultimately be output. For example, the sound processing unit 600 applies a filter (conversion of sound pressure, frequency, relative distance of sound, phase, etc.) preset for each production concept to generate an audio signal that meets the creator's intentions. Furthermore, the sound processing unit 600 performs post-processing such as a compressor to generate a final output audio signal (step S468).

[0237] The sound processing unit 600 outputs the generated audio signal to an audio output device such as a speaker 650. The sound processing unit 600 may also binauralize the audio signal (step S470) and output the audio signal to headphones 660. The sound processing unit 600 can binauralize audio using, for example, an HRTF group in a direction to be decoded in Ambisonics or an HRTF group in a direction to be gain rendered in VBAP.

[0238] 25 , by storing setting information in advance in the game engine 500, creators and users do not need to set detailed information for each piece of content, but can use common setting information for all pieces of content that use the same game engine 500. Note that the configuration shown in Fig. 25 is an example, and some or all of the processing may be executed on the game engine 500 side, or some or all of the processing may be executed on the sound processing device 100 side.

[0239] 25 is an example, and the processing procedures may be arbitrarily changed in the sound processing device 100. For example, the sound processing device 100 may perform mixing processing on the parent listener and the child listener or may perform compressor processing before the spatialization processing.

[0240] (3-8-4. Listener Generation) As described above, the sound processing device 100 generates a listener at a predetermined position, thereby realizing sound that is in line with the creator's production concept. In order to reduce the burden on creators and users, it is desirable to automatically generate a listener at an appropriate position. In this regard, the listener generation process will be described with reference to Figs. 26 and 27. Fig. 26 is a diagram (1) showing an example of the listener generation process.

[0241] FIG. 26( a ) shows an example in which, when there is an NPC 302 that speaks to the character 300 , the parent listener 320 is generated on a line connecting the character 300 and the NPC 302 .

[0242] In this example, the sound processing device 100 first references the player tag in the audio metatag information assigned to the character 300, and since a player tag has been assigned, sets the parent listener 320 near the character 300. Then, based on the production concept set by the creator, the sound processing device 100 adjusts the parent listener 320 to either match the position of the character 300 or move it closer to the NPC 302. In this example, the sound processing device 100 sets the parent listener 320 on a straight line connecting the character 300 and the NPC 302, near the midpoint between the character 300 and the NPC 302. This allows the user to hear sounds observed from a position closer to the NPC 302 than sounds heard at a position aligned with the character 300, thereby providing a stronger sense of realism.

[0243] Figure 26(b) shows an example in which, in a scene similar to that shown in Figure 26(a), a parent listener 320 is set between the character 300 and the virtual camera 310, and a child listener 330 is generated on a straight line connecting the character 300 and the NPC 302.

[0244] In this example, the sound processing device 100 first references the player tag in the audio metatag information assigned to the character 300, and since a player tag has been assigned, sets the parent listener 320 near the character 300. Furthermore, the sound processing device 100 detects the speech of the NPC 302 and sets the child listener 330 near the NPC 302. Then, based on the production concept set by the creator, the sound processing device 100 sets the parent listener 320 on a straight line connecting the virtual camera 310 and the character 300. Furthermore, the sound processing device 100 sets the child listener 330 near the NPC 302. This allows the user to hear sounds from the entire virtual space on the parent listener 320, and at the same time, to clearly hear the voice of the NPC 302 whispering to the character 300.

[0245] Fig. 27 is a diagram (2) illustrating an example of the listener generation process. Fig. 27 illustrates an example in which the sound processing device 100 sets sounds to be heard by a character 300 watching a virtual live performance held in a virtual space.

[0246] In the example of FIG. 27, the virtual live performance is taking place on the first floor 700, and the character 300 is located on the second floor 702.

[0247] In this example, the sound processing device 100 first determines the position of the virtual camera 310 in the space and the positions of each character (avatar) to be placed. Next, the sound processing device 100 acquires sound production concept information set for the virtual live performance, and places the parent listener 320 and child listeners 330 in positions that conform to the concept.

[0248] For example, in order to enable the user to hear the sounds of the entire live venue, the sound processing device 100 sets the parent listener 320 on a line connecting the virtual camera 310 and the character 300. In addition, in order to create a sense of realism in the performance, the sound processing device 100 sets the child listener 330 in front of the performance stage on the first floor 700.

[0249] Furthermore, the sound processing device 100 acquires emitter information (in this example, the performer), and sets information based on the emitter coordinates to determine the balance of listening between the parent listener 320 and the child listener 330. For example, based on the emitter information, the sound processing device 100 can prioritize listening to the sounds of conversations near the character 300 with the parent listener 320, or prioritize listening to the sounds of a performance with the child listener 330.

[0250] Thereafter, the sound processing device 100 spatializes the audio signal based on the emitter information and the listener information, and further acquires mix coefficients for the parent listener 320 and the child listener 330 based on the concept information to perform mix processing. The sound processing device 100 then outputs audio with the final determined balance.

[0251] In this way, the sound processing device 100 can easily realize effective sounds that are in line with the creator's concept by automatically and appropriately generating the listener's position.

[0252] It should be noted that listener generation is not limited to the above and may be realized by various methods. As a first example, the sound processing device 100 may generate child listeners based on concept information set for all assets (a collective term for objects and setting information that make up a game stage or scene) of a predetermined stage in a game.

[0253] First, the sound processing device 100 reads concept information setting for the initial position of the player character for the current asset of a certain stage. Note that the concept information is not limited to being set for the entire asset, but may also be set for a spatial region specified within the asset or a camera object within the asset. Note that if concept information overlaps, the sound processing device 100 may automatically select the concept information to be applied according to a separately specified priority.

[0254] Next, the sound processing device 100 reads setting information related to a listener position control processing method that corresponds to the concept information. The following describes a case where the setting information for the listener position control processing method assumes the use of a child listener. In this example, the sound processing device 100 has acquired concept information that "emphasizes ambient sounds when a specific operation is performed," and the following specific control information is set: (1) N child listeners are generated and placed on concentric circles centered on the player (the radius of the circle is, for example, a constant multiple of the distance between the player and the TPV camera) so that the angles between each of the child listeners are equal. At this time, the initial mix gain value of the child listeners is set to "0." (2) When a player performs an operation that satisfies the condition for activating "emphasize ambient sounds," (a) the parent listener disables reception of sounds arriving from outside the concentric circles centered on the player. (b) The mix gain value of the child listeners is set to "1." (c) When the player's operation no longer satisfies the conditions for activating "emphasize ambient sounds," the child listener's Mix gain value is set to "0," and the parent listener's sound receiving conditions are restored to their original state.

[0255] Thereafter, the sound processing device 100 multiplies the sound received by the child listener by a Mix gain and adds the result to the parent listener. When the asset is switched, the sound processing device 100 deletes the child listener.

[0256] Through the above processing, the sound processing device 100 can realize sound expression that makes the player feel as if he or she is listening to surrounding sounds by straining his or her ears, without relying on real physical laws.

[0257] Next, as a second example, a process in which the sound processing device 100 generates a child listener when a player moves between areas within an asset and newly enters a specific spatial region for which concept information is set will be described.

[0258] As in the first example, when the player is included in a spatial region defined in the current asset, the sound processing device 100 reads concept information set for that spatial region. Then, the sound processing device 100 reads setting information for a listener position control processing method corresponding to the concept information setting.

[0259] The following describes a case where the configuration information for the listener position control processing method assumes the use of child listeners. In this example, the concept information is "emphasize fear," and the following specific control information is set: (1) N child listeners are generated in advance. The initial value of the child listener's mix gain is set to "0." (2) Among objects tagged with "Foley / sound effects" in the audio meta tag information, the child listeners are automatically moved to the vicinity of the N objects that are closest in linear distance to the player among those not within the camera's field of view, and the mix gain is set to "1.5." This process allows the user to hear the sounds of invisible objects louder. (3) When the player moves and the distance to the object changes, the child listeners located near the object farthest from the player and not making any sound are moved to the vicinity of the new object that the player has approached that is not making any sound.

[0260] Specifically, the sound processing device 100 generates N child listeners in accordance with control rule (1). Next, the sound processing device 100 acquires information about objects within a specific spatial region that may emit sound (determined based on whether or not they have an emitter attribute, or on acoustic meta tag information, etc.).

[0261] The sound processing device 100 refers to the acquired acoustic meta tag information of the object, and for the object tagged with "Foley / sound effect", moves the child listener and sets the mix gain in accordance with control rule (2). Note that if a child listener has already been generated, the sound processing device 100 may delete it and regenerate it, or may reuse the child listener.

[0262] Thereafter, when the player moves and the distance to the object changes, the sound processing device 100 rearranges the child listeners in accordance with control rule (3). Then, the sound processing device 100 multiplies the sound received by each child listener by a Mix gain, adds the result to the parent listener, and outputs an audio signal. Note that the sound processing device 100 may delete the child listeners when the player moves out of the specific spatial region.

[0263] Next, as a third example, we will explain a process of generating a child listener in accordance with rules when concept information is set for a camera object in an asset and the currently active field of view is switched to that camera.

[0264] As an initial setting in the third example, if there is a child listener that has already been generated, the sound processing device 100 determines whether to delete it and regenerate it, or to reuse it.

[0265] Next, the sound processing device 100 reads the concept information setting assigned to the current camera. In this example, the concept information "emphasizes realism" and the following specific control information is set: (1) A child listener is generated and placed at a position of "1:9" on the line connecting the object tagged "Foley / sound effects" and the player, and the mix gain is set so that the sound is heard 6 dB louder than the parent listener. Specifically, the mix gain is set so that the volume is 6 dB higher than the volume at the parent listener position as perceived by the sound processing device 100. (2) A child listener is generated and placed at the emitter position of the "enemy object (speech)," and reverb processing is applied to the child listener, with the mix gain set so that the sound is heard 10 dB quieter than the volume heard at the parent listener. This assumes the effect of adding a reverb plug-in to the audio track obtained from each listener in addition to the characteristics originally set for the emitter.

[0266] Specifically, the sound processing device 100 acquires information about objects that have the potential to produce sound within the asset. If the asset is large, the sound processing device 100 may check the objects by limiting the objects to those within the field of view of the current camera or those close to the parent listener.

[0267] Next, the sound processing device 100 refers to the acoustic meta tag information of the acquired object, generates a child listener for the object tagged with "Foley / sound effect" in accordance with control rule (1), and sets the mix gain.

[0268] Furthermore, the sound processing device 100 generates a child listener and sets a Mix gain for an object tagged with "enemy object (speaking)" in the object's acoustic meta tag information in accordance with control rule (2).

[0269] The sound processing device 100 then multiplies the sounds received by each child listener by a Mix gain, adds the result to the parent listener, and outputs an audio signal. Note that the sound processing device 100 may delete the child listener when the camera is switched.

[0270] As described above, the sound processing device 100 can flexibly generate child listeners and adjust the mix gain (i.e., the volume of each child listener) in accordance with a set concept, allowing creators to achieve sound expression in accordance with a concept without having to manually perform detailed settings.

[0271] (3-8-5. Application Examples) The processing according to the above-described modified example can be used for various content. For example, the sound processing device 100 can apply the processing according to the above-described modified example to a so-called stealth game in which the player faces an enemy while hiding in an FPV game.

[0272] Specifically, when a player uses a scope during a game, the sound processing device 100 generates a child listener near the visual target, thereby enabling not only the target but also its sound to be amplified, creating a sense of realism. Furthermore, the sound processing device 100 can realize expressions and effects such as selectively amplifying only distant sounds by imagining an acoustic scope (a device that can amplify only sound), which is difficult to reproduce in reality.

[0273] That is, if the processing according to the modified example is not used, sounds that are far from the player require the effort of separately recording and playing back, for example, as dialogue, but by generating a child listener, the sound processing device 100 can output distant sounds without the effort of separately recording, etc. Specifically, the sound processing device 100 can output sounds that are far from the player without creating a sense of incongruity by arranging objects within the scope angle of view and child listeners in association with each other.

[0274] Furthermore, the sound processing device 100 can output natural sound even when the zoom level changes by changing the mix coefficient of the child listener along with changing the zoom magnification of the scope. Furthermore, the sound processing device 100 can realize a function such as listening to game chat of the opposing team by, for example, having a device for generating a child listener appear in the game content.

[0275] In addition, in content such as searching for a specific person in a crowd, the sound processing device 100 may narrow down the target child listeners to which output is to be made or apply compressor processing so that the player can concentrate on listening to only the sound of the character that is attracting the player's attention.

[0276] (3-8-6. Variations of Processing) In the processing according to the above-described modified example, the sound processing device 100 may have various variations.

[0277] For example, even when a child listener is generated, the sound processing device 100 may be configured to not allow the user to hear the sound observed by the child listener. Alternatively, the sound processing device 100 may mute the sound of the parent listener and allow the user to hear only the sound of the child listener. In this way, when sounds observed by the parent listener and the child listener overlap, the sound processing device 100 can prevent the user from hearing the overlapping sound.

[0278] The mute setting (such as a solo mode where only a single sound is heard) may be set by the creator or user, or may be set automatically based on the degree of overlap of sounds (volume, etc.) observed within the content.

[0279] Furthermore, the sound processing device 100 (or the game engine) may set priorities in the extended attributes set for objects that are sound generating units. The sound processing device 100 may adjust the placement of child listeners or change the mix ratio so that the user preferentially hears sounds output from sound generating units with higher priorities. For example, the sound processing device 100 may adjust the position of the listener so that the user can hear the sound from a frontal direction of a sound generating unit with higher priorities.

[0280] Furthermore, the sound processing device 100 may select sounds to be output according to priority in accordance with the processing capacity of the playback environment. Specifically, in an environment with sufficient processing capacity, the sound processing device 100 processes sounds output from a large number of sound generators. On the other hand, in an environment with insufficient processing capacity, the sound processing device 100 may process only sounds from sound generators with high priority.

[0281] The sound processing device 100 may also accept user-defined extended attributes, such as audio meta tag information, set for each object. This allows the sound processing device 100 to arrange child listeners and adjust the volume of the child listeners as desired by the user. For example, the sound processing device 100 may treat multiple child listeners as multiple tracks of an audio signal, and set sound pressure distribution in a virtual space and various filters.

[0282] (4. Other Modifications) The control device that controls the sound processing device 100 of the embodiment may be realized by a dedicated computer system or a general-purpose computer system.

[0283] For example, an acoustic processing program for executing the above-described operations is stored in a computer-readable recording medium such as an optical disk, a semiconductor memory, a magnetic tape, or a flexible disk and distributed. Then, for example, the program is installed in a computer and the above-described processing is executed to configure a control device. In this case, the control device may be a device external to the sound processing device 100 (e.g., a personal computer). Alternatively, the control device may be a device internal to the sound processing device 100 (e.g., the control unit 130).

[0284] The communication program may also be stored in a disk device provided in a server on a network such as the Internet, and may be downloaded to a computer. The above-described functions may also be realized by cooperation between an operating system (OS) and application software. In this case, the components other than the OS may be stored on a medium and distributed, or may be stored on a server and downloaded to a computer.

[0285] Furthermore, among the processes described in the above embodiments, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically using a known method. In addition, the information including the processing procedures, specific names, various data, and parameters shown in the above documents and drawings can be changed as desired unless otherwise specified. For example, the various information shown in each drawing is not limited to the information shown in the drawings.

[0286] Furthermore, the components of each device shown in the figure are conceptual functional units and do not necessarily have to be physically configured as shown. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of the devices can be functionally or physically distributed or integrated in any unit depending on various loads, usage conditions, etc. Note that this distribution or integration configuration may also be performed dynamically.

[0287] The above-described embodiments can be combined as appropriate within the scope of the present invention without causing any inconsistency in the processing content. The order of the steps shown in the flowcharts of the above-described embodiments can be changed as appropriate.

[0288] Furthermore, for example, the embodiments can be implemented as any configuration that constitutes an apparatus or system, such as a processor as a system LSI (Large Scale Integration), a module using multiple processors, a unit using multiple modules, a set in which other functions are added to a unit, or the like (i.e., a configuration of a part of an apparatus).

[0289] In the embodiments, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all the components are in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device housed in a single housing with multiple modules, are both systems.

[0290] Furthermore, for example, the embodiment may have a cloud computing configuration in which a single function is shared and processed jointly by a plurality of devices via a network.

[0291] (5. Effects of Sound Processing Device According to the Present Disclosure) A sound processing device according to the present disclosure (sound processing device 100 in the embodiment) includes an acquisition unit (acquisition unit 131 in the embodiment), a correction unit (correction unit 132 in the embodiment), and a generation unit (generation unit 133 in the embodiment). The acquisition unit acquires information relating to the movement of a character that serves as the sound receiving point of a sound source in a virtual space. The correction unit corrects parameters used to generate an audio signal at the sound receiving point based on the information relating to the character's movement. The generation unit generates an audio signal at the sound receiving point based on the parameters corrected by the correction unit.

[0292] In this way, when a character that serves as a sound-receiving point moves, the sound processing device according to the present disclosure corrects parameters for generating an audio signal by moving the sound-receiving point ahead, adjusting spatial characteristics such as reverberation time for voice synthesis, gain, delay, etc. This allows the sound processing device to make it easier for the user to sense the characteristics of the preceding space, and to provide the user with an emphasis on changes in sound that humans perceive due to changes in space, thereby enhancing the user's sense of immersion in the virtual space.

[0293] The acquisition unit acquires, as information relating to the character's movement, a direction or speed at which the character is attempting to move. The correction unit places a second sound receiving point at a position separated from the character based on the direction or speed at which the character is attempting to move. The generation unit generates a sound signal at the second sound receiving point as the sound received by the character, instead of generating a sound signal at the sound receiving point.

[0294] In this way, the sound processing device sets a virtual sound receiving point (in games and the like, the existence of such a virtual character separated from the character is sometimes called a "ghost") at a position away from the character operated by the user. That is, the sound processing device sets the second sound receiving point separated from the character as the coordinates for synthesizing the sound heard by the user and generates sound. This allows the user to recognize sounds and the like in a space at a preceding position along the direction of movement of the character, and therefore, can quickly grasp information about the space to which the character will move.

[0295] The acquisition unit also acquires information about the character's field of view as information about the character's movement. The correction unit corrects the distance or direction from the character in which the second sound receiving point is located, based on the information about the character's field of view.

[0296] In this way, the sound processing device corrects the position of the second sound receiving point based on the field of view of the character, i.e., the orientation of the virtual camera displaying the virtual space. For example, the sound processing device sets a range of coordinates in which the second sound receiving point is set, assuming an elliptical shape as shown in Fig. 13. This allows the sound processing device to prevent the second sound receiving point from being set at a position that is inconsistent with the field of view of the character, thereby making it possible to provide the user with natural sound.

[0297] Furthermore, when the acquisition unit acquires the position where the second sound receiving point is to be placed based on the direction or speed at which the character is to move, the acquisition unit acquires a localization difference, which is an angle formed between the sound source and the position of the character and the position at which the second sound receiving point is to be placed. The correction unit corrects the distance from the character at which the second sound receiving point is actually placed so that the localization difference falls within an allowable range.

[0298] In this way, the sound processing device may adjust the position of the second sound receiving point so that the angle formed by the sound source, the character, and the second sound receiving point does not exceed a certain range. This allows the sound processing device to prevent the user from feeling unnaturalness that may occur when the sound receiving point is shifted from the character.

[0299] The acquisition unit also acquires, as information related to the movement of the character, whether an object exists at the target point to which the character is to move. The correction unit corrects the distance or direction from the character in which the second sound receiving point is to be located, based on the existence of the object.

[0300] In this way, when an object exists at the target point to which the character is to move, i.e., the coordinates at which the second sound receiving point is to be set, the sound processing device adjusts the setting so that the second sound receiving point is not set at those coordinates. For example, as shown in Fig. 15, when the second sound receiving point is to be set on an object such as a wall, the sound processing device temporarily narrows the movement range of the second sound receiving point to adjust the setting so that the second sound receiving point is not set inside the wall. This allows the sound processing device to prevent unnatural voice synthesis from being performed, such as in which the user hears sounds inside the wall.

[0301] The acquisition unit acquires a first setting value directed in the front-to-back direction of the character and a second setting value directed to the side of the character as a range in which the second sound receiving point is to be placed. The correction unit places the second sound receiving point within a range defined based on the acquired first setting value and second setting value.

[0302] That is, the sound processing device acquires setting values ​​in the front, rear, and lateral directions of the character, and determines the movement range of the second sound receiving point (an elliptical shape in the example of Fig. 13) based on the acquired setting values, as shown in Fig. 13. This allows the sound processing device to determine the second sound receiving point as desired by the content creator or user, and therefore makes it possible to set sounds that the user can hear within a range that is in line with the intention of the content creator or user.

[0303] The acquisition unit acquires information about the character's movement that is related to a space predicted to be the character's destination. The correction unit corrects the characteristics of the space in which the character is located based on the information about the space predicted to be the character's destination. The generation unit generates an audio signal at the sound receiving point based on the characteristics of the space corrected by the correction unit.

[0304] In this way, the sound processing device may perform not only indirect correction processing by changing the sound receiving point of the user, but also direct correction processing by synthesizing sound using spatial parameters of the destination. This processing allows the user to more clearly recognize environmental changes accompanying the character's movement, thereby further enhancing the sense of immersion in the virtual space.

[0305] The acquisition unit acquires information about the character's movement, including a path taken by the character from the previous frame to the current frame, and the correction unit corrects the volume (gain) or delay time (delay) of an audio signal played at the character's location in the current frame, based on the path taken by the character.

[0306] In this way, the sound processing device may generate sounds that are emphasized according to the character's moving speed, direction, etc., such as when the character is moving toward the sound source at a speed faster than normal. This allows the sound processing device to further emphasize changes in the sound environment that accompany the character's movement, thereby providing the user with a tense sound environment.

[0307] The acquisition unit acquires the positional relationship between the position of the character and the arrangement of the object. The correction unit corrects the frequency or sound pressure of the sound perceived by the character at the sound receiving point based on the positional relationship between the position of the character and the arrangement of the object.

[0308] In this way, the sound processing device may perform correction processing in accordance with real physical phenomena when an object such as a wall is located near a character, as shown in Figures 19 and 20. This allows the user to experience a more realistic sound environment in the virtual space.

[0309] The correction unit may also correct the parameters so that the closer the position of the character and the position of the object are, the more the sound in a predetermined frequency band at the sound receiving point is emphasized.

[0310] In this way, the sound processing device may emphasize frequency bands in accordance with real physical phenomena. For example, the sound processing device may correct the sound so that low-frequency sounds near walls are emphasized more than usual. This allows the sound processing device to provide the user with realistic sounds, such as the sense of narrowness felt when standing close to a wall.

[0311] The acquisition unit acquires, as sound receiving points corresponding to the character, positions of a right sound receiving point corresponding to the right ear and a left sound receiving point corresponding to the left ear, and a positional relationship between the position of the object. The correction unit may correct the frequency or sound pressure of the sound perceived by the character at the right sound receiving point and the left sound receiving point, based on the positional relationship between the right sound receiving point and the left sound receiving point and the position of the object.

[0312] In this way, when a character has multiple sound receiving points set, such as a right ear and a left ear, the sound processing device can realistically perform different filter processing on each ear, thereby reproducing a more realistic sound environment.

[0313] The acquisition unit also acquires, as information regarding the character's movement, whether the character is close to a space predicted as the character's destination in the transition from the previous frame to the current frame. If the character is close to the space predicted as the character's destination, the correction unit corrects the characteristics of the space in which the character is located so as to weight the characteristics of the space predicted as the character's destination more heavily.

[0314] In this way, the sound processing device may not only set parameters gradually for the range of adjacent spaces as shown in Fig. 18, but may also emphasize changes related to movement by further adjusting parameters according to the direction of movement, etc. In this way, the sound processing device allows the user to better sense, via the character, changes in sound that accompany movement through the space.

[0315] The acquisition unit also acquires, as information related to the character's movement, a speed at which the character moves to the destination from the previous frame to the current frame. The correction unit corrects the volume or delay time of the audio signal played at the position of the character in the current frame based on the character's speed of movement.

[0316] In this way, the sound processing device may correct the volume or delay time depending on not only the path but also the character's movement speed when changing from the previous frame to the current frame. By using such processing, the sound processing device can allow the user to better sense, via the character, the change in sound that accompanies movement through space.

[0317] The correction unit may also correct at least one of a direct sound, an early reflected sound, a late reverberation sound, and a diffracted sound at a position where the character is located, based on the moving speed of the character.

[0318] 2, the sound synthesis includes multiple processes, and the sound processing device can emphasize sound elements that are easily perceived by the user, such as direct sound and diffracted sound, or can emphasize sound-characterizing elements, such as late reverberation and diffracted sound, thereby enabling the user to more easily perceive changes in sound.

[0319] The acquisition unit acquires information about the movement of objects other than the character operated by the user. The correction unit corrects the characteristics of the space in which the character is located based on information about the space predicted to be the object's destination. The generation unit generates an audio signal at the sound receiving point based on the characteristics of the space corrected by the correction unit.

[0320] In this way, even when the character is not moving, the sound processing device may perform the above-mentioned emphasis processing in response to the movement of an object or NPC that may be a sound source. This allows the sound processing device to provide the user with a more emphasized change in sound that the user perceives as the sound source approaches, thereby increasing the user's sense of immersion and creating a sense of tension or urgency in the content.

[0321] The acquisition unit also acquires information about the user's gaze area as information about movement. The correction unit corrects parameters used to generate audio signals at the sound receiving point based on the information about the gaze area. For example, the acquisition unit acquires information about a direction or position set by the user as a destination or target for the character. The correction unit corrects parameters used to generate audio signals at the sound receiving point based on the information about the direction or position set as a destination or target for the character.

[0322] In this way, even if the character does not move, the sound processing device may perform the above-mentioned correction process when the character takes an action such as turning toward a sound source (changing the direction of the camera), aiming in the direction of the line of sight, etc. This allows the sound processing device to perform a variety of effects in content, such as emphasizing sounds in the direction the user is looking at in an FPV game or the like, so that the user can sense them.

[0323] The sound processing device may also have the following configuration. That is, the sound processing device includes an acquisition unit, a generation unit, and an output control unit (output control unit 134 in this embodiment). The acquisition unit acquires information about the sound emission point in the virtual space and a first sound receiving point (parent listener 320 in this embodiment), which is the original sound receiving point. The generation unit generates a third sound receiving point (child listener 330 in this embodiment), which is a sound receiving point different from the first sound receiving point, based on information about an object placed in the virtual space. The output control unit controls the sound output to the user by synthesizing sounds observed at the first sound receiving point and the third sound receiving point.

[0324] In this way, the sound processing device may generate a third sound receiving point at an arbitrary position in addition to the first sound receiving point, and output a sound obtained by synthesizing the sounds observed at the first sound receiving point and the third sound receiving point. This allows the sound processing device to easily realize effective sound processing because it can create the sound audibility by arranging the third sound receiving point, adjusting the volume, etc., without the user having to manually create an acoustic design. Specifically, the sound processing device makes it possible to easily add the acoustic effects desired by creators and users in environments such as UGC where it is difficult to create sound.

[0325] The generation unit may generate the third sound receiving point when an object that emits a sound toward the player operated by the user is detected.

[0326] In this way, the sound processing device may generate a child listener when detecting some kind of action, such as speaking to a user, in game content, etc. This allows the sound processing device to adjust the voice speaking to the user to make it easier to hear, or to create a sense of realism as if the voice were being spoken right next to the user's ear, without the creator having to take the time to adjust the audibility.

[0327] The generation unit may also generate a third sound receiving point on a line connecting the player operated by the user and the object that emits a sound to the player.

[0328] In this way, the sound processing device can automatically generate child listeners on the lines connecting the players, thereby reproducing natural sounds that are easy for the user to hear.

[0329] The generation unit may also generate a third sound receiving point near an object in the virtual space based on attribute information assigned to the object.

[0330] In this way, the sound processing device may generate child listeners near objects that are set as sound generators in their attribute information or that are set as playing sound effects in their audio metatag information. That is, the sound processing device can set listener positions by interpolation or extrapolation using not only the user's player but also other objects. This allows the sound processing device to place child listeners in positions that are appropriate for the type and volume balance of sounds emitted by each object, even when a large number of objects are placed in a virtual space, thereby reproducing an optimized acoustic space.

[0331] The generation unit may also generate the third sound receiving point based on the content that constructs the virtual space or on concept information set in a scene within the content.

[0332] In this way, the sound processing device generates child listeners in accordance with the concept set for each content or scene of the content, such as "creating a sense of fear." This allows the sound processing device to easily and effortlessly create an acoustic environment or an environment that can serve as a starting point for creators' development work. Specifically, the sound processing device can standardize the sound development workflow by separating the project (e.g., game) and listener control. Furthermore, because the sound processing device controls listeners independently of the content, development work can be standardized across different game titles, for example, reducing creators' learning costs for the game engine.

[0333] The generation unit may also generate a third sound receiving point outside the viewing angle of a virtual camera placed in the virtual space.

[0334] In this way, by generating child listeners outside the viewing angle, the sound processing device can realize a variety of effects, such as hearing sounds emitted from positions outside the user's field of view in a stealth game, for example.

[0335] In addition, the generation unit may generate a third sound receiving point at a position beyond the normal hearing range of the player operated by the user, based on the user's operation or information set in the content that constructs the virtual space.

[0336] In this way, the sound processing device can realize sound effects that are expected to produce excellent effects in terms of presentation, even though they do not match the physical characteristics of reality, such as listening to sounds by magnifying them with a scope.

[0337] The output control unit also adjusts the synthesis ratio between the sound observed at the first sound receiving point and the sound observed at the third sound receiving point based on information set in the content that constructs the virtual space.

[0338] In this way, the sound processing device can provide the user with suitable audio by adjusting the volume and the like of the sound observed at multiple sound receiving points.

[0339] The output control unit may also adjust the synthesis ratio based on the content that constructs the virtual space or on concept information set in a scene within the content.

[0340] In this way, the sound processing device can effectively produce sound by adjusting the volume according to a set concept, such as by emphasizing the sound of child listeners near the sound producing body.

[0341] The output control unit may also adjust the synthesis ratio based on attribute information assigned to an object that is a sound generator for the sound observed at the third sound receiving point.

[0342] In this way, the sound processing device can easily reproduce sound effects in cooperation with the game engine by adjusting the volume according to the role of the object, for example, based on audio meta tag information previously set on the game engine side.

[0343] (6. Hardware Configuration Example) An information device such as the sound processing device 100 according to the above-described embodiment is realized by a computer 1000 having a configuration as shown in Fig. 21, for example. Fig. 21 is a hardware configuration diagram showing an example of the computer 1000 that realizes the functions of the sound processing device 100. The computer 1000 has a CPU 1100, a RAM 1200, a ROM (Read Only Memory) 1300, an SSD (Solid State Drive) 1400, a communication interface 1500, and an input / output interface 1600. The components of the computer 1000 are connected by a bus 1050.

[0344] The CPU 1100 operates and controls each component based on programs stored in the ROM 1300 or the SSD 1400. For example, the CPU 1100 loads the programs stored in the ROM 1300 or the SSD 1400 into the RAM 1200 and executes processing corresponding to the various programs.

[0345] The ROM 1300 stores boot programs such as a Basic Input Output System (BIOS) that is executed by the CPU 1100 when the computer 1000 is started, and programs that depend on the hardware of the computer 1000 .

[0346] The SSD 1400 is a computer-readable recording medium that non-temporarily records programs executed by the CPU 1100 and data used by such programs. Specifically, the SSD 1400 is a recording medium that records an information processing program according to the present disclosure, which is an example of the program data 1450. The SSD 1400 may be another non-temporary recording medium, such as a hard disk drive (HDD).

[0347] The communication interface 1500 is an interface for connecting the computer 1000 to an external network 1550 (e.g., the Internet). For example, the CPU 1100 receives data from other devices and transmits data generated by the CPU 1100 to other devices via the communication interface 1500.

[0348] The input / output interface 1600 is an interface for connecting the input / output device 1650 and the computer 1000. For example, the CPU 1100 receives data from input devices such as a touch panel, keyboard, mouse, microphone, and camera via the input / output interface 1600. The CPU 1100 also transmits data to output devices such as a display, speaker, and printer via the input / output interface 1600. The input / output interface 1600 may also function as a media interface for reading programs and the like recorded on a predetermined recording medium. Examples of media include optical recording media such as DVDs (Digital Versatile Discs) and PDs (Phase Change Rewritable Disks), magneto-optical recording media such as MOs (Magneto-Optical Disks), tape media, magnetic recording media, and semiconductor memories.

[0349] For example, when the computer 1000 functions as the sound processing device 100 according to the embodiment, the CPU 1100 of the computer 1000 executes an information processing program loaded onto the RAM 1200 to realize the functions of the control unit 130, etc. The information processing program according to the present disclosure and data in the storage unit 120 are stored in the SSD 1400. The CPU 1100 reads and executes the program data 1450 from the SSD 1400, but as another example, the CPU 1100 may obtain these programs from another device via an external network 1550.

[0350] Note that the present technology can also be configured as follows. (1) A sound processing device comprising: an acquisition unit that acquires information about the movement of a character that serves as a sound receiving point of a sound source in a virtual space; a correction unit that corrects parameters used for generating an audio signal at the sound receiving point based on the information about the character's movement; and a generation unit that generates an audio signal at the sound receiving point based on the parameters corrected by the correction unit. (2) The sound processing device described in (1), wherein the acquisition unit acquires, as information about the character's movement, a direction or speed at which the character is attempting to move, the correction unit places a second sound receiving point at a position separated from the character based on the direction or speed at which the character is attempting to move, and the generation unit generates, as sound received by the character, an audio signal at the second sound receiving point instead of generating an audio signal at the sound receiving point. (3) The sound processing device described in (2), wherein the acquisition unit acquires information about the character's field of view as information about the character's movement, and the correction unit corrects the distance or direction from the character at which the second sound receiving point is to be placed based on the information about the character's field of view. (4) The sound processing device according to (3), wherein, when the acquisition unit acquires the position where the second sound receiving point is to be placed based on the direction or speed at which the character is to move, the acquisition unit acquires a localization difference which is an angle formed between the sound source and the position of the character and the position at which the second sound receiving point is to be placed, and the correction unit corrects the distance from the character at which the second sound receiving point is actually placed so that the localization difference is within an allowable range. (5) The sound processing device according to (3) or (4), wherein the acquisition unit acquires, as a range in which the second sound receiving point is to be placed, a first setting value directed in the forward and backward directions of the character and a second setting value directed to the side of the character, and the correction unit places the second sound receiving point within a range defined based on the first setting value and the second setting value.(6) The sound processing device according to any of (2) to (5), wherein the acquisition unit acquires, as information regarding the movement of the character, whether or not an object exists at a target point to which the character is to move, and the correction unit corrects the distance or direction from the character in which the second sound receiving point is to be placed, based on the existence of the object. (7) The sound processing device according to any of (1) to (6), wherein the acquisition unit acquires, as information regarding the movement of the character, information regarding a space predicted to be a destination of the character, and the correction unit corrects characteristics of the space in which the character is located, based on the information regarding the space predicted to be a destination of the character, and the generation unit generates an audio signal at the sound receiving point, based on the characteristics of the space corrected by the correction unit. (8) The sound processing device according to (7), wherein the acquisition unit acquires, as information regarding the movement of the character, a path taken by the character in the transition from the previous frame to the current frame, and the correction unit corrects a volume or a delay time of an audio signal to be played at a position of the character in the current frame, based on the path taken by the character. (9) The sound processing device according to any one of (1) to (8), wherein the acquisition unit acquires a positional relationship between a position where the character is located and an arrangement of an object, and the correction unit corrects a frequency or a sound pressure of a sound perceived by the character at the sound receiving point based on the positional relationship between the position where the character is located and the arrangement of the object. (10) The sound processing device according to (9), wherein the correction unit corrects the parameter so that the closer the position where the character is located and the arrangement of the object becomes, the more emphasized the sound of a predetermined frequency band at the sound receiving point becomes.(11) The sound processing device according to any of (9) to (10), wherein the acquisition unit acquires a positional relationship between positions of a right sound receiving point corresponding to the right ear and a left sound receiving point corresponding to the left ear as sound receiving points corresponding to the character and an arrangement of an object, and the correction unit corrects the frequency or sound pressure of the sound perceived by the character at the right sound receiving point and the left sound receiving point based on the positional relationship between the right sound receiving point and the left sound receiving point and the arrangement of the object. (12) The sound processing device according to any of (7) to (11), wherein the acquisition unit acquires, as information regarding the movement of the character, whether the character is close to a space predicted as the movement destination in the transition from the previous frame to the current frame, and the correction unit corrects the characteristics of the space in which the character is located so as to weight more heavily the characteristics of the space predicted as the movement destination when the character is close to the space predicted as the movement destination. (13) The sound processing device according to any of (7) to (12), wherein the acquisition unit acquires, as information related to the movement of the character, a movement speed at which the character moves to a destination in the transition from a previous frame to a current frame, and the correction unit corrects a volume or a delay time of an audio signal played at a location of the character in the current frame based on the movement speed of the character. (14) The sound processing device according to (13), wherein the correction unit corrects at least one of a direct sound, an early reflected sound, a late reverberation sound, or a diffracted sound at a location of the character based on the movement speed of the character. (15) The sound processing device according to any of (1) to (14), wherein the acquisition unit acquires information related to movement of an object other than the character operated by a user, and the correction unit corrects characteristics of a space in which the character is located based on information related to a space predicted as a destination of the object, and the generation unit generates an audio signal at the sound receiving point based on the characteristics of the space corrected by the correction unit.(16) The sound processing device according to any one of (1) to (15), wherein the acquisition unit acquires information about a user's gaze area as information about the movement, and the correction unit corrects parameters used for generating an audio signal at the sound receiving point based on the information about the gaze area. (17) A sound processing device comprising: an acquisition unit that acquires, as information about the movement of a character that is a sound receiving point of a sound source in a virtual space, a direction or speed at which the character is attempting to move, a correction unit that corrects the position of the sound receiving point based on the direction or speed at which the character is attempting to move, to place a second sound receiving point at a position separated from the character, and a generation unit that generates an audio signal at the second sound receiving point placed by the correction unit instead of generating an audio signal at the sound receiving point. (18) A sound processing device comprising: an acquisition unit that acquires, as information about the movement of a character that is a sound receiving point of a sound source in a virtual space, information about a space that is predicted to be a destination of the character, a correction unit that corrects parameters used for generating an audio signal at the sound receiving point based on the information about the movement of the character, and a generation unit that generates an audio signal at the sound receiving point based on the parameters corrected by the correction unit, wherein the correction unit corrects characteristics of the space in which the character is located based on information about the space that is predicted to be a destination of the character, and the generation unit generates an audio signal at the sound receiving point based on the characteristics of the space corrected by the correction unit. (19) A sound processing method including: a computer acquires information about the movement of a character that is a sound receiving point of a sound source in a virtual space, corrects parameters used for generating an audio signal at the sound receiving point based on the information about the movement of the character, and generates an audio signal at the sound receiving point based on the corrected parameters.(20) A sound processing program that causes a computer to function as a sound processing device comprising: an acquisition unit that acquires information regarding the movement of a character that serves as a sound receiving point of a sound source in a virtual space; a correction unit that corrects parameters used for generating a sound signal at the sound receiving point based on the information regarding the movement of the character; and a generation unit that generates a sound signal at the sound receiving point based on the parameters corrected by the correction unit. (21) A sound processing device comprising: an acquisition unit that acquires information about a sound emission point in a virtual space and a first sound receiving point that is the original sound receiving point; a generation unit that generates a third sound receiving point that is different from the first sound receiving point based on information about an object placed in the virtual space; and an output control unit that controls sound output to a user by synthesizing sounds observed at the first sound receiving point and the third sound receiving point. (22) The sound processing device according to (21), wherein the generation unit generates the third sound receiving point when it detects an object that emits a sound toward a player operated by the user. (23) The sound processing device according to (22), wherein the generation unit generates the third sound receiving point on a line connecting a player operated by the user and an object that emits sound to the player. (24) The sound processing device according to any of (21) to (23), wherein the generation unit generates the third sound receiving point near an object in the virtual space based on attribute information assigned to the object. (25) The sound processing device according to any of (21) to (24), wherein the generation unit generates the third sound receiving point based on concept information set in content that constructs the virtual space or a scene within the content. (26) The sound processing device according to any of (21) to (25), wherein the generation unit generates the third sound receiving point outside the viewing angle of a virtual camera that is placed in the virtual space. (27) The sound processing device according to any one of (21) to (26), wherein the generation unit generates the third sound receiving point at a position beyond the normal hearing range of a player operated by the user, based on an operation by the user or information set in content that constructs the virtual space.(28) The sound processing device according to any one of (21) to (27), wherein the output control unit adjusts a synthesis ratio between the sound observed at the first sound receiving point and the sound observed at the third sound receiving point based on information set in content that constructs the virtual space. (29) The sound processing device according to (28), wherein the output control unit adjusts the synthesis ratio based on concept information set in the content that constructs the virtual space or in a scene within the content. (30) The sound processing device according to (28) or (29), wherein the output control unit adjusts the synthesis ratio based on attribute information assigned to an object that serves as a sound emitter for the sound observed at the third sound receiving point. (31) A sound processing method including: a computer acquiring information on a sound emission point in the virtual space and a first sound receiving point that is an original sound receiving point; generating a third sound receiving point that is different from the first sound receiving point based on information on an object to be placed in the virtual space; and controlling a sound to be output to a user by synthesizing the sounds observed at the first sound receiving point and the third sound receiving point. (32) A sound processing program for causing a computer to function as a sound processing device comprising: an acquisition unit that acquires information on a sound emission point in a virtual space and a first sound receiving point that is the original sound receiving point; a generation unit that generates a third sound receiving point that is a sound receiving point different from the first sound receiving point based on information on an object placed in the virtual space; and an output control unit that controls the sound output to a user by synthesizing sounds observed at the first sound receiving point and the third sound receiving point.

[0351] REFERENCE SIGNS LIST 100 Sound processing device 110 Communication unit 120 Storage unit 130 Control unit 131 Acquisition unit 132 Correction unit 133 Generation unit 134 Output control unit 140 Input unit 150 Output unit

Claims

1. An audio processing device comprising: an acquisition unit that acquires information relating to the movement of a character that is a sound receiving point of a sound source in a virtual space; a correction unit that corrects parameters used to generate an audio signal at the sound receiving point based on the information relating to the character's movement; and a generation unit that generates an audio signal at the sound receiving point based on the parameters corrected by the correction unit.

2. The sound processing device of claim 1, wherein the acquisition unit acquires, as information regarding the movement of the character, a direction or speed in which the character is attempting to move, the correction unit places a second sound receiving point at a position separated from the character based on the direction or speed in which the character is attempting to move, and the generation unit generates, instead of generating a sound signal at the sound receiving point, a sound signal at the second sound receiving point as the sound received by the character.

3. The sound processing device described in claim 2, wherein the acquisition unit acquires information regarding the field of view of the character as information regarding the movement of the character, and the correction unit corrects the distance or direction from the character in which the second sound receiving point is located based on the information regarding the field of view of the character.

4. The sound processing device described in claim 3, wherein the acquisition unit, when acquiring the position at which the second sound receiving point is to be placed based on the direction or speed at which the character is to move, acquires a localization difference which is an angle formed between the sound source and the position of the character and the position at which the second sound receiving point is to be placed, and the correction unit corrects the distance from the character at which the second sound receiving point is actually placed so that the localization difference is within an acceptable range.

5. The sound processing device of claim 3, wherein the acquisition unit acquires a first setting value directed toward the front and rear directions of the character and a second setting value directed toward the side of the character as the range in which the second sound receiving point is to be positioned, and the correction unit positions the second sound receiving point within a range defined based on the first setting value and the second setting value.

6. The sound processing device of claim 2, wherein the acquisition unit acquires, as information regarding the movement of the character, whether or not an object exists at a target point to which the character is to move, and the correction unit corrects the distance or direction from the character in which the second sound receiving point is located based on the existence of the object.

7. The sound processing device described in claim 1, wherein the acquisition unit acquires information regarding the space predicted to be the destination of the character as information regarding the movement of the character, the correction unit corrects characteristics of the space in which the character is located based on the information regarding the space predicted to be the destination of the character, and the generation unit generates an audio signal at the sound receiving point based on the characteristics of the space corrected by the correction unit.

8. The audio processing device according to claim 7, wherein the acquisition unit acquires, as information regarding the movement of the character, a path traveled by the character in the transition from a previous frame to a current frame, and the correction unit corrects the volume or delay time of an audio signal played at a location of the character in the current frame based on the path traveled by the character.

9. The sound processing device of claim 1, wherein the acquisition unit acquires the positional relationship between the position of the character and the arrangement of an object, and the correction unit corrects the frequency or sound pressure of the sound perceived by the character at the sound receiving point based on the positional relationship between the position of the character and the arrangement of an object.

10. The sound processing device according to claim 9, wherein the correction unit corrects the parameters so that the closer the position of the character and the position of the object are, the more the sound of a predetermined frequency band at the sound receiving point is emphasized.

11. The sound processing device of claim 9, wherein the acquisition unit acquires the positional relationship between the positions of a right sound receiving point corresponding to the right ear and a left sound receiving point corresponding to the left ear as the sound receiving points corresponding to the character, and an arrangement of an object, and the correction unit corrects the frequency or sound pressure of the sound perceived by the character at the right sound receiving point and the left sound receiving point, respectively, based on the positional relationship between the right sound receiving point and the left sound receiving point and the arrangement of an object.

12. The sound processing device of claim 7, wherein the acquisition unit acquires, as information regarding the movement of the character, whether or not the character is close to the space predicted as the destination when transitioning from the previous frame to the current frame, and the correction unit corrects the characteristics of the space in which the character is located so as to weight more heavily the characteristics of the space predicted as the destination when the character is close to the space predicted as the destination.

13. The audio processing device of claim 7, wherein the acquisition unit acquires, as information regarding the movement of the character, a moving speed at which the character moves to its destination in the transition from the previous frame to the current frame, and the correction unit corrects the volume or delay time of an audio signal played at the position of the character in the current frame based on the moving speed of the character.

14. The sound processing device according to claim 13, wherein the correction unit corrects at least one of a direct sound, an early reflected sound, a late reverberation sound, and a diffracted sound at a position where the character is located based on a moving speed of the character.

15. The sound processing device of claim 1, wherein the acquisition unit acquires information regarding the movement of an object other than the character operated by a user, the correction unit corrects characteristics of the space in which the character is located based on information regarding the space predicted to be the destination of the object, and the generation unit generates an audio signal at the sound receiving point based on the characteristics of the space corrected by the correction unit.

16. The sound processing device of claim 1, wherein the acquisition unit acquires information relating to the user's gaze area as the information regarding the movement, and the correction unit corrects parameters used to generate an audio signal at the sound receiving point based on the information regarding the gaze area.

17. An audio processing device comprising: an acquisition unit that acquires, as information relating to the movement of a character that is a sound receiving point of a sound source in a virtual space, a direction or speed in which the character is attempting to move; a correction unit that places a second sound receiving point at a position separated from the character by correcting the position of the sound receiving point based on the direction or speed in which the character is attempting to move; and a generation unit that generates an audio signal at the second sound receiving point placed by the correction unit instead of generating an audio signal at the sound receiving point.

18. An audio processing device comprising: an acquisition unit that acquires, as information regarding the movement of a character that is a sound receiving point of a sound source in a virtual space, information regarding a space predicted as a destination of the character; a correction unit that corrects parameters used in generating an audio signal at the sound receiving point based on the information regarding the movement of the character; and a generation unit that generates an audio signal at the sound receiving point based on the parameters corrected by the correction unit, wherein the correction unit corrects characteristics of the space in which the character is located based on information regarding the space predicted as a destination of the character, and the generation unit generates an audio signal at the sound receiving point based on the characteristics of the space corrected by the correction unit.

19. An acoustic processing method including the steps of: a computer acquiring information relating to the movement of a character that is a sound receiving point of a sound source in a virtual space; correcting parameters used to generate an audio signal at the sound receiving point based on the information about the character's movement; and generating an audio signal at the sound receiving point based on the corrected parameters.

20. An acoustic processing program for causing a computer to function as an acoustic processing device comprising: an acquisition unit that acquires information relating to the movement of a character that is a sound receiving point of a sound source in a virtual space; a correction unit that corrects parameters used to generate an audio signal at the sound receiving point based on the information relating to the character's movement; and a generation unit that generates an audio signal at the sound receiving point based on the parameters corrected by the correction unit.

21. A sound processing device comprising: an acquisition unit that acquires information on a sound emission point in a virtual space and a first sound receiving point, which is the original sound receiving point; a generation unit that generates a third sound receiving point, which is a sound receiving point different from the first sound receiving point, based on information on an object placed in the virtual space; and an output control unit that controls the sound output to a user by synthesizing sounds observed at the first sound receiving point and the third sound receiving point.

22. The sound processing device according to claim 21, wherein the generation unit generates the third sound receiving point when an object that emits a sound toward a player operated by the user is detected.

23. The sound processing device according to claim 22, wherein the generation unit generates the third sound receiving point on a line connecting a player operated by the user and an object that emits a sound to the player.

24. The sound processing device according to claim 21, wherein the generation unit generates the third sound receiving point in the vicinity of the object in the virtual space based on attribute information assigned to the object.

25. The sound processing device according to claim 21, wherein the generation unit generates the third sound receiving point based on concept information set in the content that constructs the virtual space or in a scene within the content.

26. The sound processing device according to claim 21, wherein the generation unit generates the third sound receiving point outside a viewing angle of a virtual camera arranged in the virtual space.

27. The sound processing device according to claim 21, wherein the generation unit generates the third sound receiving point at a position beyond the normal hearing range of a player operated by the user, based on the user's operation or information set in content that constructs the virtual space.

28. The sound processing device according to claim 21, wherein the output control unit adjusts a synthesis ratio between the sound observed at the first sound receiving point and the sound observed at the third sound receiving point based on information set in the content that constructs the virtual space.

29. The sound processing device according to claim 28, wherein the output control unit adjusts the synthesis ratio based on concept information set in the content that constructs the virtual space or in a scene within the content.

30. The sound processing device according to claim 28, wherein the output control unit adjusts the synthesis ratio based on attribute information assigned to an object that is a sound generating body for the sound observed at the third sound receiving point.

31. An acoustic processing method including: a computer acquiring information on a sound emission point in a virtual space and a first sound receiving point which is the original sound receiving point; generating a third sound receiving point which is different from the first sound receiving point based on information on an object placed in the virtual space; and controlling the sound output to a user by synthesizing sounds observed at the first sound receiving point and the third sound receiving point.

32. An acoustic processing program for causing a computer to function as an acoustic processing device comprising: an acquisition unit that acquires information on a sound emission point in a virtual space and a first sound receiving point, which is the original sound receiving point; a generation unit that generates a third sound receiving point, which is a sound receiving point different from the first sound receiving point, based on information on an object placed in the virtual space; and an output control unit that controls the sound output to a user by synthesizing sounds observed at the first sound receiving point and the third sound receiving point.

Citation Information

Patent Citations

  • Spatial audio processing that emphasizes sound sources close to the focal distance

    JP2019514293A

  • Method and system for handling local transitions between listening positions in a virtual reality environment

    JP2021507558A

  • Sound auralization apparatus and sound auralization program

    JP2022182625A

  • Signal processing device and method, and program

    WO2020255810A1