Signal processing device, signal processing method, and program
Patent Information
- Application Number
- JP2021080447
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-05-11
- Publication Date
- 2025-06-02
- Estimated Expiration
- 2041-05-11
AI Technical Summary
Existing systems for generating free-viewpoint sound in virtual spaces do not adequately consider the movement of virtual cameras, leading to inappropriate sound signal output.
A signal processing device that includes mechanisms for generating both fixed and free-viewpoint sound signals based on the motion information of virtual cameras, selecting between them according to specific trajectory patterns or acoustic scenes, using sound selection units and rendering units to output appropriate sound signals.
Enables the output of appropriate sound signals that match the movement of virtual cameras, providing realistic sound experiences in virtual environments.
Smart Images

Figure 00000021_0000 
Figure 00000022_0000 
Figure 00000022_0001
Abstract
Description
Technical Field
[0001] The present disclosure relates to a signal processing apparatus, a signal processing method, and a program.
Background Art
[0002] There is a system that generates a virtual camera (virtual camera) captured video, so-called free viewpoint video, by arranging a virtual camera at an arbitrary position in a virtual world constructed by CG (computer graphics) objects or the like. In such a system, there is also a system that generates an acoustic sound corresponding to the free viewpoint video (hereinafter referred to as free viewpoint sound). As a method for generating free viewpoint sound, there is a technique in which a sound source is assigned to a CG object that pronounces, and acoustic processing is performed in consideration of the direction and distance of the CG object, a shielding object, etc. as seen from the virtual camera or the listening point. Furthermore, a method of generating free viewpoint sound by synthesizing sound sources for a plurality of CG objects is generally performed.
[0003] Patent Document 1 discloses a technique for associating the L channel and the R channel of a stereo signal with different sound source objects in a virtual space, and determining the playback volume of an audio signal reproduced from a first speaker and a second speaker according to the position and orientation of a virtual microphone.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] Here, we consider live music content in a virtual space. When music is to be played, the audio signals (sound sources) for each channel may be positioned at a fixed location relative to the listener, regardless of the position of the virtual camera in the virtual space. Furthermore, depending on the movement of the virtual camera, outputting the audio signals in accordance with the movement of the virtual camera may result in a more immersive sound experience. Patent Document 1 did not consider outputting appropriate audio signals in accordance with the movement of the virtual camera. This disclosure has been made in view of these circumstances and aims to enable the output of appropriate audio signals in accordance with the movement of the virtual camera. [Means for solving the problem]
[0006] The signal processing device according to this disclosure is characterized by comprising: a first sound generation means for generating an acoustic signal of a first sound in which one or more sound sources relating to an acoustic signal are fixedly arranged at predetermined positions; a second sound generation means for generating an acoustic signal of a second sound in which one or more of the sound sources are arranged based on motion information including the position and orientation of a virtual camera in the sound generation target space; and a selection means for selecting an acoustic signal to output either the acoustic signal of the first sound or the acoustic signal of the second sound, according to the movement of the virtual camera based on the motion information of the virtual camera. [Effects of the Invention]
[0007] According to this disclosure, it is possible to output an appropriate acoustic signal in accordance with the movement of the virtual camera. [Brief explanation of the drawing]
[0008] [Figure 1] This figure shows an example of the functional configuration of a signal processing device. [Figure 2] This figure shows an example of the data structure of sound source information. [Figure 3] This figure shows an example of the hardware configuration of a signal processing device. [Figure 4] This flowchart shows an example of sound generation processing in a signal processing device. [Figure 5] This figure shows an example of the data structure of virtual camera information. [Figure 6] This figure shows an example of the data structure of virtual sound source information. [Figure 7] This flowchart shows an example of virtual camera trajectory analysis processing. [Figure 8] This flowchart shows an example of fixed sound generation processing. [Figure 9] This flowchart shows an example of virtual sound source generation processing. [Figure 10] This flowchart shows an example of free-viewpoint acoustic rendering. [Figure 11] This figure shows an example of the functional configuration of a signal processing device. [Figure 12] This flowchart shows an example of sound generation processing in a signal processing device. [Figure 13] This flowchart shows an example of acoustic scene detection processing. [Figure 14] This figure shows an example of the functional configuration of a signal processing device. [Figure 15] This flowchart shows an example of sound generation processing in a signal processing device. [Figure 16] This figure shows an example of the functional configuration of a signal processing device. [Figure 17] This flowchart shows an example of sound generation processing in a signal processing device. [Figure 18] This figure shows an example of the functional configuration of a signal processing device. [Figure 19] This figure shows an example of the functional configuration of a signal processing device. [Modes for carrying out the invention]
[0009] Embodiments of this disclosure will be described below with reference to the drawings. The embodiments described below are not limiting to this disclosure, and not all combinations of features described in these embodiments are necessarily essential configurations. The same components will be denoted by the same reference numerals.
[0010] <Embodiment 1> FIG. 1 is a diagram showing a functional configuration example of a signal processing apparatus in the present embodiment. The signal processing apparatus in the present embodiment includes a virtual camera information receiving unit 1, a virtual camera trajectory analysis unit 2, a specific trajectory database (specific trajectory DB) 3, a free viewpoint sound generation unit 4, a sound selection unit 8, a fixed sound generation unit 9, and a sound output unit 10.
[0011] The virtual camera information receiving unit 1 receives virtual camera information from the outside and outputs the received virtual camera information to the virtual camera trajectory analysis unit 2 and the virtual sound source generation unit 5. Details of the virtual camera information will be described later. The virtual camera trajectory analysis unit 2 analyzes a trajectory related to the movement of the virtual camera in the sound generation target space based on the virtual camera information. Further, the virtual camera trajectory analysis unit 2 determines whether the trajectory of the virtual camera obtained by the analysis satisfies a specific condition. The virtual camera trajectory analysis unit 2 in the present embodiment determines whether it matches a trajectory pattern based on the trajectory data stored in the specific trajectory DB 3, and outputs the result of the determination to the sound selection unit 8. The specific trajectory DB 3 stores trajectory data related to specific movements (trajectory patterns) of the virtual camera, such as rotation, concentric circles, and high-speed movement.
[0012] The free viewpoint sound generation unit 4 generates a free viewpoint sound signal from a plurality of input sound signals (sound signal 1 to sound signal n) based on the virtual camera information. The free viewpoint sound generation unit 4 generates a sound signal of a free viewpoint sound in which one or a plurality of sound sources related to the sound signals are arranged according to the movement of the virtual camera in the sound generation target space. The free viewpoint sound generation unit 4 includes a virtual sound source generation unit 5 and a free viewpoint sound rendering unit 6. The virtual sound source generation unit 5 uses the virtual camera information received from the virtual camera information receiving unit 1 to convert the absolute coordinates of the sound sources stored in the sound source information 7 into relative coordinates viewed from the virtual camera. By associating this with the input sound signals, virtual sound sources are generated. The free viewpoint sound rendering unit 6 renders a plurality of virtual sound sources output from the virtual sound source generation unit 5 according to a predetermined channel format based on the relative coordinates of the sound sources to generate a free viewpoint sound signal. The generated free viewpoint sound signal is output to the sound output unit 10.
[0013] The sound source information 7 stores various information relating to each sound source, with a one-to-one correspondence between each of the multiple acoustic signals input to the virtual sound source generation unit 5. Figure 2(a) shows an example of the data structure of the sound source information. As shown in Figure 2(a), the sound source information includes a sound source ID 201, an acoustic signal channel 202, a total number of sound source trajectory information entries 203, and multiple sound source trajectory information entries 204 (204-1 to 204-m). The sound source ID 201 is an ID number uniquely assigned to identify the sound source information. The acoustic signal channel 202 stores the channel number of the acoustic signal corresponding to this sound source information from among the multiple acoustic signals input to the virtual sound source generation unit 5. The total number of sound source trajectory information entries 203 stores the number of sound source trajectory information entries stored in this sound source information. In this example, sound source trajectory information 204 is created for all frames. Figure 2(b) shows an example of the data structure of the sound source trajectory information 204. As shown in Figure 2(b), each of the sound source trajectory information 204 includes a time code 211 and absolute coordinates 212. The time code 211 consists of hours, minutes, seconds, and frames, and indicates the time when the acoustic signal corresponding to this sound source trajectory information was acquired. The absolute coordinates 212 are the absolute coordinates of this sound source in the acoustic generation target space, and store position information (xs, ys, zs) for the X, Y, and Z axes of a three-dimensional orthogonal coordinate system uniquely defined by the acoustic generation target space.
[0014] The sound selection unit 8 selects whether to generate a fixed sound signal or a free-viewpoint sound signal based on the result of determining whether the virtual camera trajectory output from the virtual camera trajectory analysis unit 2 matches a specific trajectory pattern. The selection result by the sound selection unit 8 is notified to the virtual sound source generation unit 5, the free-viewpoint sound rendering unit 6, and the fixed sound generation unit 9, respectively. The fixed sound generation unit 9 generates a fixed sound signal in which one or more sound sources related to the sound signal are fixedly positioned at predetermined locations. The fixed sound generation unit 9 renders each input sound signal according to a specified output channel format so that the sound source of each sound signal is localized at a predetermined position, and generates a fixed sound signal by adding them channel by channel. The generated fixed sound signal is output to the sound output unit 10.
[0015] The sound output unit 10 outputs either the free-viewpoint sound signal output by the free-viewpoint sound rendering unit 6 of the free-viewpoint sound generation unit 4, or the fixed sound signal output by the fixed sound generation unit 9, to the speaker set 11. The sound output unit 10 performs amplification and correction for the effects of the speaker installation environment on the output free-viewpoint sound signal or fixed sound signal before outputting it to the speaker set 11. The speaker set 11 consists of a number of speakers according to a predetermined channel format, and each speaker converts the signals of each channel of the sound signal output by the sound output unit 10 into sound and outputs it.
[0016] Figure 3 shows an example of the hardware configuration of the signal processing device in this embodiment. The signal processing device in this embodiment includes an input / output unit 301, a CPU 302, a RAM 303, an external storage unit 304, an operation unit 305, a display unit 306, a ROM 307, a communication IF unit 308, and a bus 309. The input / output unit 301, CPU 302, RAM 303, external storage unit 304, operation unit 305, display unit 306, ROM 307, and communication IF unit 308 are connected to each other via the bus 309 so that they can communicate with one another.
[0017] The input / output unit 301 receives input from the outside, such as acoustic signals and virtual camera information, and sends them to other components via the bus 309 according to the instructions of the CPU 302. The CPU (Central Processing Unit) 302 comprehensively controls each component of the information processing device. The CPU 302 controls other components by sending control signals via the bus 309 according to the program, and also performs various calculations. In this embodiment, the CPU 302 executes each function described in Figure 1 by executing processing according to the program stored in the ROM 307 and the external storage unit 304. The RAM (Random Access Memory) 303 temporarily stores a part of the program being executed, associated data, and the calculation results of the CPU 302. The CPU 302 loads the necessary program and data into the RAM 303 and executes the program by reading and writing as needed.
[0018] The external storage unit 304 stores the program itself and data to be stored long-term. The functions of the specific trajectory DB3 are realized by the external storage unit 304. The external storage unit 304 is, for example, an HDD (hard disk drive) or an SSD (solid state drive). The operation unit 305 receives various user instructions and operations, converts them into control signals, and transmits them to the CPU 302 via the bus 309. The CPU 302 performs control instructions for the running program and other configurations according to the control signals. The display unit 306 displays the status of the running program and the output of the program to the user. The ROM (Read Only Memory) 307 stores fixed programs and fixed parameters, such as programs for starting and stopping this hardware device, and programs for controlling basic input and output. The communication interface unit 308 can input and output data to and from communication networks such as the Internet.
[0019] The sound generation process by the signal processing device in this embodiment will now be described. Figure 4 is a flowchart showing an example of the sound generation process of the signal processing device in this embodiment. In S401, the signal processing unit performs initialization. During initialization, the CPU 302 determines the values of various information according to default values stored in ROM 307 or user operations on the operation unit 305. The determined values of the various information are transferred and stored in a predetermined area of RAM 303.
[0020] In S402, the signal processing unit performs acoustic signal reception processing. In acoustic signal reception processing, multiple channels of acoustic signals input to the virtual sound source generation unit 5 and the fixed sound generation unit 9 are received via the input / output unit 301 and the communication IF unit 308, and stored in the external storage unit 304 for each input channel. Here, the stored acoustic signals include signals acquired at the time corresponding to all time codes included in the virtual camera information received in the following S403. These acoustic signals may be real-world sounds recorded by a microphone, or they may be artificially synthesized sounds according to the state of the target sound source at the time of acquisition.
[0021] In S403, the signal processing unit performs virtual camera information reception processing. Virtual camera information reception processing is the process of receiving virtual camera information transmitted from an external source via user operation of the input / output unit 301, the operation unit 305, or the communication IF unit 308, in the virtual camera information receiving unit 1. In this example, for the sake of explanation, it is assumed that information for multiple consecutive frames is received at once as virtual camera information. The virtual camera information received in S403 is stored in RAM 303.
[0022] The data structure of the virtual camera information will be explained with reference to Figures 5(a) and 5(b). Figure 5(a) is a diagram showing an example of the overall data structure of the virtual camera information. As shown in Figure 5(a), the virtual camera information includes the pixel format 501, the frame rate 502, the total number of frames 503, and virtual camera frame information 504 (504-1 to 504-n) for the number of total frames. The pixel format 501 stores information about the pixel format of the video captured by the virtual camera, such as HD (1920×1080) or SD (720×480). The frame rate 502 stores the frame rate of the captured video, for example, 59.94 fps (frames / second). The total number of frames 503 stores the number of frames included in this virtual camera information.
[0023] The virtual camera frame information 504 is information that summarizes information defined individually for each frame. Figure 5(b) shows an example of the data structure of the virtual camera frame information 504. As shown in Figure 5(b), the virtual camera frame information 504 includes the frame sequence number 511, time code 512, virtual camera absolute coordinates 513, virtual camera orientation 514, virtual camera tilt 515, and zoom magnification 516. The frame sequence number 511 is a number that continues from the frame in which shooting began, and is set to increase by 1 when one frame is advanced. Since the frame sequence number 511 is assigned one-to-one to the virtual camera frame information, it can also be used to identify individual frame information. The time code 512 consists of hours, minutes, seconds, and frame, and indicates the time when the video signal or audio signal corresponding to this virtual camera frame information was acquired. To account for cases such as changing the playback speed of the generated audio signal, playing it in reverse, or stopping the timecode from advancing to operate only the virtual camera, the same value may be added to multiple virtual camera frame information entries in timecode 512.
[0024] The virtual camera absolute coordinates 513 are the absolute coordinates of the virtual camera in the sound generation target space, and store position information (x, y, z) with respect to the X, Y, and Z axes of a three-dimensional orthogonal coordinate system uniquely determined by the sound generation target space. The virtual camera orientation 514 is information indicating the front direction of the virtual camera, and stores the angle α in the YZ plane, the angle β in the ZX plane, and the angle γ in the XY plane. The virtual camera tilt 515 stores information indicating the tilt of the virtual camera, with the state where the front-to-back direction is parallel to the X axis, the left-to-right direction is parallel to the Y axis, and the up-to-down direction is parallel to the Z axis as the reference tilt. The virtual camera tilt 515 stores the angle between the front-to-back direction of the virtual camera and the X axis as yaw, the angle between the left-to-right direction of the virtual camera and the Y axis as roll, and the angle between the up-to-down direction of the virtual camera and the Z axis as pitch. The zoom magnification 516 stores the zoom magnification of the virtual camera.
[0025] With this data structure, the trajectory related to the movement of the virtual camera can be obtained by arranging the absolute coordinates of the virtual camera in the order of the frame numbers in the virtual camera frame information 504. Similarly, by arranging specific parameters in the order of the frame numbers in the virtual camera frame information 504, it is possible to understand the temporal changes of those parameters.
[0026] In S404, the virtual camera trajectory analysis unit 2 performs virtual camera trajectory analysis processing and analyzes the trajectory related to the movement of the virtual camera in the sound generation target space based on the virtual camera information received in S403. Based on the virtual camera information, the virtual camera trajectory analysis unit 2 analyzes whether a specific trajectory pattern is occurring in the trajectory of the virtual camera and the motion velocity of the virtual camera. Details of this virtual camera trajectory analysis processing will be described later with reference to Figure 7.
[0027] In S405, the acoustic selection unit 8 determines, based on the analysis results from the virtual camera trajectory analysis unit 2 in S404, whether or not a specific trajectory corresponding to the trajectory data stored in the specific trajectory DB3 has been detected for the virtual camera trajectory. If no specific trajectory corresponding to the trajectory data stored in the specific trajectory DB3 has been detected for the virtual camera trajectory (NO in S405), the process proceeds to S406. If a specific trajectory has been detected (YES in S405), the process proceeds to S407.
[0028] In S406, the fixed sound generation unit 9 performs fixed sound generation processing and generates a fixed sound signal by rendering the acoustic signals of each input channel in a predetermined direction in a predetermined channel format. Details of this fixed sound generation processing will be described later with reference to Figure 8.
[0029] On the other hand, in S407, the acoustic selection unit 8 determines whether the motion speed of the virtual camera calculated in S404 is within a predetermined speed, based on the analysis results from the virtual camera trajectory analysis unit 2 in S404. If the motion speed of the virtual camera is within the predetermined speed (YES in S407), the process proceeds to S408; otherwise, it proceeds to S406.
[0030] In S408, the virtual sound source generation unit 5 performs virtual sound source generation processing, and generates virtual sound source information by creating a one-to-one correspondence between the sound source trajectory and the input acoustic signal, and by calculating the sound source position as seen from the virtual camera. Details of this virtual sound source generation processing will be described later with reference to Figure 9.
[0031] The data structure of the virtual sound source information will be explained with reference to Figures 6(a) and 6(b). Figure 6(a) is a diagram showing an example of the data structure of virtual sound source information. As shown in Figure 6(a), the virtual sound source information includes a sound source ID 601, a total number of frames 602, and virtual sound source frame information 603 (603-1 to 603-1) for the number of total frames. The sound source ID 601 is the same number as the sound source ID 201 stored in the sound source information shown in Figure 2(a), and indicates the sound source corresponding to this virtual sound source information. The sound source ID 601 is also used as a number to identify individual virtual sound source information and is assigned one-to-one to each piece of virtual sound source information. The total number of frames 602 stores the total number of frames included in this virtual sound source information.
[0032] The virtual sound source frame information 603 is stored in a number equal to the total number of frames. Figure 6(b) shows an example of the data structure of the virtual sound source frame information 603. As shown in Figure 6(b), the virtual sound source frame information 603 includes the frame sequence number 611, the time code 612, the virtual sound source relative coordinates 613, and the frame sound signal 614. The frame sequence number 611 and the time code 612 are the same as the frame sequence number 511 and time code 512 included in the virtual camera frame information 504 described above, and indicate the chronological relationship between the time of sound generation and recording of this virtual sound source frame information. The virtual sound source relative coordinates 613 store the position coordinates (x's, y's, z's) of the virtual sound source in a three-dimensional coordinate system based on the position and orientation of the virtual camera at the time of video capture for this frame sequence number 611. The frame sound signal 614 stores the sound signal corresponding to this virtual sound source frame information. Alternatively, the audio signal may be stored in external memory, and a pointer to the beginning of the audio signal corresponding to this virtual sound source frame information may be stored.
[0033] Returning to Figure 4, in S409, the free-viewpoint acoustic rendering unit 6 performs free-viewpoint acoustic rendering processing and generates a free-viewpoint acoustic signal based on the virtual sound source information generated by the virtual sound source generation unit 5 in S408. The free-viewpoint acoustic rendering unit 6 generates a free-viewpoint acoustic signal at the virtual camera position by rendering each virtual sound source according to a predetermined channel format based on the generated virtual sound source information. Details of this free-viewpoint acoustic rendering process will be described later with reference to Figure 10.
[0034] In S410, the audio output unit 10 corrects and amplifies the fixed audio signal generated in S408 or the free-viewpoint audio signal generated in S409, and outputs it to the speaker set 11. This enables the reproduction of actual sound.
[0035] In S411, the signal processing device determines whether the listener has given a termination command via the operation of the operation unit 105. If a termination command has been given by the listener (YES in S411), the sound generation process is terminated. If there is no termination command (NO in S411), the process returns to S402 and the sound generation process continues.
[0036] In this embodiment, the signal processing device generates a free-viewpoint acoustic signal if a specific trajectory is detected in the movement of the virtual camera based on the determination in S405 and S407, and the movement speed of the virtual camera is within a predetermined speed; otherwise, it generates a fixed-viewpoint acoustic signal. In other words, in this embodiment, a fixed-viewpoint acoustic signal is basically generated and output, but if a specific movement pattern of the virtual camera is detected, a free-viewpoint acoustic signal can be generated and output as appropriate. Therefore, it is possible to switch between fixed-viewpoint acoustic and free-viewpoint acoustic according to the movement of the virtual camera, and even in the case of music content, it is possible to provide sound that is more appropriate for free-viewpoint video.
[0037] Figure 7 is a flowchart showing an example of the virtual camera trajectory analysis process in S404 of Figure 4. All processes in the virtual camera trajectory analysis process shown in Figure 7 are executed in the virtual camera trajectory analysis unit 2.
[0038] In S701, the virtual camera trajectory analysis unit 2 generates a trajectory related to the movement of the virtual camera based on the virtual camera information transmitted from the virtual camera information receiving unit 1. The virtual camera trajectory analysis unit 2 generates the trajectory of the virtual camera by connecting the virtual camera absolute coordinates 513, virtual camera orientation 514, and virtual camera tilt 515 of the virtual camera frame information 504 included in the virtual camera information in sequential frame number order.
[0039] In S702, the virtual camera trajectory analysis unit 2 generates the trajectory of each sound source by connecting the absolute coordinates 212 of the sound source trajectory information 204 in the sound source information 7 in a sequential timecode order. In S703, the virtual camera trajectory analysis unit 2 calculates the change in distance between the virtual camera and each sound source based on the trajectory of the virtual camera generated in S701 and the trajectories of each sound source generated in S702.
[0040] In S704, the virtual camera trajectory analysis unit 2 performs pattern matching between the data stored in the specific trajectory DB3 and the virtual camera trajectory generated in S701 and the distance change between the virtual camera and the sound source calculated in S703. The patterns detected here include, for example, trajectory patterns where the virtual camera moves along the edges of a circle or square, or trajectory patterns where it moves in a spiral. There are also patterns of changes in orientation or tilt, such as the virtual camera rotating without moving its position. Furthermore, there are patterns of changes in the positional relationship between the virtual camera and the sound source, such as the virtual camera approaching the sound source from a distance, or the virtual camera moving between multiple sound sources. Note that such pattern matching is generally widely performed and publicly known, so details will not be explained. If a match is found in the pattern matching, the result is temporarily stored in RAM303.
[0041] In S705, the virtual camera trajectory analysis unit 2 calculates the motion velocity of the virtual camera based on the virtual camera trajectory and frame count generated in S701. The motion velocity includes the speed of movement as the coordinates change and the rotational speed of the virtual camera itself. In S706, the virtual camera trajectory analysis unit 2 outputs the pattern matching results from S704 and the virtual camera's motion velocity calculated in S705 to the acoustic selection unit 8. After completing the processing in S706, the virtual camera trajectory analysis process is finished and the unit returns.
[0042] Figure 8 is a flowchart showing an example of the fixed sound generation process in S406 of Figure 4. All processes in the fixed sound generation process shown in Figure 8 are performed in the fixed sound generation unit 9. In S801, the fixed sound generation unit 9 initializes the output buffer it has internally.
[0043] The following processing steps S802-S815 are loop processing steps performed on all incoming acoustic signals. In S803, the fixed sound generation unit 9 divides the acoustic signal to be processed into frame lengths. In S804, the fixed sound generation unit 9 determines whether the current processing is the first processing after switching from free-viewpoint acoustics. In this example, the fixed sound generation unit 9 makes this determination by checking the type of the previous processing recorded in RAM 303. If the result of the determination is that a switch from free-viewpoint acoustics has occurred (YES in S804), the process proceeds to S805; otherwise (NO in S804), the process proceeds to S808.
[0044] In S805, the fixed sound generation unit 9 sets the initial direction for rendering this sound signal to the virtual sound source direction received from the virtual sound source generation unit 5. The set sound source direction is stored in a specified area of RAM 303.
[0045] In S806, the fixed sound generation unit 9 calculates the difference angle between the virtual sound source direction received from the virtual sound source generation unit 5 and the predetermined sound source direction of the sound signal to be processed. In S807, the fixed sound generation unit 9 calculates the movement angle for each frame based on the difference angle calculated in S806. For example, the fixed sound generation unit 9 calculates the movement angle for each frame by dividing the difference angle calculated in S806 equally by a predetermined number of frames.
[0046] The following processing in S808-S814 is a loop processing performed on the acoustic signal for each frame length divided in S803. In S809, the fixed sound generation unit 9 determines whether the sound source direction stored in the specified area of RAM 103 is the fixed sound source direction defined in this device. Here, the fixed sound source direction is the direction in which each sound signal is localized when generating a fixed sound signal, and is predetermined in this device for each sound signal. If the result of the determination in S809 is that the sound source direction is not the predetermined fixed sound source direction (NO in S809), the process proceeds to S810; if the sound source direction is the predetermined fixed sound source direction (YES in S809), the process proceeds to S811.
[0047] In S810, the fixed sound generation unit 9 adds the frame-by-frame movement angle calculated in S807 to the sound source direction stored in RAM 303. If the sound source direction stored in RAM 103 is not the default fixed sound source direction due to the processing in S809 and S810, the sound source direction will be moved closer to the fixed sound source direction by the amount of the movement angle for each frame. Therefore, when switching from free-viewpoint sound to fixed sound, the sound source direction will be moved from the sound source direction in free-viewpoint sound to the fixed sound source direction by a predetermined number of transition frames.
[0048] In S811, the fixed sound generation unit 9 selects three output channels that are closest to the sound source direction from among the multiple channels defined in the output channel format. In S812, the fixed sound generation unit 9 determines the distribution of the acoustic signals stored in the virtual sound source to the three channels selected in S811 using the VBAP (Vector Based Amplitude Panning) method. In this embodiment, the VBAP method is used, but other three-dimensional acoustic rendering methods may also be used.
[0049] In S813, the fixed sound generation unit 9 distributes the sound signal according to the distribution to each output channel determined in S812. The sound signals thus created are added to the sound signals in the output buffers of each channel. Through this processing, the sound image of the input sound signal is rendered according to the output channel format so that it appears in the direction of the sound source.
[0050] In S814, the fixed sound generation unit 9 determines whether the loop processing from S808 to S814 for all the sound signals divided by frame length in S803 has finished. If it has not finished, the fixed sound generation unit 9 returns to S808 and continues processing for the next divided sound signal. If it has finished, it proceeds to S815. In S815, the fixed sound generation unit 9 determines whether the loop processing from S802 to S815 for all sound signals has been completed. If it has not been completed, the fixed sound generation unit 9 returns to S802 and continues processing for the next sound signal. If it has been completed, it proceeds to S816.
[0051] In S816, the fixed sound generation unit 9 outputs the acoustic signal rendered in the output buffer as a fixed acoustic signal to the acoustic output unit 10. After completing the processing in S816, the fixed sound generation process is finished and the unit returns.
[0052] Figure 9 is a flowchart showing an example of the virtual sound source generation process in S408 of Figure 4. All processes in the virtual sound source generation process shown in Figure 9 are executed in the virtual sound source generation unit 5. In S901, the virtual sound source generation unit 5 initializes the virtual sound source information list to a state where there are no elements. Here, the virtual sound source information list is a data structure that manages multiple virtual sound source information in a list, and is stored in RAM 103.
[0053] The following processing steps S902-S915 are loop processing steps performed on all sound source information included in sound source information 7. In S903, the virtual sound source generation unit 5 acquires the next sound source information to be processed in the loop processing and stores it in RAM 303. In S904, the virtual sound source generation unit 5 generates new virtual sound source information. The generated virtual sound source information is temporarily stored in RAM 303.
[0054] The following processing steps S905 to S912 are loop processing steps performed on all virtual camera frame information stored in the virtual camera information received in step S403 of Figure 4. In S906, the virtual sound source generation unit 5 adds new virtual sound source frame information to the virtual sound source information generated in S904. In S907, the virtual sound source generation unit 5 stores the frame sequence number and timecode of the virtual camera frame information directly into the virtual sound source frame information added in S906.
[0055] In S908, the virtual sound source generation unit 5 obtains the absolute coordinates of the sound source stored in the sound source trajectory information of the same time code from the sound source trajectory information stored in the sound source information to be processed, and temporarily stores them in RAM 303. In S909, the virtual sound source generation unit 5 calculates the relative coordinates of the virtual sound source as seen from the virtual camera, using the absolute coordinates of the virtual camera stored in the virtual camera frame information and the absolute coordinates of the sound source acquired in S908. If, for the target frame sequence, the absolute coordinates of the virtual camera are A(x,y,z), the orientation of the virtual camera is (α,β,γ), and the sound source coordinates of the corresponding timecode are B(xs,ys,zs), then the relative coordinates C(x's,y's,z's) of the virtual sound source are calculated by the following formula.
[0056]
number
[0057] In S910, the virtual sound source generation unit 5 stores the relative coordinates calculated in S909 into the virtual sound source frame information. In S911, the virtual sound source generation unit 5 extracts a signal corresponding to the time code of the virtual sound source frame information from the acoustic signal specified by the sound source ID of the sound source information stored in the external storage unit 304 in S402 of Figure 4, and stores it in the virtual sound source frame information. This process will store values in all components within the virtual sound source frame information.
[0058] In S912, the virtual sound source generation unit 5 determines whether the loop processing from S905 to S912 for all virtual camera frame information within the target virtual camera information has finished. If it has not finished, the virtual sound source generation unit 5 returns to S906 and processes the next virtual sound source frame information. If it has finished, it proceeds to S913. In S913, the virtual sound source generation unit 5 stores the total number of frames in the virtual sound source information generated in S904. This stores all data components of the virtual sound source information. In S914, the virtual sound source generation unit 5 adds the virtual sound source information to be processed to the virtual sound source information list.
[0059] In S915, the virtual sound source generation unit 5 determines whether the loop processing from S902 to S915 for all sound source information has finished. If it has not finished, the virtual sound source generation unit 5 returns to S903 and processes the next sound source information. If it has finished, it proceeds to S916. In S916, the virtual sound source generation unit 5 outputs a virtual sound source information list to the free-viewpoint sound rendering unit 6, which includes the virtual camera information and all the virtual sound source information generated in the previous processing.
[0060] In S917, the virtual sound source generation unit 5 outputs the final virtual sound source frame information for each virtual sound source to the fixed sound generation unit 9. After completing the processing in S917, the virtual sound source generation process is finished and the unit returns. In this embodiment, a virtual sound source is generated by converting the coordinates of the sound source trajectory given by the sound source information into relative coordinates as seen from the virtual camera, and by extracting and combining the input acoustic signal into frame units included in the virtual camera information.
[0061] Figure 10 is a flowchart showing an example of the free-viewpoint acoustic rendering process in S409 of Figure 4. All processes in the free-viewpoint acoustic rendering process shown in Figure 10 are performed in the free-viewpoint acoustic rendering unit 6. In S1001, the free-viewpoint acoustic rendering unit 6 initializes the output buffer. In this embodiment, the output buffer of the free-viewpoint acoustic rendering unit 6 is allocated in a predetermined area of the RAM 303. In S1002, the free-viewpoint sound rendering unit 6 receives virtual camera information and a virtual sound source information list output from the virtual sound source generation unit 5.
[0062] The following processing steps S1003 to S1010 are loop processing steps performed on all virtual camera frame information stored in the virtual camera information received in S1002. Furthermore, the processing in S1004 to S1009 is a loop process performed on all virtual sound source information included in the virtual sound source information list received in S1002.
[0063] In S1005, the free-viewpoint acoustic rendering unit 6 calculates the distance from the virtual camera to the virtual sound source based on the virtual sound source relative coordinates stored in the virtual sound source frame information corresponding to the target frame sequence number. Furthermore, the free-viewpoint acoustic rendering unit 6 calculates the distance attenuation rate of the acoustic signal based on the calculated distance from the virtual camera to the virtual sound source. Such processing is commonly performed in the field of acoustics and is publicly known, so a detailed explanation is omitted.
[0064] In S1006, the free-viewpoint sound rendering unit 6 selects three output channels from among the multiple channels defined in the output channel format that are closest to the direction indicated by the virtual sound source relative coordinates mentioned above. In this embodiment, the output channel format is assumed to be a three-dimensional arrangement format such as 5.1.4ch or 22.2ch, as defined in advance.
[0065] In S1007, the free-viewpoint sound rendering unit 6 determines the distribution of the sound signals stored in the virtual sound source to the three channels selected in S1006 using the VBAP method. In this embodiment, the VBAP method is used, but other three-dimensional sound rendering methods may also be used.
[0066] In S1008, the free-viewpoint acoustic rendering unit 6 adjusts the sound pressure level by multiplying the acoustic signal stored in the target virtual sound source frame information by the distance attenuation rate calculated in S1005. Furthermore, the free-viewpoint acoustic rendering unit 6 distributes the acoustic signal according to the distribution to each output channel determined in S1007. The acoustic signals thus created are added to the output buffer of each channel. Through this processing, each virtual sound source signal in the frame being processed is rendered according to the output channel format so that the sound image appears at the direction and distance specified by relative coordinates.
[0067] In S1009, the free-viewpoint sound rendering unit 6 determines whether the loop processing from S1004 to S1009 for all virtual sound source information has finished. If it has not finished, it returns to S1004 and continues processing for the next virtual sound source information. If it has finished, it proceeds to S1010. In S1010, the free-viewpoint acoustic rendering unit 6 determines whether the loop processing from S1003 to S1010 for all virtual camera frame information within the virtual camera information has finished. If it has not finished, it returns to S1003 and continues processing for the next virtual camera frame information. If it has finished, it proceeds to S1011.
[0068] In S1011, the free-viewpoint sound rendering unit 6 outputs the sound signal rendered in the output buffer as a free-viewpoint sound signal to the sound output unit 10. After completing the processing in S1011, the free-viewpoint sound rendering process is finished and the unit returns.
[0069] As described above, according to this embodiment, an acoustic signal is normally generated using fixed acoustics, and when a specific movement is detected in the movement of the virtual camera, the acoustic generation method is switched to free-viewpoint acoustics to generate an acoustic signal. This makes it possible to switch between fixed acoustics and free-viewpoint acoustics according to the movement of the virtual camera in the acoustic generation target space. Therefore, even in music content, it becomes possible to generate acoustics appropriate to the specific movement of the virtual camera.
[0070] In this embodiment, a specific movement pattern of the virtual camera was detected by pattern matching, but this is not limited to this. For example, machine learning may be performed using at least one parameter such as the virtual camera's trajectory, speed, direction, and field of view as variables to determine the pattern. Also, in this embodiment, the localization of the acoustic signal is moved from free-viewpoint acoustics to fixed acoustics over a predetermined number of transition frames, but this is not limited to this, and it may be moved at a predetermined angular velocity.
[0071] <Embodiment 2> Embodiment 2 describes an example in which a content scene is detected from an acoustic signal and a sound generation method is selected accordingly. Note that the same configuration and processing as in Embodiment 1 will not be described.
[0072] Figure 11 shows an example of the functional configuration of the signal processing device in this embodiment. In Figure 11, components having the same function as those shown in Figure 1 are denoted by the same reference numerals, and redundant explanations are omitted.
[0073] The signal processing device in this embodiment includes a virtual camera information receiving unit 1, a virtual camera trajectory analysis unit 2, a specific trajectory DB 3, a free-viewpoint sound generation unit 4, a sound selection unit 8, a fixed sound generation unit 9, a sound output unit 10, and a sound scene detection unit 21. The sound scene detection unit 21 analyzes each input sound signal and detects an overall sound scene based on the results, outputting it to the sound selection unit 8. The sound selection unit 8 selects whether to generate a fixed sound signal or a free-viewpoint sound signal according to the determination result from the virtual camera trajectory analysis unit 2 and the scene detection result from the sound scene detection unit 21.
[0074] In this embodiment, the hardware configuration example of the signal processing device is the same as in Embodiment 1, so it is omitted.
[0075] The sound generation process by the signal processing device in this embodiment will now be described. Figure 12 is a flowchart showing an example of the sound generation process of the signal processing device in this embodiment. The processes in S1201 and S1202 are the same as those in S401 and S402 in Figure 4, respectively, so their explanation will be omitted.
[0076] In S1203, the acoustic scene detection unit 21 performs acoustic scene detection processing and analyzes the acoustic signal received in S1202 to detect an acoustic scene. Details of this acoustic scene detection processing will be described later with reference to Figure 13.
[0077] In S1204, the sound scene detection unit 21 determines whether the sound scene detected in S1203 is a singing scene. If the result of the determination is that it is a singing scene (YES in S1204), the process proceeds to S1205; otherwise, the process proceeds to S1206.
[0078] The process in S1205 is the same as the fixed acoustic processing in S406 in Figure 4. Furthermore, the processes in S1206-S1208 are the same as the processes in S403-S405 in Figure 4. Also, the processes in S1209-S1213 are the same as the processes in S407-S411 in Figure 4. Therefore, explanations of these processes are omitted.
[0079] Figure 13 is a flowchart showing an example of the acoustic scene detection process in S1203 of Figure 12. All processes in the acoustic scene detection process shown in Figure 13 are executed in the acoustic scene detection unit 21. The processing in S1301 to S1314 is a loop process performed on all input acoustic signals.
[0080] In S1302, the acoustic scene detection unit 21 detects speech by analyzing whether the acoustic signal to be processed is voiced or voiceless. Since this type of processing is well known in the field of speech processing, a detailed explanation will be omitted. In S1303, the acoustic scene detection unit 21 performs tempo analysis by analyzing whether changes in sound pressure or overlaps in spectral patterns are detected in the acoustic signal at regular time intervals.
[0081] In S1304, the acoustic scene detection unit 21 performs harmonic analysis on the acoustic signal in the frequency domain to perform tonality analysis (tonality: the proportion of a certain frequency and its multiples of that frequency component in the overall components). In S1305, the acoustic scene detection unit 21 performs formant (approximate shape of frequency spectrum) analysis based on the analysis results from processing S1302 to S1304, and detects singing voices.
[0082] In S1306, the acoustic scene detection unit 21 determines whether speech was detected in S1302 and whether tempo, tonality, or singing voice was not detected in the analysis in S1303 to S1305. If the determination in S1306 is YES, the process proceeds to S1307; if the determination is NO, the process proceeds to S1308.
[0083] In S1307, the sound scene detection unit 21 detects a talk scene in the sound signal and saves the detection result to a specified area in the RAM 303. Meanwhile, in S1308, the sound scene detection unit 21 determines whether the accuracy of the tempo analyzed in S1303 or the tonality analyzed in S1304 is high. If the determination in S1308 is YES, the process proceeds to S1309; if the determination is NO, the process proceeds to S1313.
[0084] In S1309, the sound scene detection unit 21 detects a music scene in the sound signal and saves the detection result to a specified area in the RAM 303. In S1310, the sound scene detection unit 21 determines whether or not singing was detected in S1305. If singing is detected (YES in S1310), the process proceeds to S1311; if it is not detected (NO in S1310), the process proceeds to S1312.
[0085] In S1311, the sound scene detection unit 21 detects a singing scene in the sound signal and saves the detection result to a specified area in the RAM 303. Meanwhile, in S1312, the sound scene detection unit 21 detects an instrumental scene in the sound signal and saves the detection result to a specified area in the RAM 303. In S1313, the acoustic scene detection unit 21 detects that it has detected another scene in the acoustic signal and saves the detection result to a specified area in the RAM 303.
[0086] In S1314, the acoustic scene detection unit 21 determines whether the loop processing from S1301 to S1314 for all input acoustic signals has finished. If it has not finished, the acoustic scene detection unit 21 returns to S1301 and continues processing for the next acoustic signal. If it has finished, it proceeds to S1315.
[0087] In S1315, the acoustic scene detection unit 21 integrates the detection results of all acoustic signals detected up to this point and outputs the final scene detection result to the acoustic selection unit 8. The method of integrating the detection results may be by majority vote, or by weighting each acoustic signal and adding the results. After completing the process in S1315, the acoustic scene detection process is finished and the unit returns.
[0088] As described above, in this embodiment, an acoustic scene is detected from the acoustic signal, and if the acoustic scene is suitable for free-viewpoint acoustics, free-viewpoint acoustics are selected and an acoustic signal is generated. In this way, fixed acoustics are selected and an acoustic signal is generated in scenes where you want to listen to the sound carefully, and free-viewpoint acoustics are automatically selected and an acoustic signal is generated in scenes where you want to emphasize the free-viewpoint image.
[0089] Furthermore, if the sound scene is not a singing scene, the sound generation method is appropriately switched, similar to Embodiment 1, to generate either fixed sound or free-viewpoint sound signals. Therefore, it is possible to switch between fixed sound and free-viewpoint sound according to the movement of the virtual camera in the sound generation target space, and generate sound appropriate to the specific movement of the virtual camera.
[0090] In this embodiment, at S1204 in Figure 12, it is determined whether the sound scene is a singing scene or not, but it may also be determined whether it is a music scene or not. Furthermore, in the sound scene detection process of this embodiment, talk scenes, music scenes, singing scenes, instrumental scenes, and other scenes are detected, but other scenes may also be detected. These can be modified and implemented according to the directorial intent of the free-viewpoint video content, without departing from the spirit of this embodiment.
[0091] <Embodiment 3> Embodiment 3 describes an example in which a sound generation method is selected by selecting a pre-created virtual camera information template. Note that the same configurations and processes as described in the previously mentioned embodiments will not be explained.
[0092] Figure 14 shows an example of the functional configuration of the signal processing device in this embodiment. In Figure 14, components having the same function as those shown in Figure 1 are denoted by the same reference numerals, and redundant explanations are omitted.
[0093] The signal processing device in this embodiment includes a virtual camera information receiving unit 1, a free-viewpoint sound generation unit 4, a fixed sound generation unit 9, a sound output unit 10, a template database (template DB) 31, and a template selection unit 32. The template DB 31 stores templates for virtual camera information. Here, a template for virtual camera information is information that combines virtual camera information, which records various movements of the virtual camera as trajectory information, such as rotation of the virtual camera body, movement along a circle, and approaching from a distance, with a sound generation method. Multiple such templates for virtual camera information are stored in the template DB 31.
[0094] The template selection unit 32 searches the template DB 31 according to the instructions of the CPU 302 and selects a virtual camera information template that matches the search criteria. The template selection unit 32 transmits the virtual camera information 33 of the selected template to the virtual camera information receiving unit 1. The template selection unit 32 also instructs either the virtual sound source generation unit 5 or the fixed sound generation unit 9 to generate an acoustic signal according to the sound generation method stored in the selected template.
[0095] In this embodiment, the hardware configuration example of the signal processing device is the same as in Embodiment 1, so it is omitted.
[0096] The sound generation process by the signal processing device in this embodiment will now be described. Figure 15 is a flowchart showing an example of the sound generation process of the signal processing device in this embodiment. The processes in S1501 and S1502 are the same as those in S401 and S402 in Figure 4, respectively, so their explanation will be omitted.
[0097] In S1503, the template selection unit 32 searches the template DB 31 based on instructions from the CPU 102 and selects an appropriate virtual camera information template.
[0098] In S1504, the template selection unit 32 determines whether the sound generation method of the template selected in S1503 is fixed sound. If the sound generation method is fixed sound (YES in S1504), the template selection unit 32 instructs the fixed sound generation unit 9 to generate an acoustic signal and proceeds to S1505. If the sound generation method is not fixed sound (NO in S1504), the template selection unit 32 instructs the free-viewpoint sound generation unit 4 to generate an acoustic signal and proceeds to S1506.
[0099] The process in S1505 is the same as the fixed acoustic processing in S406 in Figure 4, so its explanation is omitted. Also, the processes in S1506 to S1509 are the same as the processes in S408 to S411 in Figure 4, so their explanation is omitted.
[0100] In this embodiment, the sound generation method is automatically selected by selecting a pre-created virtual camera information template. As a result, in a template that records virtual camera movements suitable for free-viewpoint sound, specifying free-viewpoint sound as the sound generation method allows for the generation of sound signals appropriate to the virtual camera movements.
[0101] <Embodiment 4> Embodiment 4 describes an example of selecting a sound generation method specified for a virtual camera when receiving multiple virtual camera information in parallel and generating multiple sounds. Note that the same configurations and processes as described in the previously mentioned embodiments will not be explained.
[0102] Figure 16 shows an example of the functional configuration of the signal processing device in this embodiment. In Figure 16, components having the same function as those shown in Figure 1 are denoted by the same reference numerals, and redundant explanations are omitted.
[0103] The signal processing device in this embodiment includes a virtual camera information receiving unit 41, an acoustic selection unit 42, a free-viewpoint acoustic generation unit 43, a fixed acoustic generation unit 44, and an acoustic output unit 45. The virtual camera information receiving unit 41 simultaneously receives multiple virtual camera information and outputs the received virtual camera information to the acoustic selection unit 42. In this embodiment, the virtual camera information includes an acoustic generation method in addition to the information shown in Figure 5.
[0104] The sound selection unit 42 checks the sound generation method attached to each of the multiple virtual camera information outputs from the virtual camera information receiving unit 41. Furthermore, the sound selection unit 42 transmits the virtual camera information to the free-viewpoint sound generation unit 43 or the fixed-viewpoint sound generation unit 44 to instruct the generation of an acoustic signal, according to the sound generation method attached to the virtual camera information.
[0105] The free-viewpoint sound generation unit 43 is configured in a multiple parallel configuration. Each of the free-viewpoint sound generation units 43 has a virtual sound source generation unit 5 and a free-viewpoint sound rendering unit 6. Each pair of the virtual sound source generation unit 5 and the free-viewpoint sound rendering unit 6 in the free-viewpoint sound generation unit 43 corresponds to virtual camera information that specifies the generation of free-viewpoint sound, and generates free-viewpoint sound signals in parallel based on multiple pieces of virtual camera information.
[0106] The fixed sound generation unit 44 is configured in a parallel arrangement. Each fixed sound generation unit 44 corresponds to the virtual camera information specified for generating fixed sound, and generates the acoustic signal of the fixed sound in parallel based on the time codes of multiple virtual camera information.
[0107] The sound output unit 45 is a collection of multiple units with the same configuration as the sound output unit 10 shown in Figure 1. It outputs individual sound signals from each of the multiple free-viewpoint sound generation units 43 or multiple fixed sound generation units 44 to the speaker set 11 individually.
[0108] In this embodiment, the hardware configuration example of the signal processing device is the same as in Embodiment 1, so it will be omitted.
[0109] The sound generation process by the signal processing device in this embodiment will now be described. Figure 17 is a flowchart showing an example of the sound generation process of the signal processing device in this embodiment. The processes in S1701 and S1702 are the same as those in S401 and S402 in Figure 4, respectively, so their explanation will be omitted. In S1703, the virtual camera information receiving unit 41 receives multiple virtual camera information simultaneously in parallel and outputs the received virtual camera information to the acoustic selection unit 42.
[0110] The following processes, S1704 to S1710, perform parallel processing on all virtual camera information received in S1703. In S1705, the sound selection unit 42 determines whether the sound generation method for the virtual camera information to be processed is fixed sound. If the sound generation method is fixed sound (YES in S1705), the sound selection unit 42 outputs the virtual camera information to the fixed sound generation unit 44 and proceeds to S1706. If the sound generation method is not fixed sound (NO in S1705), the sound selection unit 42 outputs the virtual camera information to the virtual sound source generation unit 5 of the free-viewpoint sound generation unit 43 and proceeds to S1707.
[0111] In S1706, the fixed sound generation unit 44 performs the same processing as in S406 in Figure 4 on the sound signals within the time code range added to the virtual camera information. The processes in S1707 to S1709 are the same as those in S408 to S410 in Figure 4, so their explanation will be omitted. The above processing is performed in parallel for each virtual camera information, and then the process proceeds to S1711. The process in S1711 is the same as the process in S411 in Figure 4, so its explanation will be omitted.
[0112] In this embodiment, when generating multiple sounds based on multiple virtual camera information, a suitable sound generation method is added to the virtual camera information. This makes it possible to select a sound generation method appropriate for the free-viewpoint video generated based on each virtual camera information.
[0113] <Other Embodiments> In the embodiment described above, the sound selection unit selects a sound generation method to generate either a fixed sound or a free-viewpoint sound signal. However, it is not limited to this, and both fixed sound and free-viewpoint sound signals may always be generated, and the sound selection unit may select which sound signal to output. An example of the configuration of the signal processing device in this case is shown in Figure 18. In Figure 18, components having the same function as those shown in Figure 1 are denoted by the same reference numerals, and redundant explanations are omitted.
[0114] The signal processing device shown in Figure 18 includes a virtual camera information receiving unit 1, a virtual camera trajectory analysis unit 2, a specific trajectory DB 3, a free-viewpoint sound generation unit 4, a fixed sound generation unit 9, a sound output unit 10, and a sound selection unit 51. The sound selection unit 51 selects either a generated free-viewpoint sound or a fixed sound signal based on the analysis results output by the virtual camera trajectory analysis unit 2 and outputs it to the sound output unit 10. The sound selection unit 51 may also be configured to switch between fixed sound and free-viewpoint sound according to a timetable based on a predetermined time code.
[0115] Alternatively, the system may detect the subject's position in the free-viewpoint video generated based on virtual camera information, and select a free-viewpoint audio signal if the relationship between the virtual camera position and the subject's position follows a specific pattern. An example of the signal processing device configuration in this case is shown in Figure 19. In Figure 19, components having the same function as those shown in Figure 1 are denoted by the same reference numerals, and redundant explanations are omitted.
[0116] The signal processing device shown in Figure 19 includes a virtual camera information receiving unit 1, a virtual camera trajectory analysis unit 2, a specific trajectory DB 3, a free-viewpoint sound generation unit 4, a sound selection unit 8, a fixed sound generation unit 9, a sound output unit 10, a free-viewpoint video generation unit 61, and a subject position detection unit 62. The free-viewpoint video generation unit 61 generates a free-viewpoint video based on multiple input video signals (video signal 1 to video signal h) and virtual camera information. The subject position detection unit 62 analyzes the video generated by the free-viewpoint video generation unit 61 to detect the subject position. The detected subject position is output to the virtual camera trajectory analysis unit 2. The virtual camera trajectory analysis unit 2 detects whether the relationship change pattern between the subject position and the virtual camera position is a specific pattern and outputs the detection result to the sound selection unit 8. The sound selection unit 8 selects whether to generate a fixed sound signal or a free-viewpoint sound signal by referring to the pattern detection result from the virtual camera trajectory analysis unit 2.
[0117] Furthermore, in Embodiment 4, the audio output for all virtual camera information is output in parallel to the speaker set 11. However, a video signal in which the free-viewpoint video generated based on each virtual camera information and each audio output are superimposed may be input to a switcher, and the switching may be performed manually. This allows switching between video with fixed audio and video with free-viewpoint audio as needed while checking the video.
[0118] Furthermore, this embodiment can be used for any application of generating sound at a free listening point surrounded by virtual sound sources. For example, it can be used in free-viewpoint sound generation systems and game sound generation systems, as well as as a method for controlling them.
[0119] (Other embodiments of the present invention) This disclosure can also be implemented by supplying a program that implements one or more of the functions of the embodiments described above to a system or device via a network or storage medium, and by a process in which one or more processors in the computer of that system or device read and execute the program. It can also be implemented by a circuit (e.g., an ASIC) that implements one or more functions.
[0120] Furthermore, the embodiments described above are merely examples of how the present invention may be implemented, and the technical scope of this disclosure should not be limited by them. In other words, this disclosure can be implemented in various ways without departing from its technical concept or its main features. [Explanation of Symbols]
[0121] 1, 41: Virtual camera information receiving unit 2: Virtual camera trajectory analysis unit 3: Specific trajectory DB 4: Free viewpoint sound generation unit 5: Virtual sound source generation unit 6: Free viewpoint sound rendering unit 7: Sound source information 8, 42, 51: Sound selection unit 9: Fixed sound generation unit 10, 45: Sound output unit 11: Speaker set 21: Sound scene detection unit 31: Template DB 32: Template selection unit
Claims
1. a first sound generating means for generating a first sound signal by fixedly arranging one or more sound sources related to the sound signal at predetermined positions; a second sound generation means for generating a sound signal of a second sound in which the one or more sound sources are arranged based on movement information including a position and an orientation of a virtual camera in a sound generation target space; and a selection means for selecting either the audio signal of the first audio or the audio signal of the second audio as the audio signal to be output in accordance with the movement of the virtual camera based on the movement information of the virtual camera.
2. an analysis means for analyzing the movement of the virtual camera based on the movement information of the virtual camera; 2. The signal processing device according to claim 1, wherein the selection means selects the audio signal of the second audio as the audio signal to be output when the movement of the virtual camera analyzed by the analysis means satisfies a specific condition.
3. The signal processing device according to claim 2, characterized in that the selection means selects the second audio signal as the audio signal to be output when the movement of the virtual camera analyzed by the analysis means satisfies the specific condition, and the direction or tilt of the virtual camera changes in a predetermined pattern at a specific position.
4. The signal processing device according to claim 2 or 3, characterized in that the selection means selects the second audio signal as the audio signal to be output when the movement of the virtual camera analyzed by the analysis means is a predetermined trajectory as the case where the specific condition is satisfied.
5. The signal processing device according to any one of claims 2 to 4, characterized in that the selection means selects the second acoustic signal as the acoustic signal to be output when the positional relationship between the virtual camera and the sound source changes in a predetermined pattern as the case where the specific condition is satisfied.
6. The signal processing device according to any one of claims 2 to 5, characterized in that the selection means selects the first sound as the sound signal to be output when the movement speed of the virtual camera is not within a predetermined speed as the case where the specific condition is not satisfied.
7. The signal processing device according to any one of claims 2 to 6, characterized in that the selection means selects the first audio signal as the audio signal to be output when the movement of the virtual camera analyzed by the analysis means does not match a specific pattern, as a case in which the specific condition is not met.
8. The signal processing device according to any one of claims 1 to 7, characterized in that, when switching the output acoustic signal from the acoustic signal of the second acoustic to the acoustic signal of the first acoustic, the first acoustic generation means generates the acoustic signal by moving the position of the sound source from the position when the acoustic signal of the second acoustic was selected to the predetermined position that is fixedly arranged over a predetermined time.
9. The signal processing device according to any one of claims 1 to 8, characterized in that, when switching the output acoustic signal from the acoustic signal of the second acoustic to the acoustic signal of the first acoustic, the first acoustic generation means generates the acoustic signal by moving the position of the sound source from the position when the acoustic signal of the second acoustic was selected to the predetermined position that is fixedly arranged at a predetermined angular velocity.
10. a detection means for detecting a scene from the acoustic signals of one or more of said sound sources; The signal processing device according to any one of claims 1 to 9, characterized in that the selection means selects either the acoustic signal of the first acoustic or the acoustic signal of the second acoustic as the acoustic signal to be output based on the detection result of the detection means.
11. 11. The signal processing device according to claim 10, wherein the selection means selects the audio signal of the first audio as the audio signal to be output when the detection means detects a music scene.
12. 11. The signal processing device according to claim 10, wherein the selection means selects the audio signal of the first sound as the audio signal to be output when the detection means detects a singing scene.
13. 2. The signal processing device according to claim 1, wherein the selection means selects either the first acoustic signal or the second acoustic signal as the acoustic signal to be output in accordance with a pre-specified timetable.
14. A signal processing device as described in any one of claims 1 to 13, characterized in that one of the first sound generation means or the second sound generation means generates an acoustic signal in accordance with the movement of the virtual camera based on the movement information of the virtual camera.
15. the first sound generating means and the second sound generating means generate sound signals in parallel, The signal processing device according to any one of claims 1 to 13, characterized in that the selection means selects either the audio signal of the first audio or the audio signal of the second audio as the audio signal to be output in accordance with the movement of the virtual camera based on the movement information of the virtual camera.
16. 2. The signal processing device according to claim 1, wherein the selection means selects either the audio signal of the first audio or the audio signal of the second audio as the audio signal to be output in accordance with a template that records the movement of the virtual camera.
17. 2. The signal processing device according to claim 1, wherein the selection means selects either the first audio signal or the second audio signal as the audio signal to be output in accordance with an audio generation method added to the movement information of the virtual camera.
18. The signal processing device according to claim 1, characterized in that the selection means selects either the first sound signal or the second sound signal as the output sound signal based on the position of the subject of the image generated by the virtual camera and the position of the virtual camera.
19. a first sound generating step of generating a first sound signal by fixedly arranging one or more sound sources related to the sound signal at predetermined positions; a second sound generation step of generating a sound signal of a second sound in which one or more of the sound sources are arranged based on movement information including a position and an orientation of a virtual camera in a sound generation target space; a selection step of selecting either the audio signal of the first audio or the audio signal of the second audio as the audio signal to be output in accordance with the movement of the virtual camera based on the movement information of the virtual camera.
20. A program for causing a computer to function as each of the means of the signal processing device according to any one of claims 1 to 18.