Information processing device, information processing method, program, and automatic sound synthesis system
The information processing device automates the synchronization of virtual viewpoint image and sound data from different spaces by using model acquisition and evaluation units, enhancing efficiency in data synthesis.
Patent Information
- Application Number
- JP2021157900
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-09-28
- Publication Date
- 2025-09-01
- Estimated Expiration
- 2041-09-28
AI Technical Summary
Existing methods require manual synchronization of sound data with virtual viewpoint image data, which is time-consuming and inefficient when the data is collected in different imaging spaces.
An information processing device that includes model acquisition, evaluation, and data generation units to automatically synchronize and combine virtual viewpoint image data and sound data from different imaging spaces based on posture model agreement.
Efficiently synthesizes virtual viewpoint image data and sound data from different spaces, reducing the workload and time required for manual editing.
Smart Images

Figure 0007731632000001 
Figure 0007731632000002 
Figure 0007731632000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a technique for synthesizing sound data with virtual viewpoint image data. [Background technology]
[0002] There is a technology that performs synchronized imaging from multiple viewpoints using imaging devices such as cameras placed at multiple different positions, and generates a virtual viewpoint image based on the multiple captured images (hereinafter referred to as "multiple viewpoint images") obtained by the imaging. Virtual viewpoint images allow objects in an imaging space to be viewed from various angles. Some systems that generate virtual viewpoint images synthesize data of sounds collected during imaging of the multiple viewpoint images (hereinafter referred to as "sound data") with data of the virtual viewpoint image (hereinafter referred to as "virtual viewpoint image data"). In such systems, the sound data is synthesized with the virtual viewpoint image data by synchronizing the time of the virtual viewpoint image with the time of the sounds collected during imaging of the multiple viewpoint images. Patent Document 1 discloses a technology in the field of game systems that outputs sounds corresponding to the movement of a character in a virtual space according to the distance between joints in a joint model of the character moving in a virtual space. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2020-195672 Summary of the Invention [Problem to be solved by the invention]
[0004] Combining virtual viewpoint image data generated based on multiple viewpoint images captured in a first imaging space, such as a studio, with sound data collected in a second imaging space different from the first imaging space, such as a stage, can be achieved by a method in which the user manually synthesizes the data. Specifically, a user who provides a virtual viewpoint image (hereinafter simply referred to as "user") manually synchronizes pre-acquired sound data with the virtual viewpoint image data to synthesize the sound data with the virtual viewpoint image. However, with this method, the user must manually synthesize each of the various pieces of sound data they want to synthesize with the virtual viewpoint image data while creating the virtual viewpoint image. Therefore, this method has the problem of requiring a long time for editing to synthesize the sound data with the virtual viewpoint image data.
[0005] The present disclosure is intended to solve such problems, and has an object to provide an information processing device that efficiently combines virtual viewpoint image data and sound data acquired in different imaging spaces. [Means for solving the problem]
[0006] The information processing device according to the present disclosure includes a first model acquisition means for acquiring data of a first posture model generated based on first multi-viewpoint images obtained from a plurality of imaging devices that capture a first imaging space; an image acquisition means for acquiring data of a virtual viewpoint image generated based on the first multi-viewpoint image; a second model acquisition means for acquiring data of a second posture model generated based on second multi-viewpoint images obtained from a plurality of imaging devices that capture a second imaging space different from the first imaging space; an audio acquisition means for acquiring data of audio collected when the second multi-viewpoint image is captured in the second imaging space; an evaluation means for evaluating the degree of agreement between the first posture model and the second posture model; and a data generation means for generating data including the virtual viewpoint image and audio based on the degree of agreement. [Effects of the Invention]
[0007] According to the present disclosure, virtual viewpoint image data and sound data acquired in different imaging spaces can be efficiently synthesized. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a block diagram illustrating an example of a configuration of an information processing system according to a first embodiment. [Figure 2] 2 is a block diagram showing an example of a hardware configuration of a first information processing device according to the first embodiment. FIG. [Figure 3] 3A and 3B are explanatory diagrams for explaining an example of a standard shape model and a standard posture model according to the first embodiment. [Figure 4] 10 is a flowchart showing an example of a processing flow in the second information processing device according to the first embodiment. [Figure 5] 6 is a flowchart showing an example of a processing flow in the first information processing device according to the first embodiment. [Figure 6] FIG. 3 is an explanatory diagram for explaining an example of association processing in an association unit according to the first embodiment. [Figure 7] 4 is an explanatory diagram illustrating an example of the configuration of information indicating the association between acoustic data and a posture model by an association unit according to the first embodiment. FIG. [Figure 8] 10 is a flowchart showing an example of a processing flow in a second information processing device according to the second embodiment. [Figure 9] 10 is a flowchart showing an example of a processing flow in a first information processing device according to a second embodiment. [Figure 10] FIG. 10 is an explanatory diagram for explaining an example of association processing in an association unit according to the second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. Note that the configurations shown in the following embodiments are merely examples, and the scope of the present disclosure is not limited to these configurations.
[0010] (Embodiment 1) [composition] An information processing system 1 according to embodiment 1 will be described with reference to Figures 1 to 7. Figure 1 is a block diagram showing an example of the configuration of the information processing system 1 according to embodiment 1. The information processing system 1 includes a plurality of imaging devices 11, a sound collection device 15, a plurality of imaging devices 16, a first information processing device 100, a second information processing device 150, and an output device 19.
[0011] The information processing system 1 is a system for synthesizing virtual viewpoint image data and sound data acquired in different imaging spaces. Sound has various effects on viewers viewing images. For example, in the field of performing arts, stage sounds such as applause and footsteps caused by performers' movements have a significant impact on viewers during the performance. Furthermore, in indoor sports such as sports, the sound in spaces such as gymnasiums is essential for giving viewers a sense of realism during the competition. These sounds are affected by reverberations in the imaging space, such as a stage or gymnasium, and the structure of the walls and floors that make up the imaging space. Therefore, the reverberations and structure of an imaging space, such as a studio where a system for generating virtual viewpoint images is installed, differ from the reverberations and structure of a stage or gymnasium. Therefore, sound collection in an imaging space, such as a studio, cannot acquire the sound data necessary to give viewers a sense of realism or performance. The information processing system 1 of this embodiment is a system for solving this problem.
[0012] Each of the multiple imaging devices 11 is configured with a digital video camera, a digital still camera, or the like, and is installed around a first imaging space (hereinafter referred to as the "first imaging space") such as a studio. Each of the multiple imaging devices 11 images the first imaging space and outputs data of the captured image obtained by the imaging (hereinafter referred to as "captured image data") to the first information processing device 100. Each of the multiple imaging devices 16 is configured with a digital video camera, a digital still camera, or the like, and is installed around a second imaging space (hereinafter referred to as the "second imaging space") such as a stage or a gymnasium. Each of the multiple imaging devices 16 images the second imaging space and outputs captured image data obtained by the imaging to the second information processing device 150.
[0013] The sound collection device 15 is composed of a microphone or the like, and collects sound in the second imaging space, specifically, sound generated when an object present in the second imaging space moves, converts the collected sound into an acoustic signal, and outputs it to the second information processing device 150. Hereinafter, the captured image data output by each of the multiple imaging devices 11 will be collectively referred to as multi-viewpoint image data. Similarly, the captured image data output by each of the multiple imaging devices 16 will be collectively referred to as multi-viewpoint image data (hereinafter referred to as "multi-viewpoint image data").
[0014] The second information processing device 150 acquires the acoustic signal output by the sound collection device 15 and the multi-viewpoint image data output by the plurality of imaging devices 16. The second information processing device 150 outputs the sound indicated by the acquired acoustic signal as acoustic data to the first information processing device 100. Furthermore, the second information processing device 150 generates posture model data corresponding to objects appearing in the captured images based on the plurality of captured image data constituting the acquired multi-viewpoint image data, and outputs the generated posture model data to the first information processing device 100. The second information processing device 150 includes an audio acquisition unit 151, a second image group acquisition unit 152, a second foreground acquisition unit 153, a second model generation unit 154, a model output unit 155, and an audio output unit 156. Details of each unit included in the second information processing device 150 will be described later.
[0015] Here, posture model data refers to data that represents the positions of the joints that make up an object, the connection relationships between the joints, the distances between the joints, the angles of the joints, etc. Hereinafter, it is assumed that the acoustic data includes information indicating the time at which the second information processing device 150 acquired the audio signal, etc. It is also assumed that the imaging devices 16 are time-synchronized with each other, and that the second information processing device 150 and each imaging device 16 are time-synchronized with each other. It is also assumed that the captured image data output by each imaging device 16 includes information indicating the time at which the captured image was captured (hereinafter referred to as "capture time information"). Note that the method of synchronizing time between devices is well known, so a description thereof will be omitted.
[0016] The first information processing device 100 acquires the acoustic data and posture model data output by the second information processing device 150, and the multiple-viewpoint image data output by the multiple imaging devices 16. Based on the multiple captured image data constituting the acquired multiple-viewpoint image data, the first information processing device 100 generates data of a virtual viewpoint image (hereinafter referred to as "virtual viewpoint image data") and posture model data corresponding to an object appearing in the captured image. Hereinafter, the posture model data generated by the first information processing device 100 will be referred to as first posture model data, and the posture model data generated by the second information processing device 150 will be referred to as second posture model data. The first information processing device 100 evaluates the degree of agreement between the first posture model and the second posture model, and generates data including the generated virtual viewpoint image and the acquired acoustics based on the degree of agreement.
[0017] Specifically, the first information processing device 100 generates data including the generated virtual viewpoint image and the acquired sound by synthesizing the generated virtual viewpoint image data with the acquired sound data to generate virtual viewpoint image data with the sound data. Furthermore, the first information processing device 100 outputs the synthesized virtual viewpoint image data with the acquired sound data to the output device 19. The first information processing device 100 includes a first image group acquisition unit 101, a first foreground acquisition unit 102, an image generation unit 103, a first model generation unit 104, a model acquisition unit 105, a sound acquisition unit 106, an association unit 107, an evaluation unit 108, and a data generation unit 109. Details of each unit included in the first information processing device 100 will be described later. Hereinafter, it is assumed that the image capturing devices 11 are synchronized in time with each other. It is also assumed that the captured image data output by each image capturing device 11 includes information indicating the capture time of the captured image (capture time information). Note that a method for synchronizing time between devices is well known, and therefore a description thereof will be omitted.
[0018] The output device 19 has a display output unit constituted by an LCD or the like and an audio output unit constituted by a speaker or the like, and renders the virtual viewpoint image data output by the first information processing device 100 to output the virtual viewpoint image and sound in a viewable manner. The destination to which the first information processing device 100 outputs the synthesized virtual viewpoint image data is not limited to the output device 19, and the first information processing device 100 may output the synthesized virtual viewpoint image data to a storage device not shown in Fig. 1. In this case, the first information processing device 100 writes the synthesized virtual viewpoint image data to the storage device and stores the synthesized virtual viewpoint image data in the storage device.
[0019] The following describes the hardware configuration of the first information processing device 100. The processing of each unit of the first information processing device 100 is performed by hardware such as an ASIC (Application Specific Integrated Circuit) built into the first information processing device 100. The processing may also be performed by hardware such as an FPGA (Field Programmable Gate Array). Furthermore, the processing may also be performed by software using a CPU (Central Processor Unit) or a GPU (Graphic Processor Unit) and a memory.
[0020] The hardware configuration of the first information processing device 100 when each unit included in the first information processing device 100 operates as software will be described with reference to Fig. 2. Fig. 2 is a block diagram showing an example of the hardware configuration of the first information processing device 100 according to embodiment 1. The first information processing device 100 is configured by a computer, and the computer has a CPU 201, a ROM 202, a RAM 203, an auxiliary storage device 204, a display unit 205, an operation unit 206, a communication unit 207, and a bus 208, as shown as an example in Fig. 2.
[0021] The CPU 201 controls the computer using programs or data stored in the ROM 202 or RAM 203, causing the computer to function as each unit included in the first information processing device 100 shown in FIG. 1 . The first information processing device 100 may have one or more dedicated hardware components different from the CPU 201, and at least a portion of the processing performed by the CPU 201 may be executed by the dedicated hardware components. Examples of the dedicated hardware components include an ASIC, an FPGA, and a DSP (digital signal processor). The ROM 202 stores programs and the like that do not require modification. The RAM 203 temporarily stores programs or data supplied from the auxiliary storage device 204, or data and the like that are supplied from the outside via the communication unit 207. The auxiliary storage device 204 is formed, for example, by a hard disk drive or the like, and stores various data such as image data and audio data.
[0022] The display unit 205 is configured with, for example, a liquid crystal display or an LED, and displays a GUI (Graphical User Interface) or the like for the user to operate or view the first information processing device 100. The operation unit 206 is configured with, for example, a keyboard, a mouse, a joystick, a touch panel, or the like, and receives operations by the user to input various instructions to the CPU 201. The CPU 201 also operates as a display control unit that controls the display unit 205 and an operation control unit that controls the operation unit 206.
[0023] The communication unit 207 is used for communication with devices external to the first information processing device 100. For example, when the first information processing device 100 is connected to an external device via a wired connection, a communication cable is connected to the communication unit 207. When the first information processing device 100 has a function for wireless communication with an external device, the communication unit 207 includes an antenna. The bus 208 connects the various units included in the first information processing device 100 to transmit information. In the first embodiment, the display unit 205 and the operation unit 206 are described as being present inside the first information processing device 100, but at least one of the display unit 205 and the operation unit 206 may be present as a separate device outside the first information processing device 100.
[0024] The hardware configuration of the second information processing device 150 will be described. As with the first information processing device 100, the processing of each unit included in the second information processing device 150 is performed by hardware such as an ASIC or FPGA built into the second information processing device 150. The processing may also be performed by software using a CPU or GPU and memory. When each unit included in the second information processing device 150 operates as software, the second information processing device 150 may be configured by a computer shown in FIG. 2, and the computer may function as each unit included in the second information processing device 150 shown in FIG. 1.
[0025] [Processing of each part] The processing of each unit included in the second information processing device 150 will be described. The audio acquisition unit 151 acquires an audio signal output by the sound collection device 15 and digitizes the audio signal by AD conversion to generate audio data. When generating the audio data, the audio acquisition unit 151 generates the audio data so that, for example, information indicating the time at which the audio signal was acquired is included in the audio data. The audio acquisition unit 151 stores the generated audio data in an auxiliary storage device 204 included in the second information processing device 150, and causes the auxiliary storage device 204 to hold the audio data. The audio output unit 156 outputs the audio data generated by the audio acquisition unit 151 to the first information processing device 100. Specifically, the audio output unit 156 reads out the audio data held in the auxiliary storage device 204 from the auxiliary storage device 204, and outputs the read audio data to the first information processing device 100.
[0026] The second image group acquisition unit 152 acquires captured image data output by each of the multiple imaging devices 16. That is, the second image group acquisition unit 152 acquires multi-viewpoint image data from the multiple imaging devices 16. The second foreground acquisition unit 153 extracts an image area corresponding to an object appearing in each of the multiple captured images constituting the multi-viewpoint image acquired by the second image group acquisition unit 152, and acquires the extracted image area as a foreground area. Furthermore, the second foreground acquisition unit 153 generates a foreground image showing the acquired foreground area. Here, the object corresponding to the image area extracted as the foreground area generally refers to a dynamic object (hereinafter referred to as a "moving object") that moves or whose position may change when images are captured from the same direction in a time series. Examples of moving objects include natural persons such as players or referees on the field where a sport is being played, and balls used in ball games. Furthermore, in an event such as a concert or entertainment, the moving object may be a singer, musician, performer, or master of ceremonies. The method of generating a foreground image from a captured image is well known, and therefore a description thereof will be omitted.
[0027] The second model generation unit 154 generates data of a pose model (second pose model) corresponding to an object appearing in a captured image based on foreground regions corresponding to each of the multiple captured images acquired by the second foreground acquisition unit 153. A method by which the second model generation unit 154 generates data of the second pose model will be described. The second model generation unit 154 first acquires data of a standard 3D shape model (hereinafter referred to as a "standard shape model") that imitates the shape of a standard person, and data of a standard pose model (hereinafter referred to as a "standard pose model") that corresponds to the standard shape model. The data of the standard shape model and the standard pose model are stored in advance in, for example, the auxiliary storage device 204 included in the second information processing device 150, and the second model generation unit 154 acquires these data by reading them from the auxiliary storage device 204.
[0028] The standard shape model 301 and the standard pose model 302 will be described with reference to Fig. 3. Fig. 3 is an explanatory diagram for explaining an example of the standard shape model 301 and the standard pose model 302 according to the first embodiment. The standard shape model 301 is a model expressed by a three-dimensional mesh, and the data of the standard shape model 301 includes coordinates indicating the position of each vertex and identification information such as IDs (identifiers) of vertices constituting faces of a triangle, a rectangle, or the like. The standard shape model 301 may be expressed by a set of points called voxels.
[0029] The data of the standard posture model 302 includes information indicating positions in the standard posture model 302 corresponding to joints in the human body, such as the head, shoulders, elbows, wrists, or limbs (hereinafter referred to as "joint information 303"). In addition to the joint information 303, the data of the standard posture model 302 also includes information indicating the connection relationships between the joints in the standard posture model 302 (hereinafter referred to as "connection information 304"). The connection information 304 is, for example, information indicating the distances between the joints in the standard posture model 302. In addition to the joint information 303 and the connection information 304, the data of the standard posture model 302 also includes information indicating angles at the joints in the standard posture model 302 (hereinafter referred to as "angle information"). The angle information is, for example, information indicating the angles formed between line segments connecting adjacent joints in the standard posture model 302.
[0030] The second model generation unit 154 deforms the standard shape model 301 corresponding to the standard pose model 302 shown as an example in FIG. 3 so that it matches the foreground area acquired by the second foreground acquisition unit 153. The second model generation unit 154 estimates the standard pose model 302 that best matches the standard pose model 302 as the pose model corresponding to the object (the second pose model) and generates data of the second pose model. However, the method for generating the second pose model is not limited to the above-described method. For example, the second model generation unit 154 may estimate the two-dimensional pose of the object on the two-dimensional image, and estimate a three-dimensional pose model based on the position, imaging direction, angle of view, etc. of each imaging device 16 to generate data of the second pose model.
[0031] Second model generation unit 154 stores the generated second posture model data in auxiliary storage device 204 included in second information processing device 150, and causes auxiliary storage device 204 to hold the data. When storing the second posture model data in auxiliary storage device 204, second model generation unit 154 stores information indicating the capture time of the captured image used to generate the second posture model (capture time information) in association with the second posture model data in auxiliary storage device 204. Model output unit 155 outputs the second posture model data generated by second model generation unit 154 and the capture time information associated with the second posture model data to first information processing device 100. Specifically, model output unit 155 reads the second posture model data and the capture time information from auxiliary storage device 204, and outputs the read second posture model data and the capture time information to first information processing device 100.
[0032] The processing of each unit included in the first information processing device 100 will be described. Model acquisition unit 105 acquires second posture model data output by model output unit 155 included in the second information processing device 150, and image capture time information associated with the second posture model data. Sound acquisition unit 106 acquires sound data output by sound output unit 156 included in the second information processing device 150. Correspondence unit 107 associates the second posture model data with the sound data. Details of the association performed by correlation unit 107 will be described later using FIGS. 5 and 6.
[0033] The first image group acquisition unit 101 acquires captured image data output by each of the multiple image capture devices 11. That is, the first image group acquisition unit 101 acquires multi-viewpoint image data from the multiple image capture devices 11. The first foreground acquisition unit 102 extracts an image area corresponding to an object appearing in each of the multiple captured images constituting the multi-viewpoint image acquired by the first image group acquisition unit 101, and acquires the extracted image area as a foreground area. Furthermore, the first foreground acquisition unit 102 generates a foreground image showing the acquired foreground area. Here, the object refers to a moving object, such as a natural person existing in a first imaging space such as a studio. The first model generation unit 104 generates data of a posture model (first posture model) corresponding to the object appearing in the captured image, based on the foreground area corresponding to each of the multiple captured images acquired by the first foreground acquisition unit 102. Specifically, for example, first model generation unit 104 generates data of the first posture model using a method similar to the method used by second model generation unit 154 to generate data of the second posture model, as described above.
[0034] Evaluation unit 108 evaluates the degree of match between the first posture model and the second posture model based on the data of the first posture model generated by first model generation unit 104 and the data of the second posture model acquired by model acquisition unit 105. Specifically, evaluation unit 108 evaluates the degree of match by comparing the data of the first posture model with the data of the second posture model, assuming the data of the second posture model as correct data. More specifically, evaluation unit 108 evaluates the degree of match by comparing at least one of the joint information 303, the connection information 304, and the angle information included in the data of the first posture model and the data of the second posture model. For example, evaluation unit 108 determines that the first posture model and the second posture model match when at least one of the difference value between the joint information 303, the difference value between the connection information 304, and the difference value between the angle information is equal to or less than a predetermined threshold.
[0035] Evaluation unit 108 may evaluate the degree of match between the first posture model and the second posture model by combining two or more pieces of information selected from joint information 303, connection information 304, and angle information. Evaluation unit 108 may also compare pieces of information selected from joint information 303 that indicate the positions of predetermined joint sites, such as joints at the head and extremity tips, in standard posture model 302. Similarly, evaluation unit 108 may compare pieces of connection information selected from connection information 304 that indicate the positions of predetermined joint sites in standard posture model 302, or pieces of angle information that indicate the angles of predetermined joint sites in standard posture model 302. The method used by evaluation unit 108 to evaluate the degree of match between the first posture model and the second posture model is not limited to the above.
[0036] The image generation unit 103 generates a virtual viewpoint image based on the multi-viewpoint images acquired by the first image group acquisition unit 101 and the foreground regions in each captured image constituting the multi-viewpoint images acquired by the first foreground acquisition unit 102. Specifically, for example, the image generation unit 103 generates the virtual viewpoint image using the following method. First, the image generation unit 103 generates data of a three-dimensional shape (hereinafter referred to as a "foreground model") corresponding to an object by visual hull analysis using a foreground image showing the foreground region in each captured image. The method of generating three-dimensional shape data using visual hull analysis is well known, so a description thereof will be omitted.
[0037] Next, the image generation unit 103 performs texture mapping on the foreground model using at least one of the multiple captured images that constitute the multi-viewpoint image, thereby coloring the foreground model. Furthermore, the image generation unit 103 performs texture mapping on a background model that forms the background of the foreground model using a prepared background image, thereby coloring the background model. Here, for example, the background image is a captured image obtained by capturing an image of a stadium, a stage, or the like. Finally, the image generation unit 103 generates a virtual viewpoint image by performing rendering according to the position of a viewpoint (hereinafter referred to as a "virtual viewpoint") in a three-dimensional virtual space specified by a user, etc. The method for generating the virtual viewpoint image in the image generation unit 103 is not limited to the above-mentioned method. For example, the image generation unit 103 may generate a virtual viewpoint image by performing projective transformation on the captured image without using three-dimensional shape data.
[0038] Based on the degree of match, which is the evaluation result by evaluation unit 108, data generation unit 109 generates data including the virtual viewpoint image generated by image generation unit 103 and sound indicated by the sound data acquired by sound acquisition unit 106. Specifically, data generation unit 109 combines the data of the virtual viewpoint image generated by image generation unit 103 with the sound data acquired by sound acquisition unit 106 to generate virtual viewpoint image data with sound data. In other words, the data generated by data generation unit 109, including the virtual viewpoint image and sound, is virtual viewpoint image data with sound data. More specifically, first, data generation unit 109 acquires the time corresponding to the second posture model determined by evaluation unit 108 to be a match between the first posture model and the second posture model. Here, the time corresponding to the second posture model is the capture time of the captured image from which the foreground region used to generate the second posture model was acquired.
[0039] Next, data generation unit 109 acquires, from the audio data associated with the data of the second posture model by association unit 107, audio data at a time corresponding to a second posture model determined by evaluation unit 108 to match the first posture model. Finally, data generation unit 109 synthesizes the acquired audio data at the time with the virtual viewpoint image data corresponding to the first posture model determined by evaluation unit 108 to match the first posture model, thereby generating virtual viewpoint image data with audio data. Here, the virtual viewpoint image corresponding to the first posture model is a virtual viewpoint image generated using the same foreground area as the foreground area used when generating the first posture model. Data generation unit 109 outputs the generated virtual viewpoint image data with audio data to an output device. Note that, with regard to virtual viewpoint image data corresponding to a first posture model determined by evaluation unit 108 to not match the first posture model, data generation unit 109 outputs the audio data acquired by audio acquisition unit 106 directly to the output device without synthesizing it.
[0040] [Operation flow] The operation of the second information processing device 150 will be described with reference to FIG. 4. FIG. 4 is a flowchart showing an example of the processing flow in the second information processing device 150 according to the first embodiment. In the description of FIG. 4, the symbol "S" denotes a step. First, in S401, the audio acquisition unit 151 acquires an audio signal output from the sound collection device 15 to generate acoustic data. Furthermore, the second image group acquisition unit 152 acquires captured image data output from each of the multiple imaging devices 16, i.e., multi-viewpoint image data. Next, in S402, the second foreground acquisition unit 153 acquires a foreground region corresponding to an object appearing in the captured image for each of the multiple captured images constituting the multi-viewpoint image acquired in S401, and generates a foreground image showing the foreground region. Next, in S403, the second model generation unit 154 estimates a second posture model and generates data of the second posture model.
[0041] Next, in S404, model output unit 155 outputs data of the second posture model and imaging time information associated with the data of the second posture model to the first information processing device 100. Specifically, for example, when an output instruction is received from the first information processing device 100, model output unit 155 outputs data of the second posture model and the imaging time information to the first information processing device 100. Furthermore, sound output unit 156 outputs sound data to the first information processing device 100. Specifically, for example, when an output instruction is received from the first information processing device 100, sound output unit 156 outputs the sound data to the first information processing device 100. After S404, the second information processing device 150 ends the processing of the flowchart shown in FIG. 4.
[0042] The operation of the first information processing device 100 will be described with reference to Fig. 5. Fig. 5 is a flowchart showing an example of the flow of processing in the first information processing device 100 according to the first embodiment. In the description of Fig. 5, the symbol "S" denotes a step. First, in S501, the model acquisition unit 105 acquires second posture model data and imaging time information associated with the second posture model data from the second information processing device 150. Furthermore, the sound acquisition unit 106 acquires sound data from the second information processing device 150.
[0043] Next, in S502, the association unit 107 analyzes the sound data acquired in S501. Specifically, for example, the association unit 107 analyzes the volume of the sound indicated by the sound data, i.e., the amplitude of the audio signal corresponding to the sound, and searches for a time point in the sound data where the volume reaches a maximum value. The analysis of the volume by the association unit 107 may, for example, analyze the volume of the sound corresponding to a predetermined frequency, such as 48 kilohertz (kHz), among the frequencies in the sound. Furthermore, the association unit 107 cuts out the sound data based on the analysis result in S502. Cutting out the sound data based on the analysis result may, for example, cut out a portion of the sound data acquired in S501 that corresponds to the sound of a natural person, the object, landing when jumping on stage.
[0044] In this case, sounds other than the landing sound generated from the stage floor when the character lands are unnecessary sounds. Therefore, sound data other than the period to be synthesized with the virtual viewpoint image data is deleted and extracted as sound data for synthesis, using a method such as deleting a period of time when the sound volume is below a predetermined threshold before and after the time when the sound volume reaches a maximum value. The method for extracting sound data for synthesis is not limited to the above.
[0045] Next, in S503, association unit 107 associates the audio data for synthesis with the data of the second posture model. Specifically, first, based on the analysis result in step S502, association unit 107 identifies data of the second posture model generated based on a captured image captured at the same time as the time when the volume of the audio data for synthesis reaches its maximum value. Hereinafter, the data of the posture model generated based on a captured image captured at a certain time will be referred to as frame data of a posture model frame. That is, based on the analysis result in step S502, association unit 107 identifies frame data of the second posture model frame generated based on a captured image captured at the same time as the time when the volume of the audio data for synthesis reaches its maximum value. Next, association unit 107 associates the time when the volume of the audio data for synthesis reaches its maximum value with the frame data of the identified second posture model frame. In this way, association unit 107 associates the audio data for synthesis with the data of the second posture model.
[0046] The association process in association unit 107 will be described with reference to FIG. 6. FIG. 6 is an explanatory diagram for explaining an example of the association process in association unit 107 according to the first embodiment. In FIG. 6, second posture model frames 601a to 601e arranged in time series are shown in the upper part, and acoustic data 602a to 602c shown as time-series audio signals are shown in the lower part. Note that the lower part of FIG. 6 shows, as an example, acoustic data corresponding to 48 kHz from the acoustic data as an audio signal. Here, acoustic data 602b represents acoustic data for synthesis extracted by the extraction process in S502. Furthermore, acoustic data 602a and 602c represent audio data deleted by the extraction process in S502 from the audio data acquired in S501. Note that in FIG. 6, the horizontal axis represents the time axis, and in the lower part of FIG. 6, the vertical axis represents the amplitude of the audio signal.
[0047] The sound data 602b is associated with one of the second posture model frames 601a to 601e as sound data for synthesis with a virtual viewpoint image. The analysis in S502 identifies the time when the volume of the sound data 602b is at its maximum, i.e., the time corresponding to the time when the sound data 602b has its maximum amplitude. After this identification, a second posture model frame generated based on an image captured at the same time is identified. In the example shown in FIG. 6, the time corresponding to the second posture model frame 601d coincides with the time when the sound data 602b has its maximum amplitude, and therefore, the association unit 107 associates the sound signal 602b with the second posture model frame 601d.
[0048] However, even if the second posture model frame and the acoustic data are time-synchronized, the frame rate of the second posture model frame and the sampling rate of the acoustic data may differ. In such cases, there may be no second posture model frame at the same time as the time when the acoustic data has the maximum amplitude. Therefore, in such cases, it is sufficient to identify the second posture model frame at the time closest to the time when the acoustic data has the maximum amplitude, and associate the identified second posture model frame with the time when the acoustic data has the maximum amplitude. The method for associating the acoustic data with the second posture model is not limited to the above, and any method that can synchronize the acoustic data with the second posture model will suffice.
[0049] The configuration of information indicating the association between audio data and a posture model by association unit 107 will be described with reference to FIG. 7. FIG. 7 is an explanatory diagram illustrating an example of the configuration of information indicating the association between audio data and a posture model by association unit 107 according to the first embodiment. As shown in FIG. 7, for example, pattern numbers are assigned according to the number of patterns of audio data to be synthesized with a virtual viewpoint image. Here, it is assumed that the patterns are set to audio data to be synthesized with a virtual viewpoint image for each performance of a stage performance or for each shooting scene. The number of audio data shown in FIG. 7 is input as the number of pattern numbers, i.e., the number of audio data for synthesis extracted in S502. Each piece of audio information shown in FIG. 7 stores audio data for synthesis, and the posture estimation model data shown in FIG. 7 stores frame data of a second posture model frame at a time corresponding to the time when the audio data for synthesis has the maximum amplitude.
[0050] After S503, in S511, the first image group acquisition unit 101 acquires captured image data output by each of the multiple image capture devices 11, i.e., multi-viewpoint image data. Next, in S512, the first foreground acquisition unit 102 acquires a foreground region corresponding to an object appearing in each of the multiple captured images constituting the multi-viewpoint image acquired in S511, and generates a foreground image showing the foreground region. Next, in S513, the image generation unit 103 generates a virtual viewpoint image. Next, in S514, the first model generation unit 104 estimates a first posture model and generates data of the first posture model.
[0051] Next, in S515, evaluation unit 108 evaluates the degree of match between the first posture model and the second posture model, and determines whether the first posture model and the second posture model match. If it is determined that they match in S515, then in S516, data generation unit 109 combines the virtual viewpoint image data generated in S513 with the audio data for synthesis associated with the second posture model data in S503. Data generation unit 109 then outputs the combined virtual viewpoint image data to the output device. If it is determined in S515 that the first posture model and the second posture model do not match, data generation unit 109 outputs the virtual viewpoint image data generated in S513 to the output device as is.
[0052] The first information processing device 100 determines whether or not a termination condition is satisfied in S520. Here, the termination condition is, for example, when an operation signal indicating a termination instruction is received from the user. The first information processing device 100 repeatedly executes the processes from S511 to S516 until the termination condition is satisfied, and terminates the process of the flowchart shown in FIG. 5 when the termination condition is satisfied.
[0053] As described above, the first information processing device 100 can efficiently combine virtual viewpoint image data and sound data acquired in different imaging spaces. As a result, it becomes possible to efficiently combine sound that cannot be reproduced in a studio or the like with a virtual viewpoint image, thereby reducing the workload for combining sound data with virtual viewpoint image data.
[0054] In the first embodiment, the first information processing device 100 has been described as including the first image group acquisition unit 101, the first foreground acquisition unit 102, the image generation unit 103, and the first model generation unit 104; however, the present invention is not limited to this. For example, if the first information processing device 100 includes a first posture model acquisition unit illustrated in FIG. 1 that acquires data of a first posture model generated by a device different from the first information processing device 100, the first information processing device 100 does not need to include the first model generation unit 104. Also, for example, if the first information processing device 100 includes an image acquisition unit illustrated in FIG. 1 that acquires virtual viewpoint images generated by a device different from the first information processing device 100, the first information processing device 100 does not need to include the image generation unit 103. Also, the first information processing device 100 may include each unit included in the second information processing device 150. That is, the first information processing device 100 may have the functions included in the second information processing device 150. When the first information processing device 100 has all the components that the second information processing device 150 has, the information processing system 1 does not need to have the second information processing device 150.
[0055] (Embodiment 2) An information processing system 1 according to embodiment 2 will be described with reference to Figures 8 to 10. The configuration of the information processing system 1 according to embodiment 2 is similar to the configuration of the information processing system 1 according to embodiment 1, which is shown as an example in Figure 1. That is, the information processing system 1 includes a plurality of imaging devices 11, a sound collection device 15, a plurality of imaging devices 16, a first information processing device 100, a second information processing device 150, and an output device 19.
[0056] The information processing system 1 according to the first embodiment is as follows. First, the second information processing device 150 generates second posture model data and sound data in advance based on captured image data and audio signals obtained by capturing images and collecting sounds in a second imaging space such as a stage. Next, the first information processing device 100 generates first posture model data and virtual viewpoint image data based on captured image data obtained by capturing images in a first imaging space such as a studio. Furthermore, the first information processing device 100 evaluates the degree of agreement between the first posture model data and the second posture model data, combines the virtual viewpoint image data and sound data based on the degree of agreement, and outputs the combined virtual viewpoint image data.
[0057] In contrast, an information processing system 1 according to a second embodiment (hereinafter simply referred to as "information processing system 1") is as follows. First, the second information processing device 150 generates second posture model data and sound data in advance based on captured image data and audio signals obtained by capturing and collecting sound in a second imaging space such as a stage. Physical information is also generated in advance by analyzing the second posture model data. The physical information will be described later. Next, the first information processing device 100 generates first posture model data and virtual viewpoint image data based on captured image data obtained by capturing image data in a first imaging space such as a studio. The physical information is also generated by analyzing the first posture model data. Furthermore, the first information processing device 100 evaluates the degree of agreement between the first posture model data and the corresponding physical information, and the second posture model data and the corresponding physical information. Finally, the first information processing device 100 combines the virtual viewpoint image data and the sound data based on the degree of agreement, and outputs the combined virtual viewpoint image data.
[0058] [composition] The functional block configuration of a first information processing device 100 according to the second embodiment (hereinafter simply referred to as "first information processing device 100") is similar to the functional block included in the first information processing device 100 according to the first embodiment shown as an example in FIG. 1, and therefore a description thereof will be omitted. That is, the first information processing device 100 includes: A first image group acquisition unit 101, a first foreground acquisition unit 102, an image generation unit 103, a first model generation unit 104, a model acquisition unit 105, an acoustic acquisition unit 106, an association unit 107, an evaluation unit 108, and a data generation unit 109. Furthermore, the functional block configuration of a second information processing device 150 according to the second embodiment (hereinafter simply referred to as "second information processing device 150") is similar to the functional block included in the second information processing device 150 according to the first embodiment shown as an example in FIG. 1, and therefore a description thereof will be omitted. That is, the second information processing device 150 includes a sound acquisition unit 151, a second image group acquisition unit 152, a second foreground acquisition unit 153, a second model generation unit 154, a model output unit 155, and a sound output unit 156. Furthermore, the hardware configurations of the first information processing device 100 and the second information processing device 150 are similar to those of the first information processing device 100 and the second information processing device 150 according to the first embodiment, and therefore description thereof will be omitted.
[0059] [Processing of each part] The following describes the differences between the information processing system 1 and the information processing system 1 according to the first embodiment. First, a description will be given of the processing of each unit included in the second information processing device 150. The audio acquisition unit 151, the second image group acquisition unit 152, the second foreground acquisition unit 153, and the sound output unit 156 are the same as the audio acquisition unit 151, the second image group acquisition unit 152, the second foreground acquisition unit 153, and the sound output unit 156 according to the first embodiment, and therefore descriptions thereof will be omitted.
[0060] In addition to generating second posture model data, the second model generation unit 154 also has a function of analyzing the generated second posture model data and generating physical information corresponding to the second posture model data (hereinafter referred to as "second physical information"). Here, the physical information is information indicating the velocity or acceleration of a joint corresponding to an object in a second posture model frame. Specifically, the second model generation unit 154 generates the second physical information by calculating the velocity or acceleration of the joint based on frame data of a plurality of second posture model frames. The second model generation unit 154 stores the generated second physical information in the auxiliary storage device 204 of the second information processing device 150, in addition to the second posture model data and information indicating the capture time of the captured image used to generate the second posture model (capture time information). Specifically, the second model generation unit 154 associates the second physical information with the frame data of the second posture model frame and stores the second physical information in the auxiliary storage device 204 of the second information processing device 150.
[0061] Model output unit 155 outputs the second posture model data, the imaging time information associated with the second posture model data, and the second physical information generated by second model generation unit 154 to first information processing device 100. Specifically, model output unit 155 reads the second posture model data, imaging time information, and second physical information from auxiliary storage device 204, and outputs the read second posture model data, imaging time information, and second physical information to first information processing device 100.
[0062] Next, the processing of each unit included in the first information processing device 100 will be described. The first image group acquisition unit 101, first foreground acquisition unit 102, image generation unit 103, sound acquisition unit 106, and data generation unit 109 are the same as the corresponding units according to the first embodiment, and therefore description thereof will be omitted. The first model generation unit 104 has a function of generating first posture model data, as well as a function of analyzing the generated first posture model data and generating physical information corresponding to the first posture model data (hereinafter referred to as "first physical information"). Specifically, the first model generation unit 104 generates the first physical information by calculating the velocity or acceleration of the joint based on frame data of a plurality of first posture model frames. The generated first physical information is associated with the frame data of the first posture model frames.
[0063] Model acquisition unit 105 acquires second physical information in addition to second posture model data output by model output unit 155 of second information processing device 150 and image capture time information associated with the second posture model data. The acquired second physical information is associated with frame data of a second posture model frame. Corresponding unit 107 has a function of extracting acoustic data for synthesis from acoustic data and a function of associating the acoustic data for synthesis, frame data of the second posture model frame, and second physical information corresponding to the frame data of the second posture model frame. Of the functions of correlating unit 107, the function of extracting acoustic data for synthesis from acoustic data was explained in the first embodiment, so its explanation is omitted here. Furthermore, a method for associating the acoustic data for synthesis with frame data of the second posture model frame was explained in the first embodiment, so its explanation is omitted here. Furthermore, since the second physical information is associated with frame data of the second posture model frame, its explanation is omitted here.
[0064] The evaluation unit 108 has a function of evaluating the degree of match between the first posture model and the second posture model based on the data of the first posture model and the data of the second posture model. Specifically, the evaluation unit 108 has a function of evaluating the degree of match by comparing at least one of the joint information 303, the connection information 304, and the angle information included in the data of the first posture model and the data of the second posture model. In addition to this function, the evaluation unit 108 also has a function of evaluating the degree of match between the first posture model and the second posture model based on the first physical information and the second physical information. Specifically, the evaluation unit 108 compares information indicating the velocity or acceleration of a joint corresponding to an object in the first posture model with information indicating the velocity or acceleration of a joint corresponding to the object in the second posture model that corresponds to the joint. For example, when at least one of the difference between the velocities and the difference between the accelerations is equal to or less than a predetermined threshold, the evaluation unit 108 determines that the first posture model and the second posture model match.
[0065] More specifically, for example, first, evaluation unit 108 evaluates the degree of match between the first posture model and the second posture model based on the data of the first posture model and the data of the second posture model. Next, when it is determined that the first posture model and the second posture model match based on the data of the first posture model and the data of the second posture model, evaluation unit 108 evaluates the degree of match between the first posture model and the second posture model based on the first physical information and the second physical information. Such a step-by-step evaluation can improve the accuracy of the evaluation of the degree of match between the first posture model and the second posture model. Note that when it is determined that the first posture model and the second posture model match based on the data of the first posture model and the second posture model, evaluation unit 108 may instruct first model generation unit 104 and model acquisition unit 105 to generate physical information.
[0066] [Operation flow] The operation of the second information processing device 150 will be described with reference to Fig. 8. Fig. 8 is a flowchart showing an example of a processing flow in the second information processing device 150 according to the second embodiment. In the description of Fig. 8, the symbol "S" denotes a step. In Fig. 8, the same symbols as in Fig. 4 will not be described. First, the second information processing device 150 executes the processes from S401 to S403.
[0067] After S403, in S804, second model generation unit 154 generates second physical information. Next, in S805, model output unit 155 outputs the second posture model data, image capture time information associated with the second posture model data, and the second physical information to the first information processing device 100. Specifically, for example, when an output instruction is received from the first information processing device 100, model output unit 155 outputs the second posture model data, the image capture time information, and the second physical information to the first information processing device 100. Furthermore, audio output unit 156 outputs audio data to the first information processing device 100. Specifically, for example, when an output instruction is received from the first information processing device 100, audio output unit 156 outputs the audio data to the first information processing device 100. After S805, the second information processing device 150 ends the processing of the flowchart shown in FIG. 8.
[0068] The operation of the first information processing device 100 will be described with reference to Fig. 9. Fig. 9 is a flowchart showing an example of a processing flow in the first information processing device 100 according to the second embodiment. In the description of Fig. 9, the symbol "S" denotes a step. In Fig. 9, the same symbols as in Fig. 5 will not be described. First, the first information processing device 100 executes the processes of S501 and S502.
[0069] After S502, in S903, association unit 107 associates the audio data for synthesis, the second posture model data, and the second physical information with each other. Specifically, first, based on the analysis result of step S502, association unit 107 identifies second posture model data generated based on captured images captured at the same time as the time when the volume of the audio data for synthesis reaches its maximum value. That is, based on the analysis result of step S502, association unit 107 identifies frame data of a second posture model frame generated based on captured images captured at the same time as the time when the volume of the audio data for synthesis reaches its maximum value. Next, association unit 107 associates a function of extracting audio data for synthesis from audio data with the audio data for synthesis, the frame data of the second posture model frame, and the second physical information corresponding to the frame data of the second posture model frame. In this way, association unit 107 associates the audio data for synthesis, the second posture model data, and the second physical information with each other.
[0070] Referring to FIG. 10, the matching process in matching unit 107, particularly the process of matching between audio data for synthesis and frame data of a plurality of second posture model frames, will be described. FIG. 10 is an explanatory diagram for explaining an example of the matching process in matching unit 107 according to the second embodiment. In FIG. 10, the upper part shows second posture model frames 1001a to 1001e arranged in time series, and the lower part shows audio data 1002a to 1002c represented by time-series audio signals. Note that the lower part of FIG. 10 shows, as an example, audio data corresponding to 48 kHz as an audio signal. Here, audio data 1002b represents audio data for synthesis extracted by the extraction process in S502. Furthermore, audio data 1002a and 1002c represent audio data deleted by the extraction process in S502 from the audio data acquired in S501. In FIG. 10, the horizontal axis represents time, and in the lower part of FIG. 10, the vertical axis represents the amplitude of the audio signal.
[0071] The sound data 1002b is associated with one of the second posture model frames 1001a to 1001e as sound data for synthesis with a virtual viewpoint image. The analysis in S502 identifies the time when the sound volume in the sound data 1002b is at its maximum, i.e., the time corresponding to the time when the sound data 1002b has its maximum amplitude. After this identification, a second posture model frame generated based on an image captured at the same time is identified. In the example shown in FIG. 10 , the time corresponding to the second posture model frame 1001d coincides with the time when the sound data 1002b has its maximum amplitude. Therefore, the association unit 107 associates the sound signal 1002b with the second posture model frame 1001d. If there is no second posture model frame corresponding to the time when the sound data has its maximum amplitude, the association unit 107 first identifies a second posture model frame corresponding to the time closest to the time when the sound data has its maximum amplitude, as in the first embodiment. Next, the time corresponding to the point in time when the acoustic data has the maximum amplitude may be associated with the identified second posture model frame.
[0072] After associating acoustic signal 1002b with second posture model frame 1001d, acoustic signal 1002b is associated with one or more second posture model frames included within a predetermined period before and after the time of second posture model frame 1001d. It is not necessary to associate acoustic signal 1002b with all second posture model frames included in the period. For example, acoustic signal 1002b may be associated with only those second posture model frames included in the period whose time intervals are a predetermined interval. The method of determining second posture model frames adjacent to second posture model frame 1001d to which acoustic signal 1002b is associated is not limited to the above. Furthermore, the method of associating acoustic data with a second posture model is not limited to the above, and any method capable of synchronizing audio and the second posture model may be used.
[0073] After S503, first information processing device 100 executes the processes from S511 to S514. After S514, at S911, first model generation unit 104 generates first physical information. Next, at S912, evaluation unit 108 evaluates the degree of match between the frame data of the first posture model frame to be evaluated and the frame data of the second posture model frame, and determines whether the first posture model frame and the second posture model frame match. If it is determined that there is a match at S515, evaluation unit 108 executes the process of S913. At S913, evaluation unit 108 determines whether the frame data and first physical information of the first posture model frame and the frame data and second physical information of the second posture model frame after the first posture model frame to be evaluated match.
[0074] If it is determined in S913 that there is a match, then in S516 the data generation unit 109 combines the virtual viewpoint image data generated in S513 with the audio data for synthesis associated with the second posture model data in S503. Thereafter, the data generation unit 109 outputs the combined virtual viewpoint image data to the output device. If it is determined in S912 or S913 that there is no match, the data generation unit 109 outputs the virtual viewpoint image data generated in S513 as is to the output device. In S520, the first information processing device 100 determines whether or not a termination condition is satisfied. The first information processing device 100 repeatedly executes the processes from S511 to S516 until the termination condition is satisfied, and terminates the processing of the flowchart shown in FIG. 9 when the termination condition is satisfied.
[0075] As described above, the first information processing device 100 can efficiently synthesize virtual viewpoint image data and audio data acquired in different imaging spaces. As a result, audio that cannot be reproduced in a studio or the like can be efficiently synthesized with a virtual viewpoint image, thereby reducing the workload for synthesizing audio data and virtual viewpoint image data. Furthermore, the first information processing device 100 evaluates the degree of correspondence between the first physical information corresponding to the first posture model and the second physical information corresponding to the second posture model, in addition to the degree of correspondence between the data of the first posture model and the data of the second posture model. The first information processing device 100 can improve the accuracy of automatic synthesis of virtual viewpoint image data and audio data acquired in different imaging spaces. As a result, it is possible to prevent audio data from being synthesized at an incorrect time from being synthesized with virtual viewpoint image data.
[0076] (Other embodiments) The present disclosure can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.
[0077] It should be noted that within the scope of the present disclosure, the embodiments may be freely combined, any component of each embodiment may be modified, or any component of each embodiment may be omitted. [Explanation of symbols]
[0078] 100 Information processing device 103 Image generation unit 104 First model generation unit 105 Model Acquisition Department 106 Acoustic acquisition section 108 Evaluation Department 109 Data Generation Unit
Claims
1. a first model acquisition means for acquiring data of a first posture model generated based on first multi-viewpoint images obtained from a plurality of imaging devices that capture an image of a first imaging space; an image acquisition means for acquiring data of a virtual viewpoint image generated based on the first multi-viewpoint image; a second model acquisition means for acquiring data of a second posture model generated based on second multi-viewpoint images obtained from a plurality of imaging devices that capture a second imaging space different from the first imaging space; a sound acquisition means for acquiring data of sound collected when the second multi-viewpoint image is captured in the second imaging space; evaluation means for evaluating a degree of agreement between the first posture model and the second posture model; a data generating means for generating data including the virtual viewpoint image and the sound based on the degree of coincidence; Having An information processing device characterized by:
2. a correlation means for analyzing the acoustic data and correlating the second posture model data with the analyzed acoustic data; and the data generating means identifies the first posture model corresponding to the second posture model based on the degree of coincidence, and generates data correlating the virtual viewpoint image corresponding to the identified first posture model with the sound data associated with data of the second posture model identified as corresponding to the first posture model.
2. The information processing device according to claim 1,
3. the association means analyzes the sound data, extracts sound data for synthesis from the sound data based on the analysis result, and associates the extracted sound data for synthesis with data of the second posture model corresponding to a period of the sound data for synthesis; The data generating means generates data including the virtual viewpoint image and the sound for synthesis.
3. The information processing device according to claim 2, wherein:
4. The associating means extracts the sound data to be synthesized based on the volume of the sound data.
4. The information processing device according to claim 3,
5. The association means extracts the sound data for synthesis based on the volume of the sound data corresponding to a predetermined frequency.
5. The information processing device according to claim 4,
6. the evaluation means evaluates the degree of coincidence between the first posture model and the second posture model based on at least one of information indicating positions of joints in the first posture model and information indicating positions of joints in the second posture model, and information indicating connection relationships between joints in the first posture model and information indicating connection relationships between joints in the second posture model.
6. The information processing device according to claim 1, wherein:
7. a physical information acquiring means for acquiring first physical information that is physical information of the first posture model and second physical information that is physical information of the second posture model; and the evaluating means evaluates the degree of coincidence between the first posture model and the second posture model based on the first physical information and the second physical information.
7. The information processing device according to claim 1, wherein:
8. The first physical information and the second physical information are information indicating at least one of the velocity and acceleration of a joint part in the first posture model and the second posture model, and the angle of a part in the first posture model and the second posture model.
8. The information processing device according to claim 7,
9. The first posture model is a three-dimensional model that indicates the joints and the connection relationships between the joints in an object that exists in a first imaging space, and the second posture model is a three-dimensional model that indicates the joints and the connection relationships between the joints in an object that exists in a second imaging space.
9. The information processing device according to claim 1, wherein:
10. a first image group acquisition means for acquiring the first multi-viewpoint images; a first foreground acquisition means for acquiring, for each of a plurality of captured images constituting the first multi-viewpoint image, an image region corresponding to an object in the captured image as a foreground region; a first model generation means for generating the first pose model corresponding to the object based on the foreground region acquired for each of a plurality of captured images; an image generating means for generating the virtual viewpoint image based on the first multi-viewpoint image and the foreground region acquired for each of the plurality of captured images constituting the first multi-viewpoint image; and the first model acquisition means acquires data of the first posture model generated by the first model generation means; The image acquisition means acquires data of the virtual viewpoint image generated by the image generation means.
10. The information processing device according to claim 1, wherein:
11. a second image group acquisition means for acquiring the second multi-viewpoint images; a second foreground acquisition means for acquiring, for each of a plurality of captured images constituting the second multi-viewpoint image, an image region corresponding to an object in the captured image as a foreground region; a second model generation means for generating the second pose model corresponding to the object based on the foreground region acquired for each of a plurality of captured images; a sound acquisition means for acquiring a signal of a sound collected by a sound collection device installed in the second imaging space, and digitizing the sound signal to acquire the acoustic data; and the second model acquisition means acquires data of the second posture model generated by the second model generation means; The sound acquisition means acquires data of the sound acquired by the voice acquisition means. The information processing device according to any one of claims 1 to 10,
12. a first information processing device that generates data of a first posture model and data of a virtual viewpoint image based on first multi-viewpoint images obtained from a plurality of imaging devices that capture an image of a first imaging space; a second information processing device that generates data of a second posture model based on second multi-viewpoint images obtained from a plurality of imaging devices that capture an image of a second imaging space different from the first imaging space, and generates sound data based on a sound signal obtained by sound collection by a sound collection device installed in the second imaging space; and the first information processing device acquires the second posture model and the sound data generated by the second information processing device, and generates data including the virtual viewpoint image and the sound based on a degree of coincidence between the first posture model and the second posture model. An information processing system characterized by:
13. a first model acquisition step of acquiring data of a first posture model generated based on first multi-viewpoint images obtained from a plurality of imaging devices that capture an image of a first imaging space; an image acquisition step of acquiring data of a virtual viewpoint image generated based on the first multi-viewpoint image; a second model acquisition step of acquiring data of a second posture model generated based on second multi-viewpoint images obtained from a plurality of imaging devices that capture a second imaging space different from the first imaging space; a sound acquiring step of acquiring data of sound collected when the second multi-viewpoint image is captured in the second imaging space; an evaluation step of evaluating a degree of agreement between the first posture model and the second posture model; a data generating step of generating data including the virtual viewpoint image and the sound based on the degree of coincidence; Having An information processing method comprising:
14. A program for causing a computer to operate as the information processing device according to any one of claims 1 to 11.
Citation Information
Patent Citations
Object attribute estimation device and video plotting device
JP2013120556A
Information processing unit, control method and program
JP2017211827A
Image processing system, image processing method and program
JP2019003325A
Game program and game system
JP2020195672A
Apparatus for providing sound effects according to an image and method thereof
US20050201565A1