Image generation device, image generation method, and image generation program
The image generation device synthesizes local audience video with remote viewer actions to create virtual audience frames, addressing the lack of interaction in remote viewing services and enhancing emotional engagement.
Patent Information
- Application Number
- JP2024536704
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-07-28
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2042-07-28
AI Technical Summary
Existing remote viewing services lack the ability to distribute video that reflects the actions of remote audience members, failing to recreate the sense of unity and interaction experienced at live events.
An image generation device that acquires audience seat video frames and action information of remote viewers, synthesizing them to create virtual audience video frames that are combined with event content frames, generating viewing videos that reflect the actions of remote audience members.
Enables remote viewers to experience a sense of unity and excitement by interacting with virtual audience members, enhancing the emotional engagement of remote viewing experiences.
Smart Images

Figure 0007810266000001 
Figure 0007810266000002 
Figure 0007810266000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image generation device, an image generation method, and an image generation program. [Background technology]
[0002] In remote viewing services that allow viewers to watch events such as live music concerts and sporting events remotely, displaying other spectators is an important element in recreating the emotional experiences such as the sense of unity and excitement felt when watching at the venue.
[0003] Existing remote viewing services have devised ways to display audiences, such as including footage of the audience seats in the streamed video. However, streaming the footage of the audience seats directly raises privacy concerns, such as preventing audience members' faces from appearing in the footage.
[0004] Possible ways to address this issue include virtualizing audiences, such as capturing users' movements through motion capture and representing them as avatars in a virtual space (Non-Patent Document 1), artificially creating audience movements and applying them to avatars, or representing audiences with physical penlights (Non-Patent Document 2). [Prior art documents] [Non-patent literature]
[0005] [Non-Patent Document 1] Tatsuyoshi KANEKO, Hiroyuki TARUMI, Yuki KUBOCHI, Ryota YAMAGUCHI, Keiya KATAOKA, Daiki YAMASHITA, Tomoki NAKAI, "Supporting the Sense of Unity between Remote Audiences in VR-Based Remote Live Music Support System KSA2", 2018 IEEE International Conference on Artificial Intelligence and Virtual Reality (AIVR) [Non-patent document 2] Masaki Naka, Mio Ayabe, Ren Ueda, and Motofumi Shikida: Proposal of a method to support mutual communication by visualizing excitement using penlights during online music concerts, IPSJ Technical Report, Vol. 2022-GN-115, No. 39, pp.1-7, 2022 Summary of the Invention [Problem to be solved by the invention]
[0006] Interaction between audience members is important to achieve the same sense of unity as at a live venue. Current remote viewing services only provide one-way video distribution, and do not distribute video in which the actions of remote audience members are reflected in the virtual audience.
[0007] An object of the present invention is to provide an image generation device, an image generation method, and an image generation program for generating viewing images including virtual audience images that reflect the actions of remote audience members. [Means for solving the problem]
[0008] One aspect of the present invention is a video generation device. The video generation device has an audience video acquisition unit, a content video acquisition unit, an action acquisition unit, a video generation unit, and a video output unit. The audience video acquisition unit acquires audience seat video frames of local audience members at the event venue. The content video acquisition unit acquires content video frames of the event. The action acquisition unit acquires action information of remote audience members watching the event remotely. The video generation unit generates virtual audience video frames based on the audience seat video frames and the action information of the remote audience members, and synthesizes the content video frames with the virtual audience video frames to generate viewing videos for the remote audience members. The video output unit outputs the viewing videos to a display device in the viewing environment of the remote audience members.
[0009] One aspect of the present invention is a video generation method that includes acquiring audience seating video frames of local audience members at an event venue, acquiring content video frames of the event, acquiring action information of remote audience members watching the event remotely, generating virtual audience video frames based on the audience seating video frames and the action information of the remote audience members, synthesizing the content video frames with the virtual audience video frames to generate viewing videos for the remote audience members, and outputting the viewing videos to a display device in the viewing environment of the remote audience members.
[0010] One aspect of the present invention is an image generation program that causes a processor of a computer to execute the functions of each component of the image generation device. [Effects of the Invention]
[0011] According to the present invention, there are provided an image generation device, an image generation method, and an image generation program for generating viewing images including virtual audience images that reflect the actions of remote audience members. [Brief explanation of the drawings]
[0012] [Figure 1] FIG. 1 is a block diagram illustrating an example of a functional configuration of an image generating device according to an embodiment. [Figure 2] FIG. 2 is a block diagram illustrating an example of a hardware configuration of the image generation device according to the embodiment. [Figure 3] FIG. 3 is a flowchart illustrating an example of the procedure and content of the image generation process executed by the image generation device according to the embodiment. [Figure 4] FIG. 4 is a diagram for explaining the process executed by the image generating unit of the image generating device according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0013] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.
[0014] [Function Configuration] An image generation device 10 according to one embodiment of the present invention will be described with reference to Fig. 1. Fig. 1 is a block diagram showing the functional configuration of the image generation device 10 according to one embodiment of the present invention. The image generation device 10 is a device that generates video images to be provided to remote spectators who remotely watch events such as live music concerts and sporting events.
[0015] As shown in Figure 1, an image generation device 10 according to one embodiment of the present invention has an audience image acquisition unit 11, a content image acquisition unit 12, an action acquisition unit 13, an image generation unit 14, and an image output unit 15.
[0016] The spectator video acquisition unit 11 acquires video of local spectators at the event venue via the network NW. The spectator video acquisition unit 11 acquires local spectator information based on the video of the local spectators. The spectator video acquisition unit 11 acquires spectator seating video frames based on the local spectator information. The spectator video acquisition unit 11 outputs the spectator seating video frames to the video generation unit 14.
[0017] The content video acquisition unit 12 acquires video of the event via the network NW. The content video acquisition unit 12 acquires content video frames based on the video of the event. The content video frames are video frames that do not include on-site spectators. For example, if the event is a live music concert, the content video frames are video frames of the artist, and if the event is a sporting event, the content video frames are video frames of sports scenes. For convenience, the following description will be given assuming that the event is a live music concert. The content video acquisition unit 12 outputs the content video frames to the video generation unit 14.
[0018] The action acquisition unit 13 acquires video of the remote spectators from the camera device 60. The remote spectators are users who receive the remote viewing service. In other words, they are users of the remote viewing service who watch an event remotely. The camera device 60 is installed near the remote spectators. The camera device 60 is generally a camera. The remote spectators operate the camera device 60 to capture themselves. The action acquisition unit 13 acquires action information of the remote spectators based on the video of the remote spectators. The action acquisition unit 13 outputs the action information of the remote spectators to the video generation unit 14.
[0019] The action information of the remote spectators is, for example, the waving of a penlight, the color of the penlight (and the change in color). Here, the action information of the remote spectators is described as the color of the penlight. However, the action information of the remote spectators is not limited to this, and may be information on other movements, etc.
[0020] Video generation unit 14 generates virtual audience video frames based on audience seat video frames received from audience video acquisition unit 11 and action information of remote audience members received from motion acquisition unit 13. The virtual audience video frames are video frames that virtualize the actions of local audience members whose actions are highly similar to those of the remote audience members. Video generation unit 14 synthesizes the virtual audience video frames with the content video frames received from content video acquisition unit 12 to generate viewing video frames of the remote audience members. Video generation unit 14 outputs the viewing video frames to video output unit 15.
[0021] The video output unit 15 outputs the viewing video received from the video generation unit 14 to the display device 70. The display device 70 is in the viewing environment of the remote audience. In other words, the display device 70 is installed near the remote audience. The display device 70 is, for example, a monitor or an HMD (head-mounted display). The remote audience views the viewing video frames through the display device 70.
[0022] [Hardware configuration] Next, the hardware configuration of an image generation device 10 according to an embodiment of the present invention will be described with reference to Fig. 2. Fig. 2 is a block diagram showing the hardware configuration of an image generation device 10 according to an embodiment of the present invention.
[0023] The image generation device 10 is configured by a computer. For example, the image generation device 10 is configured by a personal computer. Here, an example will be described in which the image generation device 10 is configured by a personal computer that can be operated by a remote spectator. However, the image generation device 10 is not limited to this, and may be configured by, for example, a server computer.
[0024] 2, the image generation device 10 has a hardware processor 20, a program storage unit 31, a data storage unit 32, a communication interface 41, and an input / output interface 42. The hardware processor 20, the program storage unit 31, the data storage unit 32, the communication interface 41, and the input / output interface 42 are connected to one another via a bus 50, and can exchange information among them.
[0025] The hardware processor 20 is, for example, a CPU (Central Processing Unit). The hardware processor 20 executes programs, performs data arithmetic processing, etc. The hardware processor 20 controls a program storage unit 31, a data storage unit 32, a communication interface 41, and an input / output interface 42. The hardware processor 20 also controls an image capture device 60 and a display device 70 connected to the input / output interface 42, as will be described later.
[0026] The program storage unit 31 is configured by combining a non-volatile memory such as a hard disk drive (HDD) or a solid state drive (SSD) that can be written to and read from at any time, which is a non-transitory tangible storage medium, with a non-volatile memory such as a read only memory (ROM). The program storage unit 31 stores programs to be executed by the hardware processor 20 so that the image generation device 10 can perform each process.
[0027] The data storage unit 32 is configured by combining, as a tangible storage medium, the above-mentioned nonvolatile memory with a volatile memory such as a RAM (Random Access Memory). The data storage unit 32 temporarily stores data required for the processing executed by the hardware processor 20.
[0028] The communication interface 41 includes, for example, a wireless communication interface unit, and enables transmission and reception of information between the hardware processor 20 and the communication network NW. As the wireless interface, for example, an interface that adopts a low-power wireless data communication standard such as a wireless LAN (Local Area Network) can be used.
[0029] The input / output interface 42 is connected to the image capturing device 60 and the display device 70. The input / output interface 42 enables transmission and reception of information between the hardware processor 20 and the like and the image capturing device 60 and the display device 70.
[0030] In such a hardware configuration, the functions of each part of the image generation device 10, namely the audience image acquisition unit 11, the content image acquisition unit 12, the action acquisition unit 13, the image generation unit 14, and the image output unit 15, can be implemented by the hardware processor 20 working in cooperation with the data storage unit 32 by reading and executing the program stored in the program storage unit 31.
[0031] Some or all of the components of the image generation device 10 may be configured in a variety of other forms, including integrated circuits such as an application specific integrated circuit (ASIC) or a field-programmable gate array (FPGA).
[0032] [Example of operation] Next, an example of the image generation process executed by the image generation device 10 will be described with reference to Fig. 3. Fig. 3 is a flowchart showing the procedure and content of the image generation process executed by the image generation device 10 according to the embodiment.
[0033] Here, we will explain an example of the following video generation process. A remote audience member is remotely watching video of an event venue, for example, live video. Both the local audience member at the event venue and the remote audience member wave their penlights in time with the live performance and change the color of the penlights as appropriate. The colors of the remote audience member's penlights are captured as the actions of the remote audience member. A virtual audience member image is generated by virtualizing local audience members who have many penlights of the same color as the remote audience member's penlights. The virtual audience member image is combined with content video that does not include the local audience member, and this video is provided to the remote audience member as the video they are watching.
[0034] Furthermore, as pre-settings, the image generator 14 stores pre-set parameters for generating virtual audience images. For example, the pre-set parameters include the viewing environment of the remote audience, and the cooperation speed S and cooperation probability P of the remote audience. The viewing environment of the remote audience includes the virtual audience seating arrangement and the number of virtual audience members. The pre-set parameters are not limited to these and may include other information.
[0035] In step S1, the spectator video acquisition unit 11 acquires video of local spectators at the event venue via the network NW. The spectator video acquisition unit 11 acquires local spectator information based on the video of the local spectators. The spectator video acquisition unit 11 acquires spectator seat video frames based on the local spectator information.
[0036] In step S2, the content video acquisition unit 12 acquires live video of the event venue via the network NW. The content video acquisition unit 12 acquires content video frames based on the live video.
[0037] In step S3, the action acquisition unit 13 acquires video of the remote spectator from the image capture device 60. The action acquisition unit 13 acquires action information of the remote spectator based on the video of the remote spectator. Here, the action acquisition unit 13 acquires the color of the remote spectator's penlight.
[0038] In step S4, the video generation unit 14 determines whether the color of the penlight of the remote spectator has changed. For example, the video generation unit 14 compares the previous action information (penlight color) received from the action acquisition unit 13 with the current action information (penlight color) to determine whether the color of the penlight has changed.
[0039] If the image generating unit 14 determines that the color of the penlight has changed, the process proceeds to step S5, and if it determines that the color of the penlight has not changed, the process proceeds to step S6.
[0040] In step S5, the image generating unit 14 sets the color of the penlight of the remote spectator to the master color after a delay according to the cooperative speed S.
[0041] In step S6, the image generating unit 14 extracts the spatial distribution and colors of the penlights from the audience seat image frame acquired in step S1.
[0042] More specifically, the image generating unit 14 extracts the spatial distribution and color of the penlights as follows.
[0043] S6a: Convert the audience seating video frame to grayscale, and extract areas with a certain brightness and size as areas where penlights are lit. The center coordinates of the extracted areas where penlights are lit are listed as penlight position coordinates.
[0044] S6b: For each lighted penlight part extracted in S6a, the pixel value of the color image is referenced to estimate the color of the penlight and add it to the list.
[0045] S6c: The audience seat area is specified for the audience seat video frame based on the virtual audience seat arrangement in the pre-set viewing environment of the remote audience. Homography transformation is performed on the audience seat video frame for the audience seat area, and the penlight position coordinates calculated in S6a are mapped onto the audience seat video frame without distortion.
[0046] In step S7, the video generator 14 extracts the virtual audience from the audience seat video frame in accordance with the arrangement of the virtual audience seats in the viewing environment of the remote audience, and generates a virtual audience video frame.
[0047] More specifically, the video generator 14 extracts virtual audience members and generates virtual audience video frames as follows.
[0048] S7a: The video frames of the seats of the local spectators at the event venue extracted in step S6 are matched with a virtual seating arrangement that matches the viewing audience of the remote spectators, and multiple aggregation areas are set in the video frames of the seats of the local spectators at the event venue.
[0049] S7b: For each aggregation area, count the actions of local spectators within the aggregation area, i.e., the colors of penlights. If the actions of local spectators within the aggregation area, i.e., the colors of penlights, contain the master color, i.e., the color of penlights of remote spectators, then with cooperation probability P, that master color is made the representative color of that aggregation area. In other words, the actions of local spectators within the aggregation area, i.e., the colors of penlights, are aggregated into the master color, i.e., the color of penlights of remote spectators, with cooperation probability P. On the other hand, if the penlight colors within the aggregation area do not contain the master color, then the actions of local spectators within the aggregation area, i.e., the colors of penlights, of which the most common color is made the representative color of the aggregation area. The actions of local spectators within the aggregation area, i.e., the colors of penlights, are aggregated into the most common action, i.e., the color of penlights of remote spectators.
[0050] S7c: A virtual audience image frame is generated by placing penlights of the representative colors determined in S7b in each aggregation area of the audience seat image frame.
[0051] In step S8, the video generator 14 synthesizes the virtual audience video frame generated in step S7 with the content video frame acquired in step S2 to generate a viewing video frame for the remote audience.
[0052] In step S9, the video output unit 15 outputs the viewing video frame generated in step S8 to the display device 70 in the viewing environment of the remote audience.
[0053] The image generation device 10 repeatedly performs the series of steps S1 to S9 described above.
[0054] [Example of generating a virtual audience video frame] Next, the processing executed by the image generation unit 14 will be described with reference to Fig. 4, with particular emphasis on the generation of virtual audience image frames. Fig. 4 is a diagram for explaining the processing executed by the image generation unit 14 of the image generation device 10 according to the embodiment.
[0055] The image generation unit 14 specifies the spectator seat range for the spectator seat video frame P1 captured from a bird's-eye view. Next, the image generation unit 14 converts the bird's-eye view of the spectator seat range into spectator seat video frame P2, which is a top view of the spectator seat range, using homography transformation.
[0056] Next, the video generator 14 obtains the distribution of the actions, i.e., the colors of the penlights, from the audience seating video frame P2, which is a top view of the audience seating area. Here, circle r represents red penlights, circle b represents blue penlights, and circle y represents yellow penlights.
[0057] Next, the video generator 14 sets aggregation areas that correspond to the virtual seating arrangement in the viewing environment of the remote audience for the audience seating video frame P2, which is a top view of the audience seating area from which the color distribution of the penlights was obtained. Here, as an example, nine rectangular aggregation areas are set using two vertical grids Gv and two horizontal grids Gh.
[0058] Next, the video generation unit 14 aggregates the actions of the local spectators in each aggregated area of the spectator seat video frame P3, i.e., the colors of their penlights, taking into account the action information of the remote spectators received from the action acquisition unit 13, and creates a virtual spectator video frame P4. In this example, the color of the penlights of the remote spectators is yellow.
[0059] The colors of the penlights in each aggregation area are aggregated as follows: For each aggregation area, if there are penlights of the same color as the penlights of the remote spectators, they are aggregated to the color of the penlights of the remote spectators with the cooperation probability P. If there are not penlights of the same color as the penlights of the remote spectators in each aggregation area, they are aggregated to the color of the penlights with the most number of penlights.
[0060] P5 shows a comparative example of a virtual audience image created by simply aggregating the colors of the penlights according to a majority vote. Comparing virtual audience image frames P4 and P5, we see that in virtual audience image frame P4, the aggregated areas in the upper right, center, and lower left are aggregated in yellow, the same color as the penlights of the remote audience, while in virtual audience image frame P5, the aggregated areas in the upper right, center, and lower left are aggregated in colors different from the color of the penlights of the remote audience, i.e., red, blue, and blue, respectively.
[0061] In this way, the virtual audience image frame P4 formed by the image generator 14 is an image that is in harmony with the actions of the remote audience members, that is, the colors of their penlights.
[0062] Finally, the video generation unit 14 synthesizes the virtual audience video frame P4 created as described above with the content video frame received from the content video acquisition unit 12, and outputs the synthesized video frame to the video output unit 15.
[0063] [effect] According to this embodiment, the video of the remote spectator displayed on the display device 70 includes many images of virtual spectators holding penlights of the same color as the remote spectators' penlights. This allows for cooperative interaction between the remote spectators and the virtual spectator images. As a result, the remote spectators can feel a sense of unity with the on-site spectators at the event venue and enjoy the same excitement as the on-site spectators.
[0064] In the embodiment, an example has been described in which cooperation with the actions of remote spectators is emphasized. However, the attributes of the remote spectators may be acquired in advance, and if it is determined from the attributes that the remote spectators do not like to cooperate, the cooperation probability P may be lowered. Also, the color of the aggregation area may be changed to the color of the remote spectators' penlights. For example, the color of the aggregation area may be changed to the color of the second most common penlight in the aggregation area.
[0065] In the embodiment, an example has been described in which the action information of the remote spectator is the color of the remote spectator's penlight. However, the action information of the remote spectator is not limited to this. For example, the action information of the remote spectator may be the phase of the penlight swing (the angle of the penlight), the direction of the penlight swing (vertical swing, horizontal swing), the position of the penlight swing (above the head, under the foot), the motion of the penlight swing (waving in a circular motion), etc.
[0066] The present invention is not limited to the above-described embodiments, and various modifications can be made in the implementation stage without departing from the spirit of the invention. Furthermore, the embodiments may be implemented in appropriate combinations, in which case the combined effects can be obtained. Furthermore, the above-described embodiments include various inventions, and various inventions can be extracted by combining selected elements from the disclosed elements. For example, if the problem can be solved and the desired effect can be obtained even if some elements are deleted from all elements shown in the embodiments, the configuration from which these elements are deleted can be extracted as an invention. [Explanation of symbols]
[0067] 10...Image generation device 11...Audience image acquisition unit 12...Content video acquisition unit 13...Motion acquisition section 14...Image generation unit 15...Video output section 20...Hardware processor 31...Program memory section 32...Data storage unit 41...Communication interface 42...Input / output interface 50...bus 60...Photographing equipment 70…Display device
Claims
1. an audience image acquisition unit that acquires audience image frames of local audience members at the event venue; a content video acquisition unit that acquires a content video frame of the event; an action acquisition unit that acquires action information of remote spectators who remotely watch the event; a video generation unit that generates a virtual audience video frame based on the audience seat video frame and action information of the remote audience, and synthesizes the content video frame with the virtual audience video frame to generate a viewing video of the remote audience; a video output unit that outputs the viewing video to a display device in the viewing environment of the remote audience; having Image generation device.
2. the video generation unit generates the virtual audience video frames including the local audience members whose actions are highly similar to the actions of the remote audience members, based on the action information. The image generating device according to claim 1 .
3. The image generation unit specifying an audience seat range for the audience seat video frame based on a virtual audience seat arrangement in the viewing environment of the remote audience; extracting a virtual audience from the audience seating video frame in accordance with a virtual audience seating arrangement in the viewing environment of the remote audience, and generating the virtual audience video frame; The image generating device according to claim 2 .
4. the video generation unit sets a plurality of aggregation areas in the spectator seat video frame within the spectator seat range, and aggregates the actions of the local spectators within each aggregation area by taking into account the action information of the remote spectators for each aggregation area; The image generating device according to claim 3 .
5. When the actions of the local spectators in each aggregation area include the actions of the remote spectators, the video generation unit aggregates the actions of the local spectators in that aggregation area into the actions of the remote spectators with a cooperation probability P. The image generating device according to claim 4.
6. When the actions of the local spectators in each aggregation area do not include the actions of the remote spectators, the video generation unit aggregates the actions of the local spectators in that aggregation area into the most frequent action among them. The image generating device according to claim 5 .
7. Acquiring video frames of on-site audience seating at the venue of the event; acquiring content video frames of the event; acquiring action information of remote spectators remotely viewing the event; generating a virtual audience video frame based on the audience seat video frame and action information of the remote audience, and synthesizing the content video frame with the virtual audience video frame to generate a viewing video of the remote audience; outputting the viewing video to a display device in the viewing environment of the remote audience; having Video generation method.
8. An image generation program that causes a processor of a computer to execute the functions of each component of the image generation device according to claim 1.
Citation Information
Patent Citations
Video display system, video display method, video display control program and action information transmission program
JP2013021466A
Stage direction system, direction control subsystem, method for operating stage direction system, method for operating direction control subsystem, and program
JP2013037670A
Game device and program
JP2017119031A
System for providing event-related content to user attending event and having respective user terminals
JP2018038056A
Program for providing virtual space by head mount device, method, and information processing device for executing program
JP2019050576A