Video display method and video display device

The video processing method uses AI to analyze sound signals and adjust camera settings, providing immersive event experiences by automatically adapting to solo or ensemble performances, addressing the challenge of achieving high immersion without specialized skills.

WO2025197637A1PCT designated stage Publication Date: 2025-09-25YAMAHA CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/008663
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-22
Filing Date
2025-03-10
Publication Date
2025-09-25

AI Technical Summary

Technical Problem

Achieving high levels of immersion in events without requiring advanced skills in camerawork is challenging, as it typically demands specialized knowledge and expertise.

Method used

A computer-implemented video processing method that determines the viewpoint and cropping range for event filming based on audio data, using artificial intelligence to analyze sound signals and adjust camera positions and settings, thereby creating immersive videos without the need for skilled personnel.

Benefits of technology

Enables highly immersive event experiences for remote viewers by automatically adapting camerawork to solo or ensemble performances, mimicking professional camerawork without requiring advanced skills.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025008663_25092025_PF_FP_ABST
    Figure JP2025008663_25092025_PF_FP_ABST
Patent Text Reader

Abstract

A video display method according to one embodiment is for an event and is achieved by a computer, said video display method: acquiring audio data pertaining to the sounds of the event; outputting a first video which has been processed by determining, on the basis of the audio data, a viewpoint for capturing a video of the event or a cropping range of the captured video of the event; and displaying a second video based on the first video on a display device of a listener of the event.
Need to check novelty before this filing date? Find Prior Art

Description

Image display method and image display device

[0001] An embodiment according to the present invention relates to an image display method and an image display device.

[0002] Patent Document 1 describes a karaoke device that can capture images of a singer with camerawork that matches the song.

[0003] Patent Document 2 describes a CG animation device that works in conjunction with music data.

[0004] Patent Document 3 describes a storyboard production device for producing storyboards.

[0005] JP 11-282479 A JP 2005-56101 A JP 2023-56121 A

[0006] At events, professional cameramen and editors decide the camerawork used at the event, allowing viewers to feel a strong sense of immersion in the event. However, achieving this level of immersion requires advanced skills to decide the camerawork.

[0007] An object of one embodiment of the present invention is to provide a video processing method that can realize camerawork that provides a high level of immersion in an event without requiring advanced skills.

[0008] A video display method according to one embodiment of the present invention is a computer-implemented video display method for an event, which includes: acquiring audio data relating to the sound of the event; determining a viewpoint for filming the event or a cropping range for the video of the event based on the audio data, and outputting a processed first video; and displaying a second video based on the first video on a display device for a listener of the event.

[0009] According to the video processing method of one embodiment of the present invention, camerawork that allows a high level of immersion in an event can be achieved without requiring advanced skills.

[0010] FIG. 1 is a block diagram illustrating an example of a system including a processor 10, microphones 11L and 11R, a camera 12, speakers 13L and 13R, and PCs 3 and 4. FIG. 2 is a top view illustrating an event venue 20. FIG. 3 is a block diagram illustrating an example of the configuration of the processor 10. FIG. 4 is a flowchart illustrating an example of processing executed by the processor 10. FIG. 5 is a diagram illustrating an example of two-dimensional coordinates of the event venue 20. FIG. 6 is an example of an image capturing both performers P1 and P2. FIG. 7 is an example of an image captured with performer P1 at the center. FIG. 8 is a diagram illustrating an example of a virtual space 20v. FIG. 9 is a flowchart illustrating an example of processing executed by the processor 10 according to the second embodiment. FIG. 10 is a diagram illustrating an event venue 20 in which a display 14 is further disposed. FIG. 11 is a block diagram illustrating an example of the configuration of the PC 3. FIG. 12 is a flowchart illustrating an example of processing executed by the processor 10 according to the third embodiment. FIG. 13 is a diagram illustrating an example of left-right rotation of the camera 12. FIG. 14 is a flowchart illustrating an example of processing executed by the processor 10 according to the second modification. Fig. 15 is a diagram showing an event venue 20 in which a camera 15 is further arranged. Fig. 16 is a diagram showing an event venue 20 in which a display 16 is further arranged. Fig. 17 is a diagram showing an event venue 20 in which a camera 17 is further arranged.

[0011] [First Embodiment] A processor 10 according to a first embodiment will be described below with reference to the drawings. FIG. 1 is a block diagram showing an example of a system including a processor 10, microphones 11L and 11R, a camera 12, speakers 13L and 13R, and PCs 3 and 4. FIG. 2 is a top view showing an event venue 20. Hereinafter, unless otherwise specified, all sound signals will be described as digital signals. Note that the number of microphones, speakers, and listener PCs are not limited to those shown in the first embodiment.

[0012] The processor 10 executes the video display method of the present application. The processor 10 corresponds to the video display device of the present application. The processor 10 is used, for example, to distribute events such as live music concerts.

[0013] 1 and 2, the processor 10 is connected via a LAN (Local Area Network) to microphones 11L and 11R, a camera 12, and speakers 13L and 13R that are arranged in an event venue 20 used for an event such as a live performance. The processor 10 may be connected to the microphones 11L and 11R, the camera 12, and the speakers 13L and 13R wirelessly or by wire.

[0014] Note that the event in this application does not necessarily have to be a live music concert, but may be, for example, a comedy show, a live show such as an e-sports tournament, a theatrical performance (such as an opera or musical), or a remote conference.

[0015] It should be noted that the distribution of an event in the present application does not necessarily have to be real-time distribution, and the distribution of an event in the present application may be distribution of recorded content.

[0016] The microphone 11L is placed near the performer P1 as shown in Figure 2 and acquires analog sound signals related to the singing and playing sounds of the performer P1. The microphone 11L acquires digital sound signals by AD converting the analog sound signals. The microphone 11L transmits the encoded digital sound signals to the processor 10 via the LAN as sound data.

[0017] As shown in Figure 2, the microphone 11R is placed near the performer P2 and acquires analog sound signals related to the singing and playing sounds of the performer P2. The microphone 11R acquires digital sound signals by AD converting the analog sound signals. The microphone 11R transmits the encoded digital sound signals to the processor 10 via the LAN as sound data. The singing and playing sounds of the performers P1 and P2 at the event are examples of event sounds in this application. Note that noise unrelated to the event (such as natural sounds) does not correspond to event sounds in this application.

[0018] The processor 10 receives sound data from the microphones 11L and 11R via the LAN. The processor 10 obtains sound signals by decoding the sound data and performs sound signal processing (gain adjustment, effects, etc.) on the sound signals. The processor 10 transmits the processed sound signals as sound data to the speakers 13L and 13R. The speakers 13L and 13R output sounds (singing sounds and playing sounds of performers P1 and P2) based on the received sound data. An event listener U1 at the event venue 20 hears the singing sounds and playing sounds of performers P1 and P2 output from the speakers 13L and 13R.

[0019] The camera 12 captures a video signal by capturing an image of the performers P1 and P2. The camera 12 transmits the video signal as video data to the processor 10 via the LAN. The processor 10 receives the video data from the camera 12.

[0020] The processor 10 is connected to PCs (Personal Computers) 3 and 4 of listeners (hereinafter referred to as far-end listeners) located in remote locations different from the event venue 20 where the performers P1 and P2 are located (see FIG. 1 ) via the Internet 2. The processor 10 transmits audio data and video data to the PCs 3 and 4 via the Internet 2. The PCs 3 and 4 are examples of terminals of listeners located in remote locations different from the locations of the event performers in this application.

[0021] Fig. 3 is a block diagram showing an example of the configuration of the processor 10. As shown in Fig. 3, the processor 10 includes an operation unit 100, a communication interface 101, a DSP (Digital Signal Processor) 102, a CPU (Central Processing Unit) 103, a flash memory 104, and a RAM (Random Access Memory) 105. Note that the processor 10 may further include a SoC (System on a Chip), or may include a SoC instead of the CPU 103 and the DSP 102.

[0022] The operation unit 100 includes a mouse, a touch panel, a keyboard, or the like that accepts user operations.

[0023] The communication interface 101 is an interface conforming to standards such as LAN, HDMI (registered trademark), or USB.

[0024] The DSP 102 performs audio signal processing on the audio signal, and video processing on the video signal.

[0025] The flash memory 104 stores various programs. The various programs include programs that operate the processor 10. However, the flash memory 104 does not necessarily have to store the various programs. The various programs may be stored in a separate device, such as a server. In this case, the processor 10 receives the various programs from the separate device, such as the server.

[0026] The CPU 103 executes various operations by reading out programs stored in the flash memory 104 into the RAM 105. The CPU 103 transmits the sound signals that have been subjected to sound signal processing by the DSP 102 as sound data to the speakers 13L, 13R and the PCs 3 and 4 of the listeners on the far end via the communication interface 101. The CPU 103 transmits the video signals that have been subjected to video processing by the DSP 102 as video data to the PCs 3 and 4 of the listeners on the far end via the communication interface 101. The CPU 103 corresponds to the processing unit in this application.

[0027] The processing executed by the processor 10 will be described below with reference to the drawings. Figure 4 is a flowchart showing an example of the processing executed by the processor 10.

[0028] The processor 10 starts the process when the power of the processor 10 is turned on (FIG. 4: START).

[0029] After starting, the processor 10 acquires sound data from the microphones 11L and 11R (FIG. 4: step S10), and decodes the sound data to acquire a sound signal.

[0030] After step S10, the processor 10 determines the position of the camera 12 within the event venue 20 based on the sound signal ( FIG. 4 : step S11). The processor 10 transmits information indicating the determined position of the camera 12 to the camera 12. The camera 12 receives the information and moves to the determined position based on the information. This changes the position of the camera 12. Determining the position of the camera 12 is an example of determining the "viewpoint from which to photograph the event" in this application.

[0031] As an example, the processor 10 determines the position of the camera 12 using information indicating the shape of the physical space of the event venue 20 (hereinafter referred to as spatial information). The spatial information is information indicating two-dimensional or three-dimensional positions with a predetermined position as the reference point (origin). Such spatial information is obtained, for example, based on CAD data used when the event venue 20 was designed. The flash memory 104 of the processor 10 stores the CAD data.

[0032] It should be noted that the spatial information described above does not necessarily have to be acquired based on CAD data used when designing the event venue 20. The spatial information may be acquired based on data related to the event venue 20 acquired using, for example, a 3D scanner or photogrammetry.

[0033] FIG. 5 is a diagram showing an example of two-dimensional coordinates of the event venue 20. In FIG. 5, the spatial information indicates the coordinates of wall positions and characteristic positions in the event venue 20 (e.g., position PS1: (x1, y1), position PS2: (x2, y2), position PS3: (x3, y3)), with position PS0 at the lower left edge of the event venue 20 as the origin coordinate (x0, y0). As shown in FIG. 5, position PS1 corresponds to the left front of the stage in the event venue 20 and corresponds to the position directly in front of performer P1 at the event. Position PS2 corresponds to the right front of the stage in the event venue 20 and corresponds to the position directly in front of performer P2 at the event. Position PS3 corresponds to the center front of the stage in the event venue 20. Position PS3 is a position between positions PS1 and PS2.

[0034] The processor 10 determines coordinates corresponding to the spatial information as the position of the camera 12. For example, the processor 10 determines position PS1 in FIG. 5 as the position of the camera 12. The processor 10 transmits information (coordinates) indicating the determined position PS1 to the camera 12. Upon receiving the information on the coordinates indicating position PS1, the camera 12 moves to the coordinates of the position corresponding to position PS1.

[0035] For example, the camera 12 has a built-in flash memory. The flash memory stores coordinates corresponding to the current position of the camera 12. As an example, the flash memory stores the coordinates of position PS3 as the coordinates corresponding to the current position of the camera 12. The camera 12 calculates the difference (x3-x1, y3-y1) between the coordinates of position PS3 (x3, y3) corresponding to the current position of the camera 12 and the coordinates of position PS1 (x1, y1) received from the processor 10, thereby determining the amount of movement and direction of movement required to move from the current position to position PS1. The camera 12 moves from the current position (position PS3) to position PS1 based on the amount of movement and direction of movement.

[0036] Fig. 6 is an example of a video that captures both performers P1 and P2. Fig. 7 is an example of a video that captures performer P1 as the center. Position PS3 corresponds to a position in the front center of the stage at the event venue 20. In this case, camera 12 that has moved to position PS3 acquires video that captures both performers P1 and P2, as shown in Fig. 6. Position PS1 corresponds to a position directly in front of performer P1 at the event (see Fig. 5). Therefore, camera 12 that has moved to position PS1 acquires video that captures performer P1 as the center, as shown in Fig. 7.

[0037] In this embodiment, the CPU 103 analyzes the sound signal to determine, for example, whether the performance is an ensemble or a solo, and determines the position of the camera 12 based on the determination result. If the CPU 103 determines that the performance is an ensemble, it moves the camera 12 to position PS3 where it can capture images of both performers P1 and P2. On the other hand, if the CPU 103 determines that the performance is a solo, it moves the camera 12 to a position where it can capture images of the solo performer at the center. If the CPU 103 determines that the performance is a solo performance by performer P1, it transmits coordinate information indicating position PS1 to the camera 12.

[0038] As an example of a method for determining whether a performance is an ensemble or a solo, the processor 10 uses artificial intelligence such as a deep neural network (DNN). For example, the flash memory 104 of the processor 10 stores a trained model that has been trained to determine the relationship between a sound signal and a type of sound source. The CPU 103 inputs the sound signal to the trained model. The trained model identifies the type of sound source contained in the sound signal and outputs the identification result. For example, if the sound signal contains only guitar sounds, the trained model outputs information such as "Identification result: guitar sounds only." Furthermore, if the sound signal contains both guitar sounds and vocal sounds, the trained model outputs information such as "Identification result: guitar sounds and vocal sounds."

[0039] The flash memory 104 stores data indicating the correspondence between the identification results of the trained model and the position of the camera 12. For example, the flash memory 104 stores data indicating the correspondence between "'Guitar sound only: position PS1'" and "Guitar sound and vocal sound: position PS3." The flash memory 104 may further store data indicating the correspondence between "Vocal sound only: position PS2." For example, when the CPU 103 obtains the identification result "Identification result: guitar sound only," the CPU 103 references the data indicating the correspondence to determine that "the position corresponding to 'Identification result: guitar sound only' = position PS1." The CPU 103 transmits coordinate information indicating position PS1 to the camera 12. The camera 12 moves to position PS1. This allows the camera 12 to capture an image centered on performer P1, who is playing the guitar solo.

[0040] Data indicating "singing sound only: position PS2" is stored in flash memory 104, and when CPU 103 obtains an identification result of, for example, "identification result: singing sound only," it can also transmit coordinate information indicating position PS2 to camera 12. In this case, camera 12 can capture an image centered on performer P2, who is the solo singer.

[0041] As described above, processor 10 automatically determines whether a performance is a solo or an ensemble, and if it is a solo performance, moves camera 12 so that the camera is centered on the solo performer, and if it is an ensemble performance, moves camera 12 so that the camera is centered on all performers involved in the performance. In other words, processor 10 can achieve advanced camerawork (such as camerawork manually switched by a highly skilled cameraman, editor, or the like) by changing the viewpoint from which the event is filmed (the position of camera 12) so that the camera is centered on the solo performer if it is a solo performance, but is centered on all performers involved in the performance if it is an ensemble performance. Therefore, even if a highly skilled cameraman, editor, or the like is not present at event venue 20, listeners on the far end can enjoy a customer experience in which they can feel highly immersed in the event thanks to the advanced camerawork achieved by processor 10.

[0042] Note that because FIG. 5 is a top view (plan view), only two-dimensional movement of the camera 12 has been described as an example. However, the camera 12 may also move in the vertical direction. For example, the camera 12 is provided with an elevator that moves (raises) the camera 12 in the vertical direction. The processor 10 determines the amount of elevation of the elevator and transmits information indicating the amount of elevation to the elevator. In this case, the spatial information includes a Z coordinate in addition to the X and Y coordinates. The Z coordinate corresponds to the position (height) of the camera 12 in the vertical direction. The processor 10 determines the Z coordinate position and transmits information indicating the determined Z coordinate to the camera 12. A flash memory built into the camera 12 stores the current Z coordinate (height) of the camera 12. The camera 12 calculates the elevation direction and amount of elevation of the elevator based on the Z coordinate received from the processor 10 and the current Z coordinate of the camera 12. The camera 12 raises and lowers the elevator based on the direction and amount of elevation.

[0043] The camera 12 may be attached to a flying drone. In this case, the flying drone moves in any of the up / down, front / back, left / right directions (XYZ axis directions) based on the coordinates determined by the processor 10 and the coordinates corresponding to the current position of the camera 12. The camera 12 attached to the flying drone moves in accordance with the movement of the flying drone.

[0044] After step S11, the processor 10 receives video data from the moved camera 12 (step S12 in FIG. 4). Finally, the processor 10 outputs (transmits) the audio data and video data to the PCs 3 and 4 of the listeners on the far-end side (step S13 in FIG. 4). The processor 10 repeatedly executes the processes of steps S10 to S13. The audio data transmitted to the PCs 3 and 4 of the listeners on the far-end side corresponds to the audio data for distribution in this application.

[0045] PCs 3 and 4 receive the audio data and video data. CPUs built into PCs 3 and 4 decode the received audio data to obtain audio signals. Speakers built into PCs 3 and 4 receive the audio signals from the CPUs built into PCs 3 and 4 and output audio based on the audio signals. Note that the speakers built into PCs 3 and 4 do not necessarily output audio based on the audio signals; external speakers connected to PCs 3 and 4 may output audio based on the audio signals.

[0046] The CPUs built into PCs 3 and 4 decode the received video data to obtain a video signal related to the first video. Each of the CPUs built into PCs 3 and 4 performs video processing, such as resizing or effects, on the video signal. The display devices (such as liquid crystal displays or organic EL displays) of PCs 3 and 4 receive the video signal related to the second video after the video processing from the CPUs built into PCs 3 and 4. The display devices display the second video based on the input video signal. This allows the display devices to display a video that matches the environment of the far-end listener (such as the resolution of the display). The CPUs of PCs 3 and 4 perform video processing according to the environment of the far-end listener, but the second video may be different or the same for each listener.

[0047] The CPUs of the PCs 3 and 4 do not necessarily have to perform the video processing on the video signal, and may simply output the acquired video signal relating to the first video as the second video to the display device.

[0048] It should be noted that PCs 3 and 4 do not necessarily need to perform image processing such as resizing or effects on the image signal. Processor 10 may perform image processing on the image signal of the first image, encode the image signal related to the second image after the image processing, and transmit the encoded image data to PCs 3 and 4.

[0049] The above effects include, for example, an effect (transition) for naturally connecting video cuts, an effect that changes the video speed, an effect that adjusts differences in color tone or resolution between materials to blend them into a single image, an effect (color effect) that changes the hue of the entire video or the color of a specific object, an effect that adjusts the lighting condition of the video (light effect), a blurring effect such as mosaic or blur, or a vibration effect that changes the position of the video over time. Alternatively, the effect on the video signal may be, for example, a split screen that combines multiple videos and arranges them in a grid pattern.

[0050] Effects for naturally connecting video cuts include, for example, fade, wipe, dissolve, jump cut, glitch, distortion, etc. Effects for changing the speed of video include, for example, slow motion, double speed, time lapse, replay, reverse playback, etc. Effects for adjusting differences in color tone or resolution between materials to blend them together as if they were a single image include, for example, mask, blend effect, chromakey compositing, etc.

[0051] Second Embodiment A processor 10 according to a second embodiment will be described below with reference to the drawings. Fig. 8 is a diagram showing an example of a virtual space 20v. The configuration of the processor 10 according to the second embodiment is the same as the configuration of the processor 10 according to the first embodiment, and therefore Fig. 3 will be applied mutatis mutandis in the description.

[0052] The processor 10 according to the second embodiment differs from the processor 10 according to the first embodiment in that it is used to distribute events that take place in a virtual space.

[0053] In the second embodiment, the processor 10 uses information that defines the shape of the virtual space 20v (see FIG. 8 ). This information is, for example, information expressed in three-dimensional coordinates with a predetermined position as the reference point (origin G0). This information is, for example, information based on 3D CAD data of a live music venue such as an existing concert hall. The flash memory 104 stores this information.

[0054] The CPU 103 places a wall object corresponding to the shape of the virtual space 20v. The CPU 103 also places objects such as speakers within the virtual space 20v. The objects include 3DCG images modeled after real-world objects such as speakers and position information indicating the positions where the objects are placed within the virtual space 20v. In this example, performers P1 and P2 are also represented by 3DCG images of certain characters. The flash memory 104 stores 3DCG image data. The flash memory 104 also stores position information for the objects. The CPU 103 places the objects within the virtual space 20v by reading the 3DCG image data and position information from the flash memory 104. In the example shown in FIG. 8 , the CPU 103 places a speaker object, an object for performer P1 (hereinafter referred to as performer object P1v), and an object for performer P2 (hereinafter referred to as performer object P2v) within the virtual space 20v.

[0055] The processor 10 sets the position, line of sight (direction), and field of view (range) of the listener who will watch the event in the virtual space 20v (hereinafter, the listener's position, line of sight, and field of view will be referred to as the viewpoint). The processor 10 renders a 3DCG image of each object based on the viewpoint, and generates an image of the virtual space 20v as seen from the viewpoint.

[0056] For example, the flash memory 104 stores information indicating a viewpoint G1 (see FIG. 8 ). The viewpoint G1 is a position from which the performer object P1v can be seen (a position directly in front of the performer object P1v), a line of sight, and a field of view. The CPU 103 acquires the information indicating the viewpoint G1 from the flash memory 104. In this case, the CPU 103 renders the 3DCG of each object based on the information indicating the viewpoint G1, and generates a two-dimensional image of the performer object P1v viewed from the viewpoint G1 in the virtual space 20v.

[0057] Similarly, the flash memory 104 stores information indicating viewpoint G2 and information indicating viewpoint G3. Viewpoint G2 is a position from which the view can be centered around performer object P2v (a position directly in front of performer object P2v), a line of sight, and a field of view. Viewpoint G3 is a position from which both performers P1 and P2 can be seen (a position directly in front of performer objects P1v and P2), a line of sight, and a field of view. The process of generating an image based on information indicating viewpoint G2 or viewpoint G3 is the same as the process of generating an image based on information indicating viewpoint G1, and therefore will not be described here.

[0058] The process executed by the processor 10 according to the second embodiment will be described below with reference to the drawings. Fig. 9 is a flowchart showing an example of the process executed by the processor 10 according to the second embodiment.

[0059] After starting, the processor 10 acquires sound data from the microphones 11L and 11R (FIG. 9: step S20). The processor 10 decodes the sound data to acquire a sound signal. The process of step S20 is the same as the process of step S10 shown in FIG. 4, so a description thereof will be omitted.

[0060] Next, the processor 10 acquires information about the far-end listener from the far-end listener's PCs 3 and 4 (FIG. 9: step S21). For example, the processor 10 receives information about the listener (hereinafter referred to as listener information) entered by the listener himself / herself. The listener information includes, for example, information about the listener's favorite performers.

[0061] Next, the processor 10 determines one of the viewpoints G1 to G3 as the viewpoint of the far-end listener based on the sound signal and the listener information (FIG. 9: step S22).

[0062] For example, the CPU 103 determines whether the performance is an ensemble performance or a solo performance by performer P1 using the same determination method as in the first embodiment. If the CPU 103 determines that the performance is an ensemble performance, it determines viewpoint G3, from which both performer objects P1v and P2v can be seen, as the viewpoint of the listeners of PCs 3 and 4.

[0063] On the other hand, if the CPU 103 determines that the performance is a solo performance by performer P1, it determines a viewpoint for each listener on the far end. For example, if the CPU 103 references the listener information for PC3 and obtains information that "favorite performer = performer P1," it determines viewpoint G1, which allows a view centered on performer P1, as the viewpoint for the listener of PC3. If the CPU 103 references the listener information for PC4 and, for example, the listener information for PC4 does not include information about a favorite performer, it determines viewpoint G3 as the viewpoint for the listener of PC4.

[0064] Next, the processor 10 generates video data for the first image rendered to display the virtual space 20v from the determined viewpoint (FIG. 9: step S23). Finally, the processor 10 transmits the sound data and video data to the PC 3 of the far-end listener (FIG. 9: step S24).

[0065] PC3 receives the audio data and the video data. The CPU of PC3 decodes the received video data to obtain a video signal related to the first video. The CPU of PC3 performs video processing such as resizing or effects on the video signal to generate a video signal related to the second video. The display of PC3 inputs the video signal related to the second video from the CPU of PC3. The display displays the second video based on the input video signal. The speaker of PC3 outputs sound based on the audio signal obtained by decoding the audio data, in the same manner as in the first embodiment.

[0066] As described above, the processor 10 can deliver the most suitable video for each listener based on both the sound data and the listener information. This allows the listener on the far end to watch a solo performance by a favorite performer with the favorite performer at the center of the video, or to watch the entire live performance, depending on the state of the performance, providing a highly immersive viewing experience similar to that of a real live performance.

[0067] The listener information may include past information. The past information may be, for example, information about viewpoints selected by listeners in past events of the performer. The processor 10 may accept a viewpoint selection for each listener and transmit a first video corresponding to the accepted viewpoint to the listener's PC. The processor 10 may record a viewpoint selection history for each listener. Then, the processor 10 may, for example, identify the viewpoint most frequently selected for each listener in past events and select it as the viewpoint for the listener during a solo performance in the current event.

[0068] Alternatively, the past information may be, for example, physical information such as the listener's pulse rate during a past event sensed by a sensor. For example, the processor 10 may identify the viewpoint at which the listener's pulse rate was fastest during the past event and select this viewpoint as the listener's viewpoint during the solo performance of the current event.

[0069] Third Embodiment A processor 10 according to a third embodiment will be described below with reference to the drawings. Fig. 10 is a diagram showing an event venue 20 in which a display 14 is further arranged. Since the configuration of the processor 10 according to the third embodiment is the same as the configuration of the processor 10 according to the first embodiment, the following description will be made with reference to Fig. 3 .

[0070] A display 14 (such as a liquid crystal display or organic EL display) is further disposed at the event venue 20 according to the third embodiment. As an example, the display 14 is disposed in front of the performers P1 and P2, as shown in FIG. 10 . This allows the performers P1 and P2 to view the images displayed on the display 14. The processor 10 is connected to the display 14 via a LAN. The display 14 corresponds to the display for the performers of the event in this application.

[0071] Fig. 11 is a block diagram showing an example of the configuration of the PC 3. As shown in Fig. 11, the PC 3 includes an operation unit 30, a communication interface 31, a microphone 32, a speaker 33, a camera 34, a display 35, a flash memory 36, a RAM 37, and a CPU 38. The operation unit 30, the communication interface 31, the flash memory 36, the RAM 37, and the CPU 38 are the same as the operation unit 100, the communication interface 101, the flash memory 104, the RAM 105, and the CPU 103 in the processor 10, and therefore a description thereof will be omitted.

[0072] The microphone 32 acquires sound data relating to the voice of the listener of PC 3. The speaker 33 outputs sound based on the sound data received from the processor 10. The camera 34 acquires video data (hereinafter referred to as far-end video data) by photographing the listener of PC 3. The display 35 is a liquid crystal display, an organic EL display, or the like, and displays video relating to the video data received from the processor 10.

[0073] The process executed by the processor 10 according to the third embodiment will be described below with reference to the drawings. Fig. 12 is a flowchart showing an example of the process executed by the processor 10 according to the third embodiment.

[0074] After the start, the processor 10 transmits audio data and video data to the PCs 3 and 4 of the far-end listeners (FIG. 12: step S30). The process of step S30 is the same as the process of step S13 in the first embodiment, and therefore a description thereof will be omitted.

[0075] The PCs 3 and 4 on the far end receive the audio data and video data from the processor 10. The displays 35 of the PCs 3 and 4 display images based on the video data. The speakers 33 of the PCs 3 and 4 output sounds based on the audio data. The listeners of the PCs 3 and 4 view the images and sounds. The listeners of the PCs 3 and 4 who have viewed the images and sounds react. For example, the listeners of the PCs 3 and 4 react by making a sound.

[0076] The microphones 32 of the PCs 3 and 4 acquire sound signals related to the voices of the listeners of the PCs 3 and 4 and transmit them to the processor 10 as sound data (hereinafter referred to as sound data related to the listeners' voices).

[0077] Next, the processor 10 receives sound data related to the listener's voice from the PCs 3 and 4 ( FIG. 12 : step S31). The processor 10 obtains a sound signal obtained by decoding the sound data related to the listener's voice (hereinafter referred to as the sound signal related to the listener's voice). The sound signal related to the listener's voice is an example of reaction information.

[0078] Next, the processor 10 selects a listener to be displayed on the display 14 from among the listeners of PC3 and PC4 ( FIG. 12 : step S32). As an example, the processor 10 selects the listener who is speaking the loudest among the listeners of PC3 and PC4 as the listener to be displayed on the display 14 based on the signal level of the sound signal related to the listener's voice.

[0079] After step S32, the processor 10 requests the selected listener's PC to send video data ( FIG. 12 : step S33). The listener's PC captures an image of the listener using the camera 34. The listener's PC acquires video data of the far-end side capturing the listener and transmits it to the processor 10.

[0080] The processor 10 receives the video data of the far-end side from the PC of the selected listener (FIG. 12: step S34).

[0081] The processor 10 transmits the received far-end video data to the display 14 ( FIG. 12 : step S35). The display 14 displays video based on the received far-end video data. This allows the performers P1 and P2 to play or sing while watching the listeners enjoying the event and shouting loudly, just as if they were at a real live concert in the event venue 20. As a result, the performers P1 and P2 can feel as if they are performing at the event in an interactive environment similar to a real live concert, and customers can enjoy a highly immersive event experience.

[0082] It should be noted that multiple displays for event performers may be placed within the event venue 20. In this case, the processor 10 may individually select the video data to be transmitted to each of the multiple display devices.

[0083] [Modification 1] The processor 10 according to Modification 1 will be described below with reference to the drawings. The configuration of the processor 10 according to Modification 1 is the same as the configuration of the processor 10 according to the first embodiment, and therefore, description thereof will be omitted.

[0084] The camera 12 according to the first modification has a function of rotating left and right (Pan), a function of rotating up and down (Tilt), and a function of zooming in and out (Zoom). The processor 10 of the first modification sets the Pan, Tilt, or Zoom of the camera 12. Setting the Pan, Tilt, or Zoom of the camera 12 is an example of "determining a cropping range for video of an event" in this application.

[0085] FIG. 13 is a diagram showing an example of left-right rotation of the camera 12. For example, the camera 12 rotates left-right, with the reference direction (angle = 0 degrees) being when it faces a direction D3 in which it can capture both performers P1 and P2. For example, the camera 12 rotates leftward by an angle θ1 with respect to the reference direction. In this case, the camera 12 faces the direction D1 in which it can capture performer P1 as the center, as shown in FIG. 13. On the other hand, the camera 12 rotates rightward by an angle θ2 with respect to the reference direction. In this case, the camera 12 faces the direction D2 in which it can capture performer P2 as the center, as shown in FIG. 13.

[0086] In Modification 1, the processor 10 determines the left-right angle of the camera 12 based on the sound signal. For example, in Modification 1, the flash memory 104 of the processor 10 stores data indicating the correspondence between the identification result using the trained model shown in the first embodiment and the left-right angle of the camera 12. For example, when the CPU 103 obtains the identification result "Identification result: guitar sound only," it refers to the data indicating the correspondence and rotates the camera 12 to the left by an angle θ1. This allows the camera 12 to capture an image centered on the solo performer P1.

[0087] On the other hand, when the CPU 103 obtains the identification result of "Identification result: guitar sound and singing sound," for example, it refers to the data indicating the correspondence and directs the camera 12 to the reference direction (angle = 0 degrees), thereby enabling the camera 12 to capture images of both performers P1 and P2.

[0088] The flash memory 104 of the processor 10 may store data indicating the correspondence between the identification result and the zoom magnification of the camera 12. The CPU 103 may set the zoom magnification of the camera 12 according to the identification result.

[0089] The flash memory 104 of the processor 10 may store data indicating the correspondence between the identification result and the vertical angle of the camera 12. The CPU 103 may set the vertical angle of the camera 12 according to the identification result.

[0090] The PCs 3 and 4 receive the video data acquired by the camera 12 via the processor 10. The displays 35 of the PCs 3 and 4 display images based on the video data.

[0091] As described above, the processor 10 according to the first modification sets the pan, tilt, or zoom of the camera 12 so that, for example, if a solo performance is being recorded, the camera focuses on the solo performer, but if an ensemble performance is being recorded, the camera focuses on all performers involved in the performance. Therefore, listeners on the far end can enjoy a customer experience in which they can feel highly immersed in the event thanks to the sophisticated camerawork achieved by the processor 10, even if a highly skilled cameraman, editor, or the like is not present at the event venue 20.

[0092] In the above example, the processor 10 determines the angle relative to the reference direction (angle = 0 degrees), but may instead determine, for example, a change value relative to the current angle of the camera 12. In other words, determining the cropping range may include determining a change value relative to the current cropping range.

[0093] The processing of the processor 10 according to the first modification can be applied to the processing of the processor 10 according to the first embodiment or the processing of the processor 10 according to the third embodiment.

[0094] In the above, the pan, tilt, or zoom of the camera 12 is shown as an example of determining the cropping range. However, determining the cropping range may also be a process of determining the range to be cropped from the image data captured by the camera 12 (ePTZ: digital pan, digital tilt, digital zoom).

[0095] It should be noted that the processor 10 may set both the viewpoint and the cropping range of the camera 12 .

[0096] [Modification 2] The processor 10 according to Modification 2 will be described below with reference to the drawings. The configuration of the processor 10 according to Modification 2 is the same as the configuration of the processor 10 according to the third embodiment, and therefore, FIG. 3 will be applied mutatis mutandis in the description.

[0097] The processor 10 according to the second modification acquires video data from all listeners (both PCs 3 and 4) on the far end side as reaction information, and selects the video data to be sent to the display 14 by analyzing the video data.

[0098] Fig. 14 is a flowchart showing an example of processing executed by the processor 10 according to Modification 2. The processing of step S30 shown in Fig. 14 is the same as the processing of step S30 shown in Fig. 12, and therefore description thereof will be omitted.

[0099] After step S30, the processor 10 receives video data from the PCs 3 and 4 of the listeners on the far-end side (step S41 in FIG. 14). After step S41, the processor 10 analyzes the received video data and selects the video data on the far-end side to be sent to the display 14 (step S42 in FIG. 14).

[0100] The analysis process is performed by artificial intelligence such as a deep neural network (DNN). As an example of the analysis process, the processor 10 identifies the color of an item worn by the listener. For example, the flash memory 104 of the processor 10 stores a trained model that has been trained to identify the relationship between the video signal and the color of the item worn by the listener. The CPU 103 inputs the video signal to the trained model. The trained model identifies the color of the item worn by the listener that is included in the video signal and outputs the identification result.

[0101] The processor 10 also receives the video signal from the camera 12 and identifies the color of the item worn by the performer. The processor 10 then determines the listeners who are wearing items of the same color as the item worn by the performer. For example, if the processor 10 determines that "the listener of PC3 is a listener who is wearing items of the same color as the item worn by the performer," it selects the far-end video data received from PC3 as the far-end video data to be transmitted to the display 14. The color of the item worn by the performer is an example of a performer attribute representing the performer's attribute in this application. The color of the item worn by the far-end listener is an example of a listener attribute representing the listener's attribute in this application.

[0102] If the processor 10 determines that "both the listeners of PC3 and PC4 are wearing items of the same color as the performer's clothing," it may randomly select one piece of far-end video data from the far-end video data received from PC3 and the far-end video data received from PC4 and transmit it to the far-end display 14. Alternatively, the far-end video data received from PC3 and the far-end video data received from PC4 may be combined.

[0103] The processor 10 transmits the selected far-end video data to the display 14 ( FIG. 14 : step S43). The display 14 displays video based on the received far-end video data. This allows the performer to play or sing while looking at listeners wearing items of the same color as the performer, just like in a real live performance where the audience is inside the event venue 20. As a result, the performer can feel as if they are performing at the event in an interactive environment like a real live performance, and can have a highly immersive event experience.

[0104] The processor 10 may analyze the video signal using artificial intelligence that performs character recognition processing. For example, the processor 10 identifies characters (text data) written on an item worn by the far-end listener.

[0105] The processor 10, for example, acquires text data of the performer's name in advance. If the recognized text data matches the text data of the performer's name, the processor 10 selects video data of the listener. This allows the performer to play or sing while looking at the listener wearing an item with the performer's name written on it, just like a real live performance with the audience inside the event venue 20. The performer's name is an example of a performer attribute in this application. Characters written on an item worn by the listener are an example of a listener attribute in this application.

[0106] The processing by the processor 10 according to the second modification is applicable to the processing by the processor 10 according to any one of the first embodiment, the second embodiment, the third embodiment, and the first modification.

[0107] [Modification 3] The processor 10 according to Modification 3 will be described below with reference to FIG.

[0108] The processor 10 according to the third embodiment differs from the processor 10 according to the third embodiment in that, in the processing of step S12 shown in FIG. 4, one piece of far-end video data is selected from the far-end video data acquired from PCs 3 and 4 based on the listener information and the sound data for distribution.

[0109] For example, the processor 10 according to the third modification causes listeners to register listener information in advance before delivering an event. In the third modification, the listener information further includes information about a favorite instrument.

[0110] For example, when the processor 10 obtains an identification result of "identification result: guitar sound only" and receives information about listeners who like guitar, it displays video data of the corresponding listeners on the display 14. This allows the performer P1 to play or sing while watching the reactions of listeners who like guitar when playing a guitar solo.

[0111] In addition, the processor 10 may acquire information indicating a favorite performer as listener information, and select one piece of video data from the video data on the far-end side based on the acquired listener information and the sound data to be distributed.

[0112] The processing of the processor 10 according to the third modification is applicable to the processing of the processor 10 according to the third embodiment.

[0113] [Modification 4] The processor 10 according to Modification 4 will be described below with reference to FIG.

[0114] In Variation 4, the flash memory 104 of the processor 10 records performances of performers P1 and P2 from past events. The flash memory 104 also stores, in association with the recording data, the timing at which a performance changes between a solo and an ensemble (e.g., the time from the start of the song until the change) and the position of the camera 12 (e.g., positions PS1, PS2, and PS3 shown in FIG. 5 ). The data stored in the flash memory 104 is an example of "information relating to pre-recorded sound from an event" in this application.

[0115] The CPU 103 identifies the song currently being played by the performers P1 and P2 and the time since the start of the song based on the sound data to be distributed. The CPU 103 determines the position of the camera 12 according to the time elapsed since the start of the song currently being played by referencing information related to the sound of the event recorded in advance. The CPU 103 moves the camera 12 to the determined position. The sound of the song currently being played by the performers P1 and P2 identified by the CPU 103 is an example of "sound of the event acquired in real time" in this application.

[0116] By the above processing, the processor 10 can reproduce the camerawork from a past event in the next event for performers P1 and P2.

[0117] The processing of the processor 10 according to the fourth modification is applicable to the processing of the processor 10 according to any one of the first embodiment, the third embodiment, and modifications 1 to 3.

[0118] [Modification 5] Hereinafter, a processor 10 according to Modification 5 will be described with reference to the drawings. Fig. 15 is a diagram showing an event venue 20 in which a camera 15 is further arranged.

[0119] A camera 15 is also placed at the event venue 20 according to Variation 5. As shown in Fig. 15, the camera 15 is placed at position PS3 (a position from which both performers P1 and P2 can be photographed). The rest of the configuration of the camera 15 is the same as that of the camera 12, and therefore a description thereof will be omitted.

[0120] In variant 5, the camera 12 is placed at position PS1 (a position where the performer P1 can be photographed as the center).

[0121] In the fifth modification, the processor 10 switches the video data to be sent to the far-end PCs 3 and 4 depending on whether the performance is an ensemble performance or a solo performance by performer P1. For example, if the processor 10 determines that the performance is an ensemble performance, it sends the video data received from the camera 15 to the far-end PCs 3 and 4. On the other hand, if the processor 10 determines that the performance is a solo performance by performer P1, it sends the video data received from the camera 12 to the far-end PCs 3 and 4.

[0122] Through the above processing, the processor 10 of variant example 5, like the processor 10 of the first embodiment, can achieve advanced camerawork in which, in the case of a solo performance, the filming focuses on the solo performer, while in the case of an ensemble performance, the filming focuses on all performers involved in the performance.

[0123] The processing of the processor 10 according to the fifth modification is applicable to the processing of the processor 10 according to any one of the first embodiment, the third embodiment, and modifications 1 to 4.

[0124] [Modification 6] Hereinafter, a processor 10 according to Modification 6 will be described with reference to the drawings. Fig. 16 is a diagram showing an event venue 20 in which a display 16 is further arranged.

[0125] A display 16 (such as a liquid crystal display or organic EL display) is also placed at the event venue 20 according to Variation 6. As shown in Fig. 16, the display 16 is placed behind the performers P1 and P2. In other words, the display 16 is placed in front of the camera 12. This allows the camera 12 to capture the image displayed on the display 16.

[0126] In the sixth modification, the processor 10 receives video data (video data acquired by capturing an image of the listener on the far-end side) acquired by a camera (not shown) of the far-end PCs 3 and 4 from the PCs 3 and 4. The processor 10 randomly selects one piece of video data from the video data received from the PC 3 and the video data received from the PC 4, and transmits the selected video data to the display 16.

[0127] The display 16 displays an image based on the received video data. In this case, the image displayed on the display 16 (a display placed in the event venue 20) captures the image of the far-end listener. The camera 12 captures the image of the far-end listener displayed on the display 16. Therefore, the display of the far-end PCs 3 and 4 displays the image of the listeners of the far-end PCs 3 and 4 captured on the display 16. This allows the listeners of the far-end PCs 3 and 4 to see their own image displayed on the display 16 in the event venue 20. This allows the listeners of the far-end PCs 3 and 4 to feel like they are participating in the event, and provides a more immersive viewing experience.

[0128] The processing of the processor 10 according to the sixth modification is applicable to the processing of the processor 10 according to any one of the first embodiment, the third embodiment, and modifications 1 to 5.

[0129] [Modification 7] Hereinafter, a processor 10 according to Modification 7 will be described with reference to the drawings. Fig. 17 is a diagram showing an event venue 20 in which a camera 17 is further arranged.

[0130] As shown in Fig. 17 , a camera 17 is further placed at the event venue 20 according to the seventh modification. The camera 17 is placed in front of the listener U1 at the event venue 20. The camera 17 transmits video data acquired by photographing the listener U1 to the processor 10. The other configuration of the camera 17 is the same as the configuration of the camera 12, and therefore a description thereof will be omitted. Note that the event venue 20 may be provided with a plurality of cameras for photographing the listener U1.

[0131] In the seventh modification, the processor 10 transmits the video data received from the camera 17 to the display 16. As a result, the image displayed on the display 16 shows the listener U1. Therefore, the listener on the far end can see the image of the listener U1 (a listener who is in the event venue 20) displayed on the display 16. This allows the listener on the far end to feel as if he or she is participating in the event, resulting in a more immersive experience.

[0132] Instead of placing the camera 17 at the event venue 20, the pan function of the camera 12 shown in the first modified example may be used to capture an image of the listener U1.

[0133] The processing of the processor 10 according to the seventh modification is applicable to the processing of the processor 10 according to any one of the first, second, and third embodiments and modifications 1 to 6.

[0134] [Modification 8] The processor 10 according to Modification 8 will be described below with reference to Fig. 17. The processor 10 according to Modification 8 differs from the processor 10 according to Modification 7 in that it measures the signal levels of the sound signals acquired from the microphones 11L and 11R, and switches the video data to be transmitted to the PCs 3 and 4 based on the magnitude of the signal levels.

[0135] For example, when the signal level of the sound signal acquired from the microphones 11L, 11R is equal to or greater than a predetermined threshold, the processor 10 transmits the video data received from the camera 12 to the PCs 3, 4. On the other hand, when the signal level of the sound signal acquired from the microphones 11L, 11R is less than the predetermined threshold, the processor 10 transmits the video data received from the camera 17 capturing the image of the listener U1 to the PCs 3, 4.

[0136] Through the processing described above, the processor 10 displays footage of the performers P1 and P2 on the displays of the PCs 3 and 4 while the performers P1 and P2 are performing, and displays footage of the listener U1 in the venue on the displays 35 of the PCs 3 and 4 when the performers P1 and P2 are not performing (such as between songs).

[0137] The processing of the processor 10 according to the eighth modification is applicable to the processing according to any one of the first, second, and third embodiments and modifications 1 to 7.

[0138] [Modification 9] The processor 10 according to Modification 9 will be described below with reference to FIG.

[0139] In the process of step S32 shown in FIG. 12, the processor 10 according to the ninth modification selects a listener based on the tempo of the song and the tempo obtained from sensing information obtained by sensing the listener.

[0140] In a ninth variation, the processor 10 identifies the tempo of a song based on, for example, high-level sounds (e.g., drum sounds) that occur at regular beats.

[0141] The tempo of the song may be stored in advance in the flash memory 104. In this case, the CPU 103 identifies the song being played based on the sound data to be distributed, and identifies the tempo of the corresponding song from the flash memory 104.

[0142] The processor 10 also determines the tempo of the listener's arm or head movements as a reaction of the far-end listener. For example, an acceleration sensor is installed on a light stick, hat, or the like worn by the far-end listener. The acceleration sensor senses the listener's arm or head movements. The processor 10 receives data related to the output of the acceleration sensor, and determines the tempo of the far-end listener's arm movements or head movements based on the output of the acceleration sensor.

[0143] The processor 10 may, for example, perform image analysis on the video data received from the far-end PCs 3 and 4 using artificial intelligence or the like to identify the tempo of the arm or head movements of the far-end listener.

[0144] The processor 10 may also identify the tempo of the listener on the far-end side based on the sound of the listener's clapping (high-level sound occurring at a regular beat), for example.

[0145] The CPU 103 calculates the degree of coincidence between the tempo of the song and the tempo of the listener's arm or head movements. The CPU 103 selects, for example, the listener with the highest degree of coincidence as the listener to be displayed on the display 14. This allows the performers P1 and P2 to enjoy the customer experience of performing while watching the listeners reacting enthusiastically by waving their arms or shaking their heads in time with the tempo of the song.

[0146] The processing of the processor 10 according to the ninth modification is applicable to the processing according to any one of the first, second, and third embodiments and modifications 1 to 8.

[0147] [Modification 10] The processor 10 according to Modification 10 will be described below with reference to Fig. 12. The configuration of the processor 10 according to Modification 10 is the same as the configuration of the processor 10 according to the third embodiment, and therefore, description thereof will be omitted.

[0148] In the process of step S34 shown in FIG. 12, the processor 10 according to the tenth modification synchronizes the timing of the beats of the performances of the performers P1 and P2 with the timing of the beats of the far-end listener displayed on the display 14.

[0149] In Modification 10, the processor 10 identifies the timing of beats in the performances of the performers P1 and P2. For example, when the same drum sounds (first and second performance sounds) are repeatedly input, the processor 10 identifies the timing at which the first performance sound and the timing at which the second performance sound are input as the timing of beats.

[0150] In addition, the processor 10 identifies the timing of high-level sounds (e.g., clapping sounds) occurring at regular intervals in the sound data received from the far-end PCs 3 and 4 as the timing of the listener's beats.

[0151] The processor 10 calculates the amount of deviation between the timing of the beats of the performers and the timing of the listener's beats. The processor 10 adjusts the timing of transmitting video data to the display 14 based on the amount of deviation, thereby synchronizing the timing of the beats of the performers P1 and P2 and the timing of the beats of the far-end listener displayed on the display 14.

[0152] As a result, the timing of the far-end listener's beats displayed on the display 14 is synchronized with the timing of the beats of the performers P1 and P2. Performers P1 and P2 can see the images of the listeners clapping along to the timing of their own performance. This allows performers P1 and P2 to enjoy the customer experience of being able to perform their own performance without feeling uncomfortable due to a mismatch in the timing of the far-end listener's beats.

[0153] The processing of the processor 10 according to the tenth modification is applicable to the processing according to any one of the first, second, and third embodiments and modifications 1-9.

[0154] [Modification 11] The processor 10 according to Modification 11 will be described below with reference to FIG.

[0155] In variant example 11, for example, the processor 10 identifies the timing of high-level sounds (e.g., drum sounds) that occur at regular intervals in the performances of performers P1 and P2 as the timing of beats in the performance.

[0156] The processor 10 determines the number of beats in the performance by counting the number of beats from the start of the performance until the next strongest beat is detected. For example, if the processor 10 determines that the time signature is 4 / 4, it determines the period from the start of the song until four beats are counted as one measure, and determines the fifth beat as the beginning of the measure.

[0157] The processor 10 determines the position of the camera 12 at the beginning of a measure, for example (transmits information indicating the position to the camera 12). Determining the position of the camera 12 at the beginning of a measure of a song is an example of "determining a viewpoint based on musical elements of sound data" in this application.

[0158] Through the above processing, the processor 10 can realize advanced camerawork, changing the viewpoint (position of the camera 12) from which the event is filmed at the beginning of each bar of the song (the timing between bars), even if a highly skilled cameraman, editor, etc. is not present at the event venue 20.

[0159] The processor 10 may determine the position of the camera 12, for example, in accordance with the timing of the development of the music. For example, a pop song develops in 16-bar or 32-bar segments. Therefore, the processor 10 may count the number of bars of the song being played, and determine the position of the camera 12 in accordance with the development of the pop song, at the start of the 17th bar or the 33rd bar.

[0160] The processing of the processor 10 according to the eleventh modification is applicable to the processing according to any one of the first, second, and third embodiments and modifications 1-10.

[0161] The above-described embodiments and modifications should be considered to be illustrative in all respects and not restrictive. The scope of the present invention is defined not by the above-described embodiments or modifications, but by the claims. Furthermore, the scope of the present invention is intended to include all modifications that are equivalent to and within the scope of the claims.

[0162] For example, the processor 10 according to the first embodiment (the processor that determines the viewpoint from which to film an event based on sound data) may be used in an event held in a virtual space as shown in the second embodiment. That is, the processor 10 may determine the viewpoint from which to film an event in the virtual space (the viewpoint in the virtual space) based on sound data. For example, the processor 10 may perform the process of determining whether the performance is an ensemble or a solo as shown in the first embodiment, and determine the viewpoint in the virtual space based on the result of this determination.

[0163] 10: Processor, 100, 30: Operation unit, 101, 31: Communication interface, 102: DSP, 103, 38: CPU, 104, 36: Flash memory, 105, 37: RAM, 11L, 11R, 32: Microphone, 12, 15, 17, 34: Camera, 13L, 13R, 33: Speaker, 14, 16, 35: Display, 2: Internet, 3, 4: PC

Claims

1. A computer-implemented method for displaying video of an event, comprising: acquiring audio data relating to the sound of the event; determining a viewpoint for filming the event or a cropping range for the video of the event based on the audio data, and outputting a processed first video; and displaying a second video based on the first video on a display device for a listener of the event.

2. A computer-implemented method for displaying a video of an event in a virtual space, comprising: acquiring sound data relating to the sound of the event; acquiring listener information relating to a listener of the event; determining a viewpoint within the virtual space based on the sound data and the listener information; generating a first image rendered to display the virtual space from the viewpoint; generating a second image based on the first image; and displaying the second image on a display device for the listener.

3. A computer-implemented method for displaying video of an event, comprising: transmitting audio data for distribution relating to the sound of the event to terminals of listeners located in remote locations different from the locations of the performers of the event; acquiring, for each listener, reaction information of the listeners to the event; selecting at least one piece of reaction information from the reaction information acquired for each listener; and displaying video based on the selected reaction information on a display for the performers of the event.

4. The video display method according to claim 1, wherein the second video is transmitted to a terminal of a listener located in a remote location different from the location of the performers of the event.

5. The image display method according to claim 1 or claim 4, wherein the second image obtained by performing image processing on the first image is displayed on the display device.

6. A video display method according to claim 1 or claim 4, wherein the viewpoint or the cropping range is determined based on the sound of the event acquired in real time and information relating to the sound of the event recorded in advance.

7. The video display method according to claim 1 or claim 4, wherein video of the listeners of the event is displayed on a display device located in the venue of the event.

8. The video display method according to claim 1 or claim 4, wherein the viewpoint or the cropping range is determined based on musical elements of the sound data.

9. The video display method according to claim 1 or claim 4, wherein determining the cropping range includes determining a change value for the current cropping range.

10. The video display method according to claim 1 or claim 4, wherein determining the viewpoint includes switching between the videos captured by each of a plurality of cameras.

11. A video display method as described in claim 1 or claim 4, wherein determining the viewpoint includes switching between filming the performers of the event and filming listeners at the event venue other than the performers, based on the level of the sound data.

12. The video display method according to claim 1 or claim 4, wherein the second video is different for each of the listeners.

13. The video display method according to claim 2, wherein the listener information includes past information, and the viewpoint is determined based on the past information.

14. The video display method according to claim 3, wherein the at least one piece of reaction information is selected based on listener information relating to the listener and the audio data to be distributed.

15. The video display method according to claim 3 or claim 14, wherein the reaction information includes a video of the listener.

16. A video display method as described in claim 3 or claim 14, wherein the at least one reaction information is selected based on the degree of coincidence between the tempo obtained from sensing information obtained by sensing the listener and the tempo obtained from the sound data for distribution.

17. A video display method according to claim 3 or claim 14, wherein the at least one piece of reaction information is selected based on a performer attribute representing an attribute of a performer of the event and a listener attribute representing an attribute of the listener.

18. The video display method according to claim 17, wherein the reaction information includes a video of the listener, and the listener attributes are extracted by analyzing the video of the listener.

19. The video display method according to claim 3 or claim 14, further comprising: acquiring first beat information from the sound data for distribution; acquiring second beat information from the reaction information; and synchronizing the video with the sound of the event based on the first beat information and the second beat information.

20. A video display device comprising a processing unit that acquires sound data related to the sound of an event, determines a viewpoint for filming the event or a cropping range of video of the event based on the sound data, outputs a processed first video, and displays a second video based on the first video on a display of a listener of the event.

21. A video display device comprising a processing unit that acquires sound data relating to the sound of an event in a virtual space, acquires listener information relating to a listener of the event, determines a viewpoint in the virtual space based on the sound data and the listener information, generates a first image rendered to display the virtual space from the viewpoint, generates a second image based on the first image, and displays the second image on a display device of the listener.

22. A video display device comprising a processing unit that transmits audio data for distribution related to the sound of an event to a terminal of a listener located in a remote location different from the location of the performers of the event, acquires reaction information of the listeners to the event for each of the listeners, selects at least one piece of reaction information from the reaction information acquired for each of the listeners, and displays video based on the selected reaction information on a display for the performers of the event.

Citation Information

Patent Citations

  • Image provision system

    JP2017216667A

  • Computer system and viewing video generation control method

    JP2024018678A

  • Information processing device and method, display control device and method, reproduction device and method, programs, and information processing system

    WO2016009865A1

  • Information processing device, information processing method, and program

    WO2017154953A1

  • Control device for mobile imaging device, control method for mobile imaging device and program

    WO2018088037A1