Video processing method, video processing device, and video processing program
The video processing method adjusts viewpoint positions based on sound changes during live events by using three-dimensional spatial calculations and acoustic processing, enhancing the viewing experience with immersive and responsive visuals and audio.
Patent Information
- Application Number
- JP2025159219
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-25
- Publication Date
- 2025-12-16
AI Technical Summary
Existing information processing devices do not automatically adjust the viewpoint position in response to changes in sound, such as venue shifts during live events, leading to less meaningful viewing experiences.
A video processing method that places objects in a three-dimensional space, calculates relative positions based on sound-related factors, and adjusts the viewpoint position accordingly, incorporating acoustic processing to simulate changes in venue and sound localization.
Enables meaningful changes in viewpoint position that enhance the live event experience both visually and audibly, providing a more immersive and responsive viewing experience.
Smart Images

Figure 2025183409000001_ABST
Abstract
Description
[Technical Field]
[0001] An embodiment of the present invention relates to a video processing method, a video processing device, and a video processing program. [Background technology]
[0002] The information processing device for displaying virtual viewpoint images in Patent Document 1 includes, as viewpoint change requirements, relational conditions that define the positional relationship of a target object and information on two or more objects existing in the viewpoint source space that define the specified viewpoint associated with the relational conditions, and discloses a configuration in which the viewpoint is changed when the relational conditions are satisfied and the viewpoint change requirements are met. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2020-190978 Summary of the Invention [Problem to be solved by the invention]
[0004] The information processing device of Patent Document 1 does not change the viewpoint position in response to a change in sound, such as a change in venue at a live event.
[0005] An object of one embodiment of the present invention is to provide a video processing method that automatically changes the viewpoint position in response to factors that cause changes in the sound of a live event, thereby realizing changes in the viewpoint position that are meaningful for the live event. [Means for solving the problem]
[0006] A video processing method according to one embodiment of the present invention places an object in a three-dimensional space, and outputs video of a live event viewed from a viewpoint set in the three-dimensional space. The video processing method outputs video of the live event viewed from a first viewpoint, receives factors related to sound of the live event, calculates a relative position with respect to the object in accordance with the factors, changes the first viewpoint to a second viewpoint based on the calculated relative position, and outputs video of the live event viewed from the second viewpoint. [Effects of the Invention]
[0007] According to one embodiment of the present invention, it is possible to realize a change in viewpoint position that is meaningful for a live event. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a block diagram showing the configuration of a video processing device 1. FIG. [Figure 2] FIG. 2 is a diagram showing a three-dimensional space R1. [Figure 3] FIG. 10 is a diagram showing an example of a video of a live event viewed from a viewpoint position set in a three-dimensional space R1. [Figure 4] 4 is a flowchart showing the operation of the processor 12. [Figure 5] FIG. 10 is a perspective view showing an example of a three-dimensional space R2 of a different live music venue. [Figure 6] FIG. 10 is a diagram showing an example of a video of a live event viewed from a viewpoint position set in a three-dimensional space R2. [Figure 7] FIG. 1 is a diagram illustrating the concept of coordinate transformation. [Figure 8] FIG. 1 is a diagram illustrating the concept of coordinate transformation. [Figure 9] FIG. 10 is a diagram illustrating a modified example of coordinate transformation. [Figure 10] 10 is a flowchart showing the operation of the processor 12 when performing acoustic processing with the first viewpoint position 50A and the second viewpoint position 50B as listening points. [Figure 11]FIG. 10 is a perspective view showing an example of a three-dimensional space R1 according to Modification 1. [Figure 12] 10 is a diagram showing an example of a video of a live event viewed from a second viewpoint position 50B according to Modification 1. FIG. [Figure 13] FIG. 13 is a perspective view showing an example of a three-dimensional space R1 in Modification 4. [Figure 14] FIG. 13 is a perspective view showing an example of a three-dimensional space R1 in Modification 4. DETAILED DESCRIPTION OF THE INVENTION
[0009] 1 is a block diagram showing the configuration of a video processing device 1. The video processing device 1 includes a communication unit 11, a processor 12, a RAM 13, a flash memory 14, a display 15, a user I / F 16, and an audio I / F 17.
[0010] The video processing device 1 is composed of a personal computer, a smartphone, a tablet computer, or the like.
[0011] The communication unit 11 communicates with other devices such as a server, etc. The communication unit 11 has a wireless communication function such as Bluetooth (registered trademark) or Wi-Fi (registered trademark), and a wired communication function such as USB or LAN.
[0012] The display 15 is made up of an LCD, an OLED, etc. The display 15 displays the image output by the processor 12.
[0013] The user I / F 16 is an example of an operation unit. The user I / F 16 is composed of a mouse, a keyboard, a touch panel, or the like. The user I / F 16 accepts operations by the user. The touch panel may be stacked on the display 15.
[0014] The audio I / F 17 has an analog audio terminal or a digital audio terminal, etc., and is connected to an audio device. In this embodiment, the audio I / F 17 is connected to headphones 20 as an example of audio device, and outputs an audio signal to the headphones 20.
[0015] The processor 12 is composed of a CPU, a DSP, an SoC (System on a Chip), or the like, and corresponds to the processing unit of the present invention. The processor 12 performs various operations by reading a program from a flash memory 14, which is a storage medium, and temporarily storing the program in RAM 13. Note that the program does not need to be stored in the flash memory 14. The processor 12 may download the program from another device, such as a server, when necessary and temporarily store it in RAM 13, for example.
[0016] The processor 12 acquires data related to the live event via the communication unit 11. The data related to the live event includes spatial information, object model data, object position information, object motion data, and sound data of the live event.
[0017] Spatial information is information that indicates the shape of a three-dimensional space corresponding to a live venue such as a live music club or concert hall, and is expressed in three-dimensional coordinates with a certain position as the origin. The spatial information may be coordinate information based on 3D CAD data of a live venue such as an actual concert hall, or may be logical coordinate information (information normalized between 0 and 1) of a fictitious live venue.
[0018] The model data of an object is three-dimensional CG image data, and is made up of multiple image parts.
[0019] The object position information is information that indicates the position of the object in three-dimensional space. In this embodiment, the object position information is expressed as three-dimensional coordinates in a time series according to the elapsed time from the start of the live event. Some objects, such as devices such as speakers, do not change their position from the start to the end of the live event, while others, such as performers, change their position over time.
[0020] The motion data of an object is data that indicates the positional relationship of the plurality of image parts for moving the model data, and the motion data in this embodiment is time-series data that corresponds to the elapsed time from the start of a live event. The motion data is set for objects that change position over time, such as performers.
[0021] The sound data of a live event is, for example, a stereo (L, R) channel sound signal. The processor 12 plays back the sound data of the live event and outputs it to the headphones 20 via the audio I / F 17. The processor 12 may perform effect processing such as reverb processing on the sound signal. Reverb processing is a process that simulates the reverberation (reflected sound) of a room by adding a predetermined delay time to the received sound signal and generating level-adjusted pseudo-reflected sound. The processor 12 applies effect processing appropriate for the live venue. For example, if the live venue indicated by the spatial information is small, the processor 12 generates pseudo-reflected sound with a short delay time and reverberation time and a low level. On the other hand, if the live venue indicated by the spatial information is large, the processor 12 generates pseudo-reflected sound with a long delay time and reverberation time and a high level.
[0022] Furthermore, the sound data of the live event may include sound signals associated with each object that emits sound. For example, speaker objects 53 and 54 are objects that emit sound. The processor 12 may perform localization processing to localize the sounds of the speaker objects 53 and 54 at the positions of the speaker objects 53 and 54, and generate stereo (L, R) channel sound signals to be output to the headphones 20. The localization processing will be described in detail later.
[0023] Processor 12 places objects in a three-dimensional space R1 as shown in Fig. 2 based on object position information from data related to the live event. Processor 12 also sets a viewpoint position within three-dimensional space R1. Processor 12 renders 3DCG images of each object based on the spatial information, object model data, object position information, object motion data, and information on the set viewpoint position, and generates a video of the live event viewed from the set viewpoint position (see Fig. 3).
[0024] Fig. 2 is a perspective view showing an example of a three-dimensional space R1. Fig. 3 is a diagram showing an example of a video of a live event viewed from a viewpoint set in the three-dimensional space R1. Fig. 4 is a flowchart showing the operation of the processor 12. The operation of Fig. 4 is executed by an application program read by the processor 12 from the flash memory 14.
[0025] 2 shows, as an example, a rectangular parallelepiped space. The processor 12 places a performer object 51, a performer object 52, a speaker object 53, a speaker object 54, and a stage object 55 in the three-dimensional space R1. The performer object 51 is an object of a guitar player. The performer object 52 is an object of a singer.
[0026] The processor 12 places the stage object 55 in the center in the left-right direction (X direction), the rearmost in the front-back direction (Y direction), and the lowest in the height direction (Z direction) of the three-dimensional space R1. The processor 12 also places the speaker object 54 in the leftmost in the left-right direction (X direction), the rearmost in the front-back direction (Y direction), and the lowest in the height direction (Z direction) of the three-dimensional space R1. The processor 12 also places the speaker object 53 in the rightmost in the left-right direction (X direction), the rearmost in the front-back direction (Y direction), and the lowest in the height direction (Z direction) of the three-dimensional space R1.
[0027] The processor 12 also places the performer object 51 on the left side of the three-dimensional space R1 in the left-right direction (X direction), at the rearmost position in the front-back direction (Y direction), and above the stage object 55 in the height direction (Z direction). The processor 12 also places the performer object 52 on the right side of the three-dimensional space R1 in the left-right direction (X direction), at the rearmost position in the front-back direction (Y direction), and above the stage object 55 in the height direction (Z direction).
[0028] Then, processor 12 receives a first factor, which is a factor related to an object (S11), and sets first viewpoint position 50A based on the first factor (S12). For example, the first factor is a factor related to stage object 55. In the first process after starting the application program, processor 12 receives a designation of "a position in front of stage object 55 at a predetermined distance" as the first factor.
[0029] Alternatively, the user may specify a specific object via the user I / F 16. For example, the processor 12 displays a schematic diagram of the three-dimensional space R1 of FIG. 2 on the display 15, and displays various objects. The user specifies an arbitrary object via, for example, a touch panel. For example, when the user specifies a stage object 55, the processor 12 sets the first viewpoint position 50A to a position in front of the stage object 55 and a predetermined distance away.
[0030] The first viewpoint position 50A set as described above corresponds to the position of the user viewing the content within the three-dimensional space R1. The processor 12 outputs a video of the live event viewed from the first viewpoint position 50A as shown in Fig. 3 (S13). The video viewed from the first viewpoint position 50A set at "a position in front of the stage object 55 at a predetermined distance" is a video that shows all of the performers of the live event, as shown in Fig. 3.
[0031] The video may be output to the display 15 of the device itself for display, or may be output to another device via the communication unit 11.
[0032] Next, processor 12 determines whether a second factor, which is a factor related to the sound of the live event, has been accepted (S14). For example, the user performs an operation via user I / F 16 to change the current live venue to a different live venue.
[0033] The processor 12 repeats the process of outputting the video based on the first viewpoint position until it determines in S14 that the second factor has been received (S14: NO → S13). As described above, the object position information and motion data are time-series data corresponding to the elapsed time from the start of the live event. Therefore, the positions and movements of the performer objects 51 and 52 change as the live event progresses.
[0034] A second factor related to the sound of a live event is, for example, a change in the live venue. As described above, the processor 12 generates artificial reflected sounds according to the size of the live venue indicated by the spatial information. Therefore, when the live venue changes, the sound of the live event also changes. Therefore, a change in the live venue is a factor related to the sound of the live event.
[0035] Fig. 5 is a perspective view showing an example of a three-dimensional space R2 of a live concert venue different from the three-dimensional space R1. Fig. 6 is a diagram showing an example of a video of a live event viewed from a viewpoint set in the three-dimensional space R2.
[0036] 5 shows an example of a three-dimensional space R2 in the shape of an octagonal prism. The three-dimensional space R2 is larger than the three-dimensional space R1. The user changes the live venue from the three-dimensional space R1 to the three-dimensional space R2 via the user I / F 16.
[0037] If the processor 12 determines in S14 that it has received an operation to change from three-dimensional space R1 to three-dimensional space R2 (S14: YES), it calculates the relative position with respect to a certain object (stage object 55 in the above example) and changes the first viewpoint position 50A to the second viewpoint position 50B based on the calculated relative position (S15).
[0038] Specifically, first, the processor 12 determines the coordinates of the performer object 51, the performer object 52, the speaker object 53, the speaker object 54, and the stage object 55 in the three-dimensional space R2.
[0039] 7 and 8 are diagrams illustrating the concept of coordinate transformation. As an example, Fig. 7 and Fig. 8 show an example in which the coordinates of speaker object 53 are transformed from three-dimensional space R1 to three-dimensional space R2. However, for ease of explanation, Fig. 7 and Fig. 8 show three-dimensional space R1 and three-dimensional space R2 as two-dimensional spaces viewed from above.
[0040] The processor 12 transforms the coordinates of the speaker object 53 from first coordinates in the three-dimensional space R1 to second coordinates in the three-dimensional space R2. In the example of Fig. 7, there are eight reference points 70A (x1, y1), 70B (x2, y2), 70C (x3, y3), 70D (x4, y4), 70E (x5, y5), 70F (x6, y6), 70G (x7, y7), and 70H (x8, y8) before the transformation, and eight reference points 70A (x'1, y'1), 70B (x'2, y'2), 70C (x'3, y'3), 70D (x'4, y'4), 70E (x'5, y'5), 70F (x'6, y'6), 70G (x'7, y'7), and 70H (x'8, y'8) after the transformation.
[0041] The processor 12 determines the center of gravity G of the eight reference points before the transformation and the center of gravity G' of the eight reference points after the transformation, and generates a triangular mesh centered on these centers of gravity.
[0042] The processor 12 transforms the internal space of the triangle before transformation and the internal space of the triangle after transformation using a predetermined coordinate transformation. The transformation uses, for example, an affine transformation. The affine transformation is an example of a geometric transformation. The affine transformation expresses the x coordinate (x') and y coordinate (y') after transformation as functions of the x coordinate (x) and y coordinate (y) before transformation, respectively. That is, the affine transformation performs coordinate transformation using the equations x' = ax + by + c and y' = dx + ey + f. The coefficients a to f can be uniquely determined from the coordinates of the three vertices of the triangle before transformation and the coordinates of the three vertices of the triangle after transformation. The processor 12 transforms the first coordinates into the second coordinates by determining the affine transformation coefficients for all triangles in the same way. The coefficients a to f may also be determined using the least squares method.
[0043] The mesh may be a polygonal mesh other than a triangle, or a combination thereof. For example, the processor 12 may generate a quadrilateral mesh as shown in FIG. 9 and perform coordinate transformation. The transformation method is not limited to the affine transformation. For example, the processor 12 may transform the quadrilateral mesh based on the following equation to perform coordinate transformation (where x0, y0, x1, y1, x2, y2, x3, and y3 are the coordinates of the transformation points, respectively): x'=x0+(x1-x0)x+(x3-x0)y+(x0-x1+x2-x3)xy y'=y0+(y1-y0)x+(y3-y0)y+(y0-y1+y2-y3)xy The transformation method may also be other geometric transformations such as isometry, similarity, or projective transformation. For example, projective transformation is expressed by the equations x'=(ax+by+c) / (gx+hy+1) and y'=(dx+ey+f) / (gx+hy+1). The coefficients are calculated in the same way as for the affine transformation described above. For example, the eight coefficients (a to h) that make up the projective transformation of a rectangle can be uniquely calculated using eight simultaneous equations. Alternatively, the coefficients may be calculated using, for example, the least squares method.
[0044] The above example shows a coordinate transformation in two-dimensional space (X,Y), but coordinates in three-dimensional space (X,Y,Z) can also be transformed using the same method.
[0045] As a result, each object is transformed into coordinates that match the shape of the three-dimensional space R2.
[0046] In this example, the processor 12 also converts the coordinates of the first viewpoint position 50A based on the above conversion method, thereby determining the second viewpoint position 50B, which is a relative position with respect to the stage object 55 in the three-dimensional space R2.
[0047] Then, the processor 12 outputs (S16) a video of the live event viewed from the second viewpoint position 50B as shown in Fig. 5. In this example, the video viewed from the second viewpoint position 50B is a video that shows all of the performers in the live event as shown in Fig. 6.
[0048] However, the coordinates of performer object 51, performer object 52, speaker object 53, speaker object 54, and stage object 55 are transformed in accordance with changes in the live venue, so that the objects are placed at positions separated from one another in accordance with three-dimensional space R2, which is larger than three-dimensional space R1.
[0049] In this way, the video processing device 1 of this embodiment can automatically change the viewpoint position in response to the cause of the change in sound of a live event, and can realize a change in the viewpoint position that is meaningful for the live event.
[0050] (About acoustic treatment) The processor 12 acquires sound signals related to a live event and performs acoustic processing on the acquired sound signals. When the live venue changes, the processor 12 generates reflected sounds appropriate for the new live venue. For example, the processor 12 generates high-level artificial reflected sounds with long delay times and reverberation times appropriate for a three-dimensional space R2 that is larger than the three-dimensional space R1.
[0051] Furthermore, the processor 12 may perform sound processing with the first viewpoint position 50A and the second viewpoint position 50B as listening points. Sound processing with the first viewpoint position 50A and the second viewpoint position 50B as listening points is, for example, the localization processing described above.
[0052] 10 is a flowchart showing the operation of processor 12 when performing acoustic processing with first viewpoint position 50A and second viewpoint position 50B as listening points. Components that are the same as those in FIG. 4 are given the same reference numerals and their explanations will be omitted.
[0053] When outputting video of a live event viewed from first viewpoint position 50A, processor 12 performs, for example, localization processing such that the sounds of speaker objects 53, 54 are localized at the positions of speaker objects 53, 54 (S21). Note that the processing of S13 and S21 may be performed in either order, or may be performed in parallel. Furthermore, because performer object 51 and performer object 52 are also objects that emit sound, processor 12 may perform localization processing such that the sounds of performer object 51 and performer object 52 are localized at the positions of performer object 51 and performer object 52, respectively.
[0054] The processor 12 performs localization processing based on, for example, an HRTF (Head Related Transfer Function). The HRTF represents a transfer function from a virtual sound source position to the user's right ear and left ear. As shown in FIGS. 2 and 3 , the speaker object 53 is located to the front left as viewed from the first viewpoint position 50A, and the speaker object 54 is located to the front right of the first viewpoint position 50A. The processor 12 performs binaural processing to convolve an HRTF onto a sound signal corresponding to the speaker object 53 so as to localize the sound to the front left of the user. The processor 12 also performs binaural processing to convolve an HRTF onto a sound signal corresponding to the speaker object 54 so as to localize the sound to the front right of the user. This allows the user to perceive as if they were at the first viewpoint position 50A in the three-dimensional space R1 and listening to sounds from the speaker object 53 located to the front left and the speaker object 54 located to the front right of the user.
[0055] Then, after transforming the coordinates of each object, processor 12 performs acoustic processing with second viewpoint position 50B as the listening point (S22). Note that the processes of S16 and S22 may be performed in either order, or may be performed in parallel.
[0056] The processor 12 performs binaural processing to convolve an HRTF onto a sound signal corresponding to the speaker object 53 such that the sound is localized to a position in front of and to the left of the user. However, the position of the speaker object 53 as seen from the second viewpoint position 50B in the three-dimensional space R2 is further to the left and in the depth direction than the position of the speaker object 53 as seen from the first viewpoint position 50A in the three-dimensional space R1. Therefore, the processor 12 performs binaural processing to convolve an HRTF onto a sound signal corresponding to the speaker object 53 such that the sound of the speaker object 53 is localized to a position on the left and farther in the depth direction. Similarly, the processor 12 performs binaural processing to convolve an HRTF onto a sound signal corresponding to the speaker object 54 such that the sound of the speaker object 54 is localized to a position on the right and farther in the depth direction.
[0057] In this way, the visual position of the object changes in response to changes in the live venue, and the localized position of the sound of the object also changes. Therefore, the video processing device 1 of this embodiment can realize changes in the viewpoint position that are meaningful both visually and audibly as part of a live event.
[0058] Furthermore, the processor 12 may also perform acoustic processing in the process of generating reflected sounds in the reverb process, with the first viewpoint position 50A and the second viewpoint position 50B as listening points.
[0059] The reflected sounds in reverb processing are generated, for example, by convolving an impulse response measured in advance at an actual live venue with a sound signal. The processor 12 performs binaural processing to convolve HRTFs with the sound signal of the live event based on the impulse response measured in advance at the live venue and the position of each reflected sound corresponding to the impulse response. The positions of reflected sounds in an actual live venue can be obtained by placing multiple microphones at listening points in the actual live venue and measuring the impulse responses. Alternatively, the reflected sounds may be generated based on a simulation. The processor 12 may calculate the delay time, level, and direction of arrival of the reflected sounds based on the position of the sound source, the position of the walls of the live venue based on 3D CAD data, etc., and the position of the listening point.
[0060] When the live venue changes, the processor 12 acquires an impulse response corresponding to the new live venue and performs binaural processing to generate reflected sounds corresponding to the new live venue. When the live venue changes, the position of each reflected sound in the impulse response also changes. The processor 12 performs binaural processing to convolve the HRTF on the sound signal of the live event based on the position of each reflected sound after the change.
[0061] As a result, the live venue changes visually, and the reflected sound also changes to sound that matches the changed live venue. Therefore, the video processing device 1 of this embodiment can realize a change in viewpoint that is meaningful for a live event both visually and aurally, and can provide users with a new customer experience.
[0062] The reflected sounds in the reverb processing include early reflections and late reverberation sounds. Therefore, the processor 12 may process the early reflections differently from the late reverberation sounds. For example, the early reflections may be generated by simulation, and the late reverberation sounds may be generated based on impulse responses measured at a live venue.
[0063] (Variation 1) In the above embodiment, a change in the live venue was shown as an example of the second factor, which is a factor related to the sound of the live event. In Modification 1, the second factor is a factor of a musical change in the live event. A factor of a musical change is, for example, a transition from a state in which all performers are playing or singing to a solo part by one performer. Alternatively, a factor of a musical change may include a transition from a solo part by one performer to a solo part by another performer.
[0064] In Variation 1, data related to a live event includes information indicating the timing of musical changes and information indicating the content of the musical changes. Processor 12 calculates a relative position with respect to the object based on the information, and changes first viewpoint position 50A to second viewpoint position 50B based on the calculated relative position. For example, if data related to a live event includes information indicating the timing of a transition to a guitar solo, processor 12 calculates a relative position with respect to performer object 51, and changes first viewpoint position 50A to second viewpoint position 50B based on the calculated relative position.
[0065] 11 is a perspective view showing an example of a three-dimensional space R1 according to Modification 1. For example, as shown in FIG. 11, the processor 12 sets a second viewpoint position 50B at a position in front of and a predetermined distance away from the performer object 51.
[0066] Fig. 12 is a diagram showing an example of a video of a live event viewed from second viewpoint position 50B according to Modification 1. Processor 12 outputs a video of the live event viewed from second viewpoint position 50B as shown in Fig. 12. Since second viewpoint position 50B is set at a position in front of performer object 51 at a predetermined distance, the video viewed from second viewpoint position 50B shows performer object 51 in a large size.
[0067] In this way, the video processing device 1 of variant example 1 sets the second viewpoint position 50B to a relative position with respect to the guitar performer object 51 at the timing of the transition to a guitar solo, so that when a factor that causes a musical change in the live event occurs, such as a transition to a solo part, it is possible to realize a change in the viewpoint position that is meaningful for the live event.
[0068] In the above embodiment, stage object 55 was specified as the first factor, but for example, processor 12 may accept a selection from the user of, for example, a favorite vocal performer object 52 in the processing of S11 in Fig. 4, and set first viewpoint position 50A to the relative position of the selected vocal performer object 51. In this way, even if the user selects a specific object according to preference, when a factor that causes a musical change in the live event, such as a transition to a solo part, occurs, a change in viewpoint position that is meaningful for the live event can be realized.
[0069] (Variation 2) The video processing device 1 of the second modification example receives an operation to adjust the first viewpoint position 50A or the second viewpoint position 50B from a user. The processor 12 changes the first viewpoint position 50A or the second viewpoint position 50B in accordance with the received adjustment operation. For example, when the processor 12 receives an operation to adjust the height of the second viewpoint position 50B to a higher position, the processor 12 changes the second viewpoint position 50B to a higher position.
[0070] (Variation 3) The video processing device 1 of the third modification records information indicating the viewpoint position as time-series data.
[0071] For example, in the processing of S11 in Figure 4, when the processor 12 receives a selection from the user of, for example, a favorite vocal performer object 52, the processor 12 may set a first viewpoint position 50A to the relative position of the selected vocal performer object 52 and record the first viewpoint position 50A as time-series data.
[0072] Alternatively, as in the second modification, when processor 12 receives an adjustment operation for second viewpoint position 50B from the user, processor 12 may record second viewpoint position 50B after the adjustment as time-series data.
[0073] This allows users to create camerawork data that automatically changes the viewpoint position in response to sound-related elements of the live event, and also adds their own adjustment history. Users can make the created camerawork data public, or obtain other users' public camerawork data and enjoy watching the live event using that other users' camerawork data.
[0074] (Variation 4) In the above embodiment, the first viewpoint position 50A and the second viewpoint position 50B correspond to the position of the user viewing the content in three-dimensional space. In contrast, in the video processing device 1 of Modification 4, the first viewpoint position 50A and the second viewpoint position 50B are set at positions different from the position of the user viewing the content.
[0075] 13 and 14 are perspective views showing an example of a three-dimensional space R1 in Modification 4. The video processing device 1 according to Modification 4 receives an operation from a user to specify the user's position in the three-dimensional space R1. The processor 12 controls objects in the three-dimensional space according to the received user position. Specifically, the processor 12 places a user object 70 at the received user position.
[0076] In this example, the processor 12 also acquires information indicating the user positions of other users and places the other user objects 75. The user positions of other users are received by the video processing devices of the respective users. The user positions of other users are acquired via the server.
[0077] 13, the processor 12 receives a selection of a performer object 51 from the user and sets the first viewpoint position 50A to the relative position of the selected vocal performer object 51. For example, when the user performs an operation to change the position of the first viewpoint position 50A to the rear of the three-dimensional space R1, the processor 12 outputs an image showing the user object 70 and another user object 75.
[0078] Here, as shown in variant example 1, even if the processor 12 sets the second viewpoint position 50B to a relative position with respect to the guitar performer object 51 at the timing of the transition to a guitar solo, the positions of the user object 70 and other user objects 75 do not change.
[0079] For example, at a live event, multiple participants may complete a specific event at a certain timing. For example, a participatory event may be held in which participants light up their penlights in specified colors at specified positions to form a specific pattern throughout the venue.
[0080] The video processing device 1 according to the fourth modification may accept such a multi-participation event as a second factor related to the live event, and may change the first viewpoint position to the second viewpoint position in accordance with the second factor. For example, when all users participating in the live event are positioned at designated user positions, the video processing device 1 sets the second viewpoint position 50B to a position where the entire live venue can be viewed from above.
[0081] The video processing device 1 according to the fourth modification accepts a user position separately from the first viewpoint position 50A and the second viewpoint position 50B, and places the user object 70 and the other user object 75. This allows a bird's-eye view of an event that forms a specific pattern throughout the entire live venue, such as that described above.
[0082] The description of the present embodiment should be considered to be illustrative in all respects and not restrictive. The scope of the present invention is defined not by the above-described embodiments but by the claims. Furthermore, the scope of the present invention includes the scope equivalent to the claims.
[0083] For example, in this embodiment, a live musical event featuring a singer and a guitarist is used as an example of a live event. However, live events also include plays, musicals, lectures, readings, and game tournaments. For example, a game tournament event may be set up as a virtual live venue with multiple player objects, a screen object displaying the game screen, an emcee object, audience objects, and speaker objects emitting game sounds. When a user specifies a first factor specifying a first player object, the video processing device outputs a video of the live venue showing the first player. Then, when a second player wins as a second factor related to the live event, the video processing device outputs a video of the second player. In this way, the video processing device can realize meaningful viewpoint changes for other types of events as live events. [Explanation of symbols]
[0084] 1: Video processing device 11: Communications Department 12: Processor 13: RAM 14: Flash memory 15:Display unit 16: User I / F 17: Audio I / F 20: Headphones 50A: First viewpoint position 50B: Second viewpoint position 51: Performer object 52: Performer object 53: Speaker object 54: Speaker object 55: Stage object 70: User object 70A: Reference point 75: Other user object R1: 3D space R2: 3D space
Claims
1. 1. A video processing method for arranging an object in a three-dimensional space and outputting a video of a live event viewed from a viewpoint position set in the three-dimensional space, comprising: outputting a video of the live event viewed from a first viewpoint; receiving a sound-related parameter for the live event; calculating a relative position with respect to the object in accordance with the factor, and changing the first viewpoint position to a second viewpoint position based on the calculated relative position; outputting a video of the live event viewed from the second viewpoint position; Image processing method.
2. the objects include a stage object for the live event; The factor includes information specifying a front position that is a predetermined distance away from the stage object. The video processing method according to claim 1 .
3. the objects include performer objects for the live event; the factors include information specifying a position corresponding to the performer object; The video processing method according to claim 1 .
4. the second viewpoint position includes a bird's-eye view position overlooking a live venue of the live event, The video processing method according to any one of claims 1 to 3.
5. The factor includes information indicating that the event is a participation event for participants; changing the first viewpoint position to the bird's-eye view position when a condition of the participatory event is satisfied; The video processing method according to claim 4 .
6. 1. A video processing device comprising: a processing unit that arranges an object in a three-dimensional space and outputs a video of a live event viewed from a viewpoint position set in the three-dimensional space, The processing unit outputting a video of the live event viewed from a first viewpoint; receiving a sound-related parameter for the live event; calculating a relative position with respect to the object in accordance with the factor, and changing the first viewpoint position to a second viewpoint position based on the calculated relative position; outputting a video of the live event viewed from the second viewpoint position; Image processing device.
7. a video processing device including a processing unit that arranges an object in a three-dimensional space and outputs a video of a live event viewed from a viewpoint position set in the three-dimensional space; outputting a video of the live event viewed from a first viewpoint; receiving a sound-related parameter for the live event; calculating a relative position with respect to the object in accordance with the factor, and changing the first viewpoint position to a second viewpoint position based on the calculated relative position; outputting a video of the live event viewed from the second viewpoint position; A video processing program that executes the processing.
Citation Information
Patent Citations
Video distribution method and server
JP2019220994A
Video Conferencing System
US20210274129A1
Information processing apparatus, information processing method, and program
JP2020190978A