Video processing method, video processing system, and video processing program
The video processing system facilitates the generation and distribution of videos from virtual spaces, addressing underutilization by enabling customizable video creation and sharing, thereby enhancing user engagement.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- GLEE HOLDINGS CO LTD
- Filing Date
- 2025-12-29
- Publication Date
- 2026-04-14
AI Technical Summary
Videos and still images generated from virtual spaces, such as computer games and video distribution services, are underutilized beyond simple distribution to users.
A video processing system and method that allows users to generate and distribute videos from virtual spaces, enabling the creation of new videos using format elements from existing videos, and supports various rendering methods including client, browser, and server rendering.
Enables novel utilization of videos from virtual spaces by allowing users to create and share videos with customizable formats, enhancing user creativity and interaction.
Smart Images

Figure 2026065034000001_ABST
Abstract
Description
Technical Field
[0001] The disclosure of this specification mainly relates to a video processing method, a video processing system, and a video processing program.
Background Art
[0002] There is a known system that stores reproduction information necessary for reproducing a video together with the images constituting the video and reproduces the video based on the reproduction information. For example, in Japanese Patent Application Laid-Open No. 2017-029509 (Patent Document 1), reproduction information necessary for reproducing the game scene at the time of capture is included as metadata in an image file generated by capturing the game screen of a game being executed by a certain user, so that another user who designates the image file can start the game from the scene reproduced by the reproduction information included in the image file. A game system is described.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Services using videos in virtual spaces such as computer games and video distribution services are used by many users.
[0005] However, videos and still images generated from virtual spaces are hardly used in any way other than distribution to users.
[0006] The object of various inventions described in this specification is to provide a new usage method for videos generated from virtual spaces.
[0007] The various inventions disclosed herein may solve or alleviate at least some of the problems described in or that can be understood from the descriptions of "Modes for Carrying Out the Invention" herein, in lieu of or in addition to the above-mentioned problems. Where the effects of an embodiment are described herein, the problems of the invention corresponding to that embodiment can be understood based on the description of those effects. [Means for solving the problem]
[0008] One aspect of the present invention relates to a video processing method performed by one or more processors. The video processing method in one aspect comprises the steps of: obtaining one or more format elements used to generate a first video including a first avatar in a virtual space; and generating a second video including a second avatar different from the first avatar using one or more selected format elements selected from the one or more format elements. [Effects of the Invention]
[0009] Embodiments of the present invention provide a novel method for utilizing videos generated from virtual space. [Brief explanation of the drawing]
[0010] [Figure 1] Block diagram of video processing system 1 according to one embodiment. [Figure 2] This diagram schematically illustrates the format data applied to videos distributed by video processing system 1. [Figure 3] This is a block diagram showing the user device 10 included in the video processing system 1. [Figure 4] This is a block diagram showing the server 20 of the video processing system 1. [Figure 5] This figure illustrates the video management information stored in the video processing system 1. [Figure 6]This figure illustrates the format data management information stored in the video processing system 1 shown in Figure 1. [Figure 7] This flowchart illustrates the process of creating a video using format data. [Figure 8] This figure shows an example of an image displayed on a user device. [Figure 9] This diagram schematically shows an example of a selection screen for choosing a data format. [Figure 10] This diagram schematically shows an example of a screen displayed during the process of selecting an avatar item as format data. [Figure 11] This diagram schematically shows a screen containing a list of videos that use the same virtual space as the received video. [Figure 12] This diagram schematically shows a screen containing a list of videos that use the same music as the received video. [Figure 13] This figure shows an example of a video list display, which shows a list of videos being streamed by server 20. [Figure 14] This figure shows an example of a display of a list of format data associated with video data generated by user A. [Figure 15] This diagram schematically shows a screen for setting whether or not format data can be used by others. [Modes for carrying out the invention]
[0011] Hereinafter, embodiments of the various inventions disclosed in this specification (sometimes simply referred to as "the present invention") will be described with reference to the drawings as appropriate. Components common to multiple drawings are denoted by the same reference numerals throughout the drawings. The embodiments of the present invention described below are not intended to limit the invention as defined in the claims. The elements described in the following embodiments are not necessarily essential to the solution of the invention.
[0012] 1. Outline of Video Processing System 1 1-1 System Configuration FIG. 1 is a block diagram showing a video processing system 1 according to an embodiment. The video processing system 1 shown in FIG. 1 includes user devices 10a, 10b, 10c, and a server 20. The server 20 is communicably connected to the user devices 10a, 10b, and 10c via a network 5. The network 5 may be a single network or may be configured by connecting a plurality of networks. The network 5 is, for example, the Internet, a mobile communication network, or a combination thereof. The network 5 may be any network that enables communication between electronic devices.
[0013] The server 20 provides services using a virtual space (virtual space platform) to the user devices 10a, 10b, and 10c. In FIG. 1, only three user devices are shown for simplicity of explanation, but the server 20 can provide services using a virtual space to a large number of four or more user devices. In this specification, for convenience of explanation, it is assumed that the user device 10a is used by user A, the user device 10b is used by user B, and the user device 10c is used by user C. Also, in this specification, the device used by a user to utilize the video distribution service provided by the server 20 is collectively referred to as a "user device 10" or "user device", and a user who utilizes the service of the server 20 using a "user device" may be collectively referred to as a "user". The user devices 10a, 10b, and 10c are all examples of "user devices", and the users A, B, and C are all examples of "users" who utilize the service of the server 20.
[0014] In the virtual space provided by the server 20, a world defined in a three-dimensional global coordinate system may be set. For example, user A can place their own avatar in the virtual space provided by the server 20. When the avatar of user A is placed in the virtual space, the server 20 generates video data representing the avatar and objects included in the virtual space in a three-dimensional model by capturing the virtual space with a virtual camera so as to include the avatar of the user, and transmits the video data to the user device 10a of user A. The user device 10a can generate a video including the avatar of user A by rendering the received video data, and display this video on the display of the user device 10a. User A can move the avatar within the virtual space by operating the avatar while viewing the video of the virtual space displayed on the display. Other users' avatars may be placed in the virtual space.
[0015] In the virtual space, the virtual camera may be placed at a position away from the user's avatar. In this case, the virtual camera can capture the visual field area of the virtual space including the avatar from a position away from the avatar at a predetermined viewing angle. In this case, the virtual camera can capture the user's avatar from a third-person perspective. When the avatar moves within the virtual space, the virtual camera may move so as to follow the avatar while keeping the distance and viewing angle with the avatar constant.
[0016] The virtual camera may be placed at the same position as the user's avatar placed in the virtual space, and capture the visual field area within the virtual space from that position. In this case, the virtual camera captures the visual field area of the virtual space from the viewpoint seen from the user's avatar (that is, the first-person perspective).
[0017] The virtual camera may be set to be switchable between shooting from a third-person perspective and shooting from a first-person perspective. The shooting from a third-person perspective and the shooting from a first-person perspective may be switched according to an operation from the user input via the user terminal.
[0018] The position of the virtual camera within the virtual space may be changed in response to user input via the user terminal. Settings other than the virtual camera position, such as gaze position, gaze direction, and / or field of view, may also be changed in response to user input via the user terminal.
[0019] Server 20 can distribute virtual space video data to the user devices of each user participating in the virtual space. Each user can view the virtual space images displayed on their respective user devices and control their own avatar. Users can interact with other users within the virtual space through their avatars. Users can purchase items, play mini-games, participate in events, and utilize various other functions provided by the virtual space.
[0020] The virtual space provided by Server 20 may be a metaverse space. Multiple users can participate in this metaverse space simultaneously through their respective avatars. In this specification, "virtual space" also includes "metaverse space." In the metaverse space, interactions between users, work involving multiple users, play involving multiple users, and other social activities in the real world are virtually reproduced. Users can participate in the metaverse space through their respective avatars. The world of the metaverse space may be defined in a three-dimensional global coordinate system. User avatars may be able to freely move around within the world of the metaverse space and communicate with each other.
[0021] The video processing system 1 may be configured so that one avatar (character object) among multiple avatars in the virtual space can distribute video as the character object of the distribution user. In other words, the video processing system 1 may be configured to perform one-to-many video distribution in a many-to-many metaverse virtual space. In such a space, there may be no particular distinction between the distribution user and the viewing user.
[0022] The virtual space provided by Server 20 does not necessarily have to have a three-dimensional world defined. Server 20 may provide the user with a home space, which is a type of virtual space in which a three-dimensional world is not defined. The user's avatar may be placed in the home space along with objects placed within the home space (for example, objects that make up the background, objects gifted by other users).
[0023] Server 20 may provide both a virtual space with a world coordinate system and a home space without a world coordinate system. Users may use both the virtual space with a world coordinate system and the home space without a world coordinate system.
[0024] Server 20 may provide an editor function for users to combine objects to construct a virtual space (world). Users may send virtual space definition data that defines the virtual space they have created to Server 20. The server may store the virtual space definition data that defines the virtual space, along with the user ID of the user who created the virtual space. Virtual spaces created by users may be available to other users.
[0025] 1-2 Video distribution and viewing When a user is using a virtual space with an avatar, they can distribute a video containing a field of view defined by a virtual camera within that virtual space to other users via server 20. When user A starts distributing a video using the virtual space, server 20 may create a room corresponding to that distribution. Other users can view user A's video by accessing that room. In this way, users of video processing system 1 can distribute videos of the virtual space captured by a virtual camera, and can also view videos distributed by other users. A distributing user may start distributing a video even when they are not using the virtual space. If a distributing user starts distributing a video when they are not using the virtual space, their avatar may be placed in the virtual space, and a video of the virtual space may be generated by capturing the virtual space in which the avatar is placed, and that video may be distributed. If a distributing user starts distributing a video when they are not using the virtual space, their avatar may be placed in a virtual space selected by the distributing user from among the available virtual spaces, or it may be placed in the distributing user's home space.
[0026] In this specification, a user who creates a video and distributes it via Server 20 is sometimes referred to as a "distributing user," and a user who watches a video distributed from Server 20 is sometimes referred to as a "viewing user." The distinction between a distributing user and a viewing user is not fixed. A user who distributes a video is a distributing user while distributing that video, but is a viewing user while watching a video distributed by another user without distributing their own video. Also, while a user is using the virtual space provided by Server 20 without distributing or watching videos (for example, while interacting with other users via an avatar), that user is neither a distributing user nor a viewing user. Thus, no specific user is fixedly designated as a distributing user; rather, a user who distributes a video among the users participating in the virtual space temporarily becomes a distributing user while distributing that video, and a user who watches a distributed video among the users participating in the virtual space temporarily becomes a viewing user while watching that video.
[0027] A streaming user may stream a video alone or together with other users. When a streaming user streams a video together with other users, the avatars of those other users may also be placed in the virtual space. When a streaming user streams a video together with other users, the video for streaming may be generated by using a virtual camera to capture the field of view in the virtual space, including the streaming user's avatar and the avatars of those other users. An event with multiple users participating may be held in the virtual space, and a video created by filming that event may be streamed from server 20. A streaming user may also be the organizer of such an event with multiple users participating. In other words, an organizer user (sometimes called an event organizer or host user) who organizes an event with multiple users participating in the virtual space can be interpreted as a streaming user. Events with multiple users participating in the virtual space include video chat, voice chat, and virtual events in the virtual space (concerts, gatherings, parties, etc.).
[0028] A streaming user refers to a user who transmits video and / or audio information. For example, a streaming user may be a user who hosts or organizes a solo video stream, a collaborative stream with multiple participants, a video or voice chat with multiple participants and / or viewers, or an event (such as a party) in a virtual space with multiple participants and / or viewers—in other words, a user who primarily performs these activities. Therefore, in this disclosure, a streaming user can also be referred to as a host user, organizer, or event organizer.
[0029] Viewers can not only watch videos but also provide reactions to them. Viewers may also be guest users (users other than the host user) participating in events with multiple users in a virtual space. Viewers can support streamers by sending comments and gifts, and are sometimes called supporters.
[0030] A viewing user refers to a user who receives information related to video and / or audio. However, a viewing user may also be a user who can react to the above information. For example, a viewing user may be a user who watches video streams or collaborative streams, or a user who participates in and / or watches video chats, voice chats, or events. Therefore, in this disclosure, a viewing user can also be referred to as a guest user, participating user, listener, spectator, or supporter.
[0031] The viewer's user device may display a video of a virtual space captured by the broadcaster's virtual camera. The viewer's avatar does not have to be placed in the same virtual space as the broadcaster's avatar. The video displayed on the viewer's user device does not have to include the viewer's avatar.
[0032] The viewer's avatar may be positioned within the virtual space so that it can move around within that virtual space. In this case, the viewer's terminal may display a video representing the field of view of the virtual space captured by a virtual camera associated with the viewer, rather than the video being rendered on the broadcaster's terminal. Along with the video representing the field of view of the virtual space captured by the virtual camera associated with the viewer, the audio being broadcast by the broadcaster may be played.
[0033] 1-3 Video generation In the video processing system 1, by capturing a virtual space with a virtual camera, three-dimensional model data representing a three-dimensional model of the field of view of the virtual space captured according to the setting information of the virtual camera may be generated. The three-dimensional model data may include setting information of the virtual camera (for example, the position of the virtual camera in the virtual space, the gaze position, the gaze direction, the magnification, and the field of view), light source data indicating the position and light intensity of the light source, a depth map, a normal map, and other information necessary to generate a three-dimensional model of the field of view of the virtual space captured by the virtual camera. In this specification, the data representing a three-dimensional model of the field of view of the virtual space generated by capturing the virtual space with a virtual camera is referred to as "video data". By rendering the video data representing the three-dimensional model with a rendering engine, a two-dimensional video (or two-dimensional video frame) that can be displayed on a monitor (display) is generated.
[0034] In the video processing system 1, video data generated on the user device 10 of a distributing user may be distributed to the user devices of other users via the server 20. Alternatively, the user device 10 may send a video (video frame) generated by rendering the video data to the server 20 instead of the video data, and the server 20 may distribute this video. Thus, the user device 10 of a distributing user may distribute three-dimensional video data generated by capturing a virtual space with a virtual camera, or it may distribute a two-dimensional video (video frame) generated by rendering this video data.
[0035] Videos created by users may be streamed in real time. Videos created by users may be stored in storage (for example, storage 25 of server 20) and streamed on demand upon request from viewing users. Videos streamed in real time may be archived on server 20.
[0036] 1-4 Rendering The rendering process, which renders video data and converts it into a two-dimensional video, may be performed by any device included in the video processing system 1. In one embodiment, the rendering of video data may be performed by the user device 10 of the viewing user. The method of rendering video data and generating a two-dimensional video on the user device 10 of the viewing user may be referred to in this specification as the "client rendering method". The video processing system 1 can adopt the client rendering method. In the client rendering method, the user device 10 may obtain a rendering engine from, for example, an application distribution platform before generating the video. The user device 10 may hold avatar display data for representing the appearance of an avatar even before starting the video playback process. In the client rendering method, the user device 10 receives object data, avatar identification information (avatar ID), motion data for representing the movement of the avatar, audio data, and other information necessary for rendering as needed from the server 20, and can render video data based on the information received from the server 20 and the information previously held to generate a two-dimensional video, and display the generated video on the display.
[0037] In another embodiment, rendering of video data may be performed by the user device 10 of the viewing user obtaining a web page from the server 20 and executing a computer program (e.g., Javascript) contained in this web page using a browser installed in the user device 10. This web page may describe the storage locations of data necessary for generating video (e.g., object data, avatar display data, motion data, etc.). For example, the Javascript contained in the web page can obtain various data from the data storage locations described in the web page and generate a two-dimensional video based on this obtained data. A browser is software for viewing web pages written in HTML and is separate from the rendering engine. In addition to web pages obtained from the server 20, the browser can be used to view web pages provided by various servers. In this specification, the method of rendering video data in a browser to generate a two-dimensional video is sometimes referred to as the "browser rendering method". The video processing system 1 can adopt the browser rendering method.
[0038] In yet another embodiment, the rendering of the video data may be performed on the server 20. When rendering is performed on the server 20, the two-dimensional video generated by rendering on the server 20 is transmitted from the server 20 to the user device 10. The user device 10 plays (displays) the video received from the server 20. The method in which the rendering of the video data is performed on the server 20 is sometimes referred to as the "server rendering method" in this specification. In the server rendering method, the video generated on the server 20 is transmitted to the user device 10.
[0039] In yet another embodiment, the rendering of the video data may be performed on the user device 10 of the distribution user. In this case, the user device 10 of the distribution user renders the video data to generate a two-dimensional video, and this generated video is distributed to the user device 10 of the viewing user via the server 20. The method in which the rendering of the video data is performed on the user device 10 of the distribution user may be referred to as the "video distribution method" in this specification.
[0040] As described above, the video processing system 1 can employ any of the following methods: client rendering, browser rendering, server rendering, and video distribution. Therefore, transmitting "video data" from the server 20 to the user device 10 includes transmitting the "video" generated from the video data by the server 20. For example, if the video processing system 1 employs the client rendering method, the server 20 transmits unrendered video data to the user device 10, whereas if the video processing system 1 employs the server rendering method, the server 20 transmits the two-dimensional video obtained by rendering the video data to the user device 10 in frame units. In other words, in this specification, the "video data" transmitted from the server 20 to the user device 10 may be unrendered video data or the video after rendering. Similarly, the "video data" transmitted from the user device 10 to the server 20 may be unrendered video data or the video after rendering.
[0041] 2. Format video associated with format data As described above, in the user device 10, video data is generated by capturing the field of view in the virtual space with a virtual camera. In the video processing system 1, one or more format data can be associated with the video data generated by capturing the virtual space or the video generated by rendering said video data. In this specification, a video or video data associated with one or more format videos may be referred to as a "format video".
[0042] In this specification, unless it is necessary to distinguish between three-dimensional video data and two-dimensional video, both "three-dimensional video data" and "two-dimensional video" generated by rendering such video data may be simply referred to as "video." In this specification, "video" may refer to either the video data before rendering or the two-dimensional video generated by rendering the video data, unless the context requires one interpretation or the other. In this usage, format data is associated with a video generated by capturing a virtual space. Format data may also be stored as metadata for the video, associated with the video.
[0043] When generating video data by capturing the field of view of a virtual space using a virtual camera, the format used to generate that video data can be generated as the format data of the video data. The format data of a given video format can represent the format used when the video is being shot and / or played back. The format data of a video generated by filming a virtual space includes, for example, virtual space identification information that identifies the virtual space being filmed, coordinate information indicating the position of the avatar within the virtual space where it is placed during filming, movement information that represents the coordinates corresponding to the avatar's position as it moves within the virtual space in a time series, avatar direction information indicating the direction the avatar is facing during filming, area identification information that identifies a predetermined area within the virtual space (e.g., inside a specific building within the virtual space, a plaza within the virtual space, an event space within the virtual space, and other areas), virtual camera setting information indicating the settings of the virtual camera to identify the camera work of the virtual camera during filming (e.g., the position of the virtual camera within the virtual space, the gaze position, the gaze direction, the magnification, and the field of view), motion information that identifies the motion of the avatar during filming, object information that identifies an object (a single object or a set or coordination of multiple objects) placed within the field of view of the virtual camera during filming, music information that identifies the music played along with the video, effect information that identifies the effects displayed with the video, filter information that identifies the filter applied to the video (e.g., blur), and insertion data information that identifies text and graphics inserted into the video. The format data for a formatted video may be determined by the user device of the user who generates the formatted video. The format data associated with a formatted video may be sent from the user device 10 that generated the formatted video to the server 20 together with the formatted video, or associated with the formatted video, when the formatted video is sent to the server 20. The user device of a viewing user may obtain the format data associated with the formatted video along with the formatted video to be viewed.
[0044] The format data associated with a formatted video may be used by other users when creating formatted videos. The generation of formatted videos using format data will be explained with reference to Figure 2. Videos M1, M2, and M3 shown in Figure 2 are formatted videos generated by the first user, second user, and third user, respectively. Video M2 is generated using a portion of the multiple format data associated with video M1, and video M3 is generated using a portion of the multiple format data associated with video M2. Videos M1 to M3 may all be stored on server 20 so that they can be viewed by other users. Since Figure 2 is referenced to explain the use of format data, the illustration of elements constituting the virtual space, such as objects, in videos M1 to M3 in Figure 2 is largely omitted. The generation of video M1 to video M3 will be explained step by step below.
[0045] Video M1 is generated by capturing the field of view of the virtual space where the first user's avatar A1 is located using a virtual camera. Video M1 is associated with format data A to D. For example, format data A is coordinate information indicating the position of avatar A1 in the virtual space at the time of shooting video M1, format data B is motion information that identifies the movement of avatar A1 at the time of shooting, format data C is virtual camera setting information that identifies the camera work of the virtual camera at the time of shooting, and format data D is music information indicating the music played together with video M1.
[0046] On the user device of a user viewing video M1, format data A to D are displayed for the user to select. A second user viewing video M1 can generate their own video M2 by using one or more of the format data A to D associated with video M1 as the format. For example, if the second user selects format data A and starts generating a video, the second user's avatar A2 is placed in the same virtual space where avatar A1 is placed in video M1, and at the same position as avatar A1 is placed in video M1. Then, a video of the field of view of the virtual space where avatar A2 is placed is generated. In this way, by selecting format data A associated with video M1, the second user can place their own avatar A2 at the same position as avatar A1 was placed in video M1 and start recording video M2. In other words, by selecting format data A associated with video M1 generated by the first user, the second user can generate their own video M2 using the format defined by format data A. To place avatar A2 at the same position as avatar A1 in video M1 without using format data A, and to generate a video of the virtual space where avatar A2 is placed, the second user must identify the virtual space where avatar A1 is placed from the metadata set in video M1, make avatar A2 appear in that virtual space, and then move avatar A2 within that virtual space to find the location where avatar A1 was placed in video M1. Since format data A indicates coordinate information showing the position of avatar A1 in the virtual space at the time video M1 was shot, the second user can select format data A associated with video M1, thereby saving the trouble of finding the shooting location in the virtual space and allowing avatar A2 to be placed at the same position as avatar A1, and generating a format video that includes avatar A2 placed at that position.
[0047] Similarly, a second user can generate a video of a virtual space that includes their own avatar A2 performing the same movements as avatar A1 in video M1 by selecting format data B and starting video generation. In other words, when a second user selects format data B and starts generating video M2, the motion information that identifies the movements of avatar A1 can be used to make the motion of the second user's avatar A2 in video M2 the same as the motion of avatar A1 in video M1.
[0048] In this way, a second user viewing video M1 generated by a first user can select their preferred format data from among the multiple format data associated with video M1 and generate video M2 using that selected format data. By selecting their preferred format data from among the format data associated with the first user's video M1 and generating video M2 using that format data, the second user can generate video M2 in the desired format more easily and quickly. Furthermore, instead of reproducing all the formats applied to video M1, the second user can select only the format data representing their preferred format from among the formats set in video M1 and generate video M2 using that selected format data. This allows the second user to express their own creativity when creating video M2, rather than simply reproducing the format of an existing video.
[0049] In the example in Figure 2, for the generation of video M2, format data A and B are selected from format data A to D associated with video M1, but format data C and D are not selected. Therefore, when shooting the virtual space to generate video M2, the camera work of the virtual camera represented by format data C is not used. Also, when playing video M2, the music corresponding to the music information indicated by format data D is not played. Thus, the second user can generate video M2 by setting the camera work of the virtual camera for shooting avatar A2 and shooting avatar A2 according to that camera work. The second user can also select the music to be played with video M2. In this way, the second user can easily generate video M2 by using the format data set in video M1 in terms of the placement and motion of avatar A2, while also being able to express their own individuality and creativity in terms of specifying the camera work and selecting the music.
[0050] If the format of video M2 is independently defined by a second user, format data representing the format independently defined by the second user may be associated with video M2. In the example shown in Figure 2, virtual camera setting information that identifies the camera work of a virtual camera determined by the second user when video M2 was shot is associated with video M2 as format data E, and music information selected by the second user is associated with video M2 as format data F.
[0051] When format data representing a format independently defined by a second user is associated with video M2, the format data associated with the second user's video M2 includes inherited format data from video M1 (format data A and B in the example in Figure 2) and original format data independently determined by the second user at the time of shooting video M2 (format data E and F in the example in Figure 2).
[0052] The format data associated with video M2, generated as described above, can be used by other users. In the example shown in Figure 2, it is assumed that video M2 will be viewed by a third user. The third user can generate video M3 using one or more of the format data A-B and E-F associated with video M2. In the example in Figure 2, video M3 is generated using format data A and E from among the multiple format data associated with video M2. Format data G and H are format data corresponding to a format independently determined by the third user. In other words, among the format data associated with video M3, format data A and E are format data inherited from video M2, while format data G and H are original format data representing a format independently determined by the third user. Of the inherited format data, format data A is inherited from video M1 for two generations. Format data E is inherited from video M2 for only one generation.
[0053] In the example shown in Figure 2, the format data for video M1 is used by a second user, and the format data for video M2 is used by a third user. However, the format data for both video M1 and video M2 may be used by two or more users. The first user of video M1 may use the format data for video M2 that was generated using the format data for video M2 but was not used in video M1.
[0054] Server 20 can manage formatted videos and the format data associated with those videos. Server 20 can also store user IDs that identify the user who independently created the format data, associating them with the format data. Furthermore, Server 20 can structurally store the usage relationships of format data so that it can understand the inheritance relationships of the format data.
[0055] In the example shown in Figure 2, only three users are featured for the sake of simplicity, so the inheritance of format data is explained up to two generations. However, format data may be inherited for three or more generations.
[0056] Details regarding the generation and management of format data, and the generation of format videos using format data, will be further explained below, along with the configuration and functions of the user device 10 and the server 20.
[0057] 3. User device 10 (User device 10a) 3-1 Configuration of User Device 10 Next, we will describe user device 10a with further reference to Figure 3. For the sake of brevity, we will describe user device 10a below, but the description of user device 10a also applies to user device 10b, user device 10c, and other user devices.
[0058] User device 10a is an information processing device capable of playing video data or video transmitted from server 20. More specifically, user device 10a is a smartphone, personal computer (PC), mobile phone, tablet terminal, personal computer, e-book reader, wearable computer, game console, head-mounted display, or any other information processing device.
[0059] The user device 10a includes a processor 11, memory 12, user interface 13, communication interface 14, and storage 15.
[0060] The processor 11 is an arithmetic unit that loads the operating system and various other programs from storage 15 or other storage into memory 12 and executes the instructions contained in the loaded programs. The processor 11 may be, for example, a CPU, MPU, DSP, GPU, various other arithmetic units, or a combination thereof. The processor 11 may also be implemented by an integrated circuit such as an ASIC, PLD, FPGA, or MCU.
[0061] Memory 12 is used to store instructions executed by the processor 11 and various other data. Memory 12 is a main memory that the processor 11 can access at high speed. Memory 12 is composed of RAM such as DRAM or SRAM.
[0062] The user interface 13 comprises an input interface for receiving user input and an output interface for outputting various information under the control of the processor 11. The input interface is a keyboard, a pointing device such as a mouse, a touch panel, or any other information input device capable of receiving user input. The output interface is, for example, a liquid crystal display, an organic EL (Electro-Luminescence) display, a display panel, or any other information output device capable of outputting the calculation results of the processor 11.
[0063] The communication interface 14 is implemented as hardware, firmware, or communication software such as a TCP / IP driver or a PPP driver, or a combination thereof. The user device 10a can send and receive data with other information devices, including the server 20, via the communication interface 14.
[0064] Storage 15 is an external storage device accessed by the processor 11. Storage 15 is, for example, a magnetic disk, an optical disk, a semiconductor memory, or any other storage device capable of storing data. Storage 15 may store a video processing application 15a for generating video from received video data. The video processing application 15a may be downloaded to the user device 10a from an application distribution platform (not shown).
[0065] The user device 10a may also include hardware not specifically shown in Figure 3. For example, the user device 10a may include a camera. This camera may be a 3D camera capable of detecting the depth of a person's face in order to detect the feature points of the user's face.
[0066] 3-2 Functions of User Device 10a The processor 11 of the user device 10a functions as a virtual space display unit 11a, a video playback unit 11b, a format data selection unit 11c, a video generation unit 11d, a video transmission unit 11e, and a video editing unit 11f by executing the instruction set included in the video processing application 15a and other instruction sets.
[0067] 3-2-1 Virtual Space Display Unit 11a When user A of user device 10a starts using a virtual space provided by server 20 and places an avatar within the virtual space, the virtual space display unit 11a acquires video data generated by capturing the virtual space with a virtual camera from server 20. This video data is three-dimensional data representing the user's avatar, other users' avatars, objects in the virtual space, and / or other components of the virtual space as three-dimensional models. If the video data was generated by capturing the field of view within the virtual space, including the avatar of user device 10a, from a third-person perspective, the virtual space display unit 11a can render this video data to generate a video of the virtual space including user A's avatar and display this video on the display of user device 10a. If the video data is generated from a first-person perspective captured by the avatar of the user device 10a, the virtual space display unit 11a can render this virtual space video data to generate a video of the virtual space that does not include user A's avatar, and display this video on the display of the user device 10a.
[0068] Various objects may be placed within the virtual space. Objects placed within the virtual space include, for example, objects representing natural terrain such as mountains, hills, rivers, and forests; objects representing structures such as buildings and bridges; furniture placed indoors; objects representing gifts given to the user; and other objects that constitute the virtual space. The video data processed by the virtual space display unit 11a includes objects placed within the virtual space that are located in a field of view determined according to the virtual camera settings.
[0069] 3-2-2 Video Playback Unit 11b As described above, a user using a virtual space with an avatar can distribute a video containing a field of view defined by a virtual camera within that virtual space to other users. The video playback unit 11b can play videos distributed by other users via the server 20. For example, the video playback unit 11b can obtain a list of videos available for distribution from the server 20 and play a video selected by user A from this list. The video playback unit 11b can, for example, receive video data of a video selected by the user as a viewing target from the server 20 and generate a video by rendering the received video data. The image of the generated video is output to the display of the user device 10a, and the sound and / or music played along with the video is output to the speaker of the user device 10a. User A can view the image output to the display and the sound and music output to the speaker. The server 20 may distribute the rendered video (sequence of video frames) to the user device 10a. The video playback unit 11b may play the video by outputting the video frames received from the server 20 to the display.
[0070] The video playback unit 11b can obtain a list of viewable video formats from the server 20 and play a video format selected by user A from this list. The video format may be sent from the server 20 to the user device 10a as rendered video frames. When the video playback unit 11b obtains a video format from the server 20, it also obtains the format data associated with the video format.
[0071] 3-2-3 Format Data Selection Section 11c As described above, each formatted video has format data associated with it that represents the format of that video. The format data selection unit 11c can, for example, display on the display one or more format data associated with a formatted video selected by user A as a viewing target or a formatted video that user A is viewing on user device 10a, and select the format data specified by user A from the displayed one or more format data as the selected format data.
[0072] As described above, format data may be sent from the server 20 to the user device 10a together with the format video being viewed, or associated with the format video data. For example, format data may be set as metadata for the format video. In this case, format data associated with the format video is sent from the server 20 to the user device 10a together with the format video.
[0073] For example, if a song playing in conjunction with the currently viewed format video is selected as format data, the format data selection unit 11c identifies the song information that identifies the selected song as selected format data. The format data selection unit 11c can identify one or more of the multiple format data associated with the currently viewed format video as selected format data. The format data selection unit 11c may also identify two or more of the multiple format data associated with the currently viewed format video as selected format data. For example, in addition to song information, the format data selection unit 11c can identify coordinate information that identifies the position of the avatar displayed in the currently viewed format video in virtual space as selected format data.
[0074] 3-2-4 Video generation unit 11d Next, the functions of the video generation unit 11d will be described. The video generation unit 11d can generate video data representing the field of view of the virtual space captured by a virtual camera placed in the virtual space. This video data is distributed to other user devices via the server 20, as will be described later. Furthermore, if format data is selected by the format data selection unit 11c, the video generation unit 11d can generate a formatted video using the selected format data. If no format data is selected by the format data selection unit 11c, the video generation unit 11d can generate a formatted video without using any format data.
[0075] First, let's explain the generation of videos for distribution. When video distribution is initiated by user A, the video generation unit 11d starts generating video data for distribution. For example, the video generation unit 11d adjusts the virtual camera settings so that the field of view of the virtual space captured by the virtual camera includes user A's avatar. The video generation unit 11d generates video data representing the field of view of the virtual space captured by the virtual camera.
[0076] While the video generation unit 11d is capturing images of the virtual space, user A's avatar can perform predetermined motions, change facial expressions, or move within the virtual space based on instructions input to the user device 10a or detection information detected by the user device 10a. The video generation unit 11d can include data representing the avatar's facial expression changes and motions in the video data so that the avatar's facial expression changes and motions can be reproduced in the rendering engine. For example, the video data can include avatar motion data to represent the avatar's movements and facial expression changes. If the avatar's motion and facial expressions in the virtual space are controlled based on motion data that represents the user's body and face movements in chronological order, acquired by the user device 10a or other sensors, the motion data acquired by the user device 10a or sensors may be used as avatar motion data. The sensors for acquiring motion data may be sensors attached to a part of the user's body.
[0077] The field of view of the virtual space captured by the video generation unit 11d may include not only User A's avatar but also avatars of other users. In this case, the video generation unit 11d may generate video data that includes not only avatar motion data representing the movements and facial expressions of User A's avatar, but also avatar motion data representing the movements and facial expressions of other users' avatars. The avatar motion data of other users may be transmitted from the server 20 to the user device 10a.
[0078] The field of view of the virtual space captured by the video generation unit 11d may not include user A's avatar, but may include avatars of other users. This allows user A to distribute a video that does not include their own avatar but includes avatars of other users. The video generation unit 11d can capture the field of view in the virtual space that includes the avatars of other users by adjusting the virtual camera settings to include those avatars.
[0079] The field of view of the virtual space captured by the video generation unit 11d may contain not only the user's avatar but also characters moving within the virtual space (for example, non-player characters whose actions are controlled by a computer). The video generation unit 11d can generate video data such that characters included in the field of view of the virtual space captured by the virtual camera are rendered by the rendering engine.
[0080] The field of view of the virtual space captured by the virtual camera may include objects associated with User A's avatar. Objects associated with User A's avatar in the virtual space may include objects representing avatar items (e.g., clothing, accessories) worn or attached by User A's avatar, objects representing vehicles that User A's avatar rides, objects representing pets that User A's avatar has with it, and other objects.
[0081] The video generation unit 11d can move the virtual camera in the virtual space or change the settings of the virtual camera in response to instructions input from the user via the user device 10a. The settings of the virtual camera may include the position of the virtual camera in the virtual space, the gaze position, the gaze direction, the magnification, and the field of view. For example, the video generation unit 11d can move the virtual camera in the virtual space or change the settings of the virtual camera in response to instructions from the user to perform zoom shooting, where the virtual camera enlarges and films the user's avatar; panning shooting, where the virtual camera films around the user's avatar; overhead shooting, where the virtual camera films the user's avatar from an overhead position; and other camera work.
[0082] Multiple virtual cameras may be installed within the virtual space. In this case, the video generation unit 11d may switch between the multiple virtual cameras and create video data showing the field of view of the virtual space captured by the active virtual camera. For example, if virtual camera A and virtual camera B are installed in the virtual space, it is possible to switch between virtual camera A and virtual camera B. While virtual camera A is selected and active, video data representing the field of view of the virtual space as seen from the position of virtual camera A may be generated, and while virtual camera B is selected and active, video data showing the field of view of the virtual space as seen from the position of virtual camera B may be generated. Switching between multiple virtual cameras may be performed according to user instructions. The camera work of the virtual cameras may be realized by switching between the active virtual cameras among the multiple virtual cameras installed in the virtual space.
[0083] As described above, the video generation unit 11d can generate video data for distribution by capturing the field of view of the virtual space with a virtual camera.
[0084] Next, the generation of formatted video will be described. When user A initiates the generation of formatted video, the video generation unit 11d captures the field of view of the virtual space using a virtual camera and generates a formatted video. If a format data has been selected by the format data selection unit 11c, the video generation unit 11d uses the selected format data to generate the formatted image. The video generation unit 11d may store the generated formatted video in the storage 15 of the user device 10a, or it may upload it to the server 20. The formatted video uploaded to the server 20 may be stored in the storage 25. The formatted video may be stored in the form of three-dimensional video data, or in the form of two-dimensional video frames.
[0085] The video generation unit 11d can generate a formatted video in the same way as a video for distribution, except that it can use selected format data. For example, the video generation unit 11d may generate a formatted video that includes changes in the avatar's facial expressions and motion. The video generation unit 11d may also generate a formatted video that includes avatars of other users but does not include user A's avatar. The video generation unit 11d may also generate a formatted video that includes characters moving in the virtual space and / or objects associated with user A's avatar. The video generation unit 11d may also generate a formatted video by moving a virtual camera in the virtual space, changing the settings of a virtual camera, or switching the active virtual camera in response to instructions input from the user via the user device 10a.
[0086] When a format data is selected by the format data selection unit 11c, the video generation unit 11d can generate a video in a virtual space using the selected format data from among various format data. Several examples of video generation using the selected format data are described below.
[0087] When the format data selection unit 11c selects virtual space identification information as the selected format data, the video generation unit 11d generates video data by capturing the field of view of the virtual space identified by the selected virtual space identification information with a virtual camera. For example, user A's avatar may be placed in a virtual space identified by the virtual space configuration information selected by the format data selection unit 11c, and video data may be generated by capturing the field of view of the virtual space including the avatar with a virtual camera.
[0088] If the format data selection unit 11c selects coordinate information indicating the position in the virtual space where the avatar is placed during shooting as the selected format data, the video generation unit 11d can place the user's avatar at the position in the virtual space specified by the coordinate information selected by the format data selection unit 11c, and generate video data by capturing the field of view in the virtual space, including the position where the avatar is placed, with a virtual camera.
[0089] If the format data selection unit 11c selects region identification information, which identifies the area in the virtual space where the avatar will be placed during shooting, as the selected format data, the video generation unit 11d can place the user's avatar in the specific area in the virtual space identified by the region identification information selected by the format data selection unit 11c, and generate video data by shooting the field of view area in the virtual space, including the position where the avatar is placed, with a virtual camera.
[0090] When the format data selection unit 11c selects movement information, which represents the coordinates corresponding to the position of an avatar moving in the virtual space in a time series, as the selected format data, the video generation unit 11d can move the user's avatar along the coordinates represented in the time series indicated by the movement information selected by the format data selection unit 11c, and generate video data by capturing the field of view in the virtual space, including the moving avatar, with a virtual camera.
[0091] If the format data selection unit 11c selects avatar direction information, which indicates the direction the avatar is facing during shooting, as the selected format data, the video generation unit 11d positions the user's avatar in the virtual space so that it is facing the direction indicated by the avatar direction information selected by the format data selection unit 11c. For example, if shooting is performed from a first-person perspective, the video generation unit 11d can generate video data by using a virtual camera positioned at the avatar's location to capture the field of view in the virtual space that is facing the direction indicated by the avatar direction information.
[0092] When the format data selection unit 11c selects virtual camera setting information, which specifies the camera work of the virtual camera used during shooting, as the selected format data, the video generation unit 11d can generate video data by shooting the virtual space using the virtual camera settings specified by the selected virtual camera setting information. This allows the virtual space to be shot by reproducing the camera work of the virtual camera within the virtual space specified by the virtual camera setting information. For example, if the virtual camera setting information specifies camera work corresponding to zoom shooting, the video generation unit 11d generates video data by shooting the field of view within the virtual space using zoom shooting. When generating video data, the video generation unit 11d can reproduce not only zoom shooting, but also panning shooting, overhead shooting, virtual camera movement, and camera work by switching between multiple virtual cameras. In other words, when virtual camera setting information is selected as the selected format data, video data is generated so as to reproduce the camera work specified by the virtual camera setting information.
[0093] When the format data selection unit 11c selects motion information that identifies the avatar's motion during shooting as the selected format data, the video generation unit 11d controls the avatar's movement in the virtual space to perform the motion identified by the selected motion information, and generates video data by capturing the field of view in the virtual space, including the avatar performing that motion, with a virtual camera. The motion information can identify, for example, dance choreography. The motion identified by the motion information may last for only a short time, such as a few seconds. For example, the motion information may identify dance choreography that lasts for about 10 seconds. When motion information is selected as the selected format data, video data is generated so that the avatar's movements identified by the motion information are reproduced.
[0094] If the format data selection unit 11c selects object information that identifies an object placed within the field of view of the virtual camera at the time of shooting as the selected format data, the video generation unit 11d can generate video data by placing the object identified by the selected object information within the field of view of the virtual camera and shooting the field of view area in the virtual space including that object with the virtual camera. If the object identified by the object information is an avatar item worn by an avatar, the video generation unit 11d can generate video data by attaching the avatar object identified by the object information to the avatar and shooting the field of view area in the virtual space including the avatar wearing this object with the virtual camera.
[0095] When the format data selection unit 11c selects song information to identify a song as the selected format data, the video generation unit 11d sets the song identified by the selected song information as the song to be played simultaneously with the video. When the video is played, the selected song is played together with the video corresponding to the field of view of the virtual space captured by the virtual camera.
[0096] When the format data selection unit 11c selects effect information that specifies an effect to be displayed along with the video as the selected format data, the video generation unit 11d generates video data that includes the selected effect information. When rendering the video data, the effect corresponding to the effect information is displayed together with the video corresponding to the field of view of the virtual space captured by the virtual camera. The effect may be, for example, an animation displayed in the background of the video. The effect may be, for example, an animation of fireworks being launched displayed in the background of the video.
[0097] If the format data selection unit 11c selects filter information that identifies a filter to be applied to the video as the selected format data, the video generation unit 11d generates video data that includes filter identification data that identifies the said filter. When rendering the video data, a video is generated in which the filter identified by the filter identification data is applied to the video corresponding to the field of view area of the virtual space captured by the virtual camera. The video generated by rendering may be blurred as a result of the filter being applied to the video.
[0098] When the format data selection unit 11c selects insertion data information that identifies text and graphics to be inserted into the video as the selected format data, the video generation unit 11d generates video data that includes the selected insertion data information. During rendering of the video data, the text and graphics identified by the insertion data information are superimposed on the video corresponding to the field of view of the virtual space captured by the virtual camera.
[0099] The video generation unit 11d can generate the format used to generate the format video data as the format data for the video data. For example, the global coordinates of an avatar placed within the field of view of the virtual camera at the time of shooting may be used as format data (coordinate information) that identifies the position where the avatar is placed at the time of shooting. Information that identifies the area in the virtual space included in the field of view of the virtual camera at the time of shooting may be used as format data (area identification information) that identifies the area where the avatar is placed at the time of shooting. If the avatar moves in the virtual space at the time of shooting, the coordinates in the virtual space represented in chronological order to correspond to the movement of the avatar may be used as format data (movement information) that indicates the movement of the avatar. The direction in which the face of the avatar placed within the field of view of the virtual camera at the time of shooting is facing may be used as format data (avatar direction information) that indicates the direction in which the avatar is facing at the time of shooting. If the video data includes avatar motion data that defines the motion of the avatar, the avatar motion data may be used as format data (motion information) that identifies the motion of the avatar. Data that can identify the movement of the avatar generated from the avatar motion data may be used as motion information. If the video data includes information regarding the settings of a virtual camera corresponding to camera work such as zoom shooting, this information regarding the virtual camera settings may be used as format data (virtual camera setting information) that identifies the camera work of the virtual camera used during shooting. For example, the virtual camera settings during shooting (e.g., virtual camera position, gaze position, gaze direction, magnification, and field of view) may be recorded in chronological order, and the information indicating the virtual camera settings recorded in this chronological order may be used as virtual camera setting information. Alternatively, the virtual camera settings that are changed in response to a series of motions of the avatar and changes in its position in the virtual space during shooting (e.g., virtual camera position, gaze position, gaze direction, magnification, and field of view) may be recorded in chronological order, and the information indicating the virtual camera settings recorded in this chronological order may be used as virtual camera setting information.For example, if the virtual camera pans when the avatar turns around during first-person perspective shooting, the virtual camera may record the changes in the direction of gaze and field of view of the virtual camera as the avatar turns around in chronological order, and this information indicating the virtual camera settings recorded in chronological order may be used as virtual camera setting information. If the field of view of the virtual space captured by the virtual camera includes an object associated with the avatar, the object ID that identifies the object may be used as format data (object information) that identifies the object placed within the field of view of the virtual camera at the time of shooting. The object information may include the local coordinates of the object included in the field of view of the virtual space captured by the virtual camera. The object information may also identify the avatar items worn by the avatar included in the field of view. As described above, when the video generation unit 11d captures the field of view of the virtual space with the virtual camera and generates video data, it can set information indicating various formats used to generate the video data as format data associated with the video data.
[0100] When generating a formatted video by capturing a field of view within a virtual space using a virtual camera, the virtual space ID that identifies the virtual space can be used as formatted data (virtual space identification information) that identifies the virtual space to be captured. The formatted data related to the format used during capture is not limited to those explicitly described herein.
[0101] If the format image is generated by the format data selection unit 11c using selected format data, the video generation unit 11d may use the selected format data as the format data for the format image. As will be described later, the format data may be assigned a format data ID that uniquely identifies the format data. If the video generation unit 11d generates a format video using the selected format data, the format data may be associated with the format data that was assigned to the selected format data.
[0102] The format data generated by the video generation unit 11d may be transmitted to the server 20 by the video transmission unit 11e, described later, together with the format video, or in association with the format video. For formats set in the format video, the format data ID assigned to the selected format data can be used as the format data for the format specified using the selected format data. Therefore, the format data associated with the format video may include the format data ID assigned to the selected format data used when the format video was generated, and format data that identifies a newly created format when the format video was generated (for example, coordinate information, avatar motion data, etc., generated as described above). In this specification, the format data that identifies a newly created format without using the selected format data when the format video was generated is sometimes referred to as "original format data".
[0103] The video generation unit 11d can generate format data as metadata for the video data. Therefore, the video generation unit 11d can store the format data associated with the video data together with the video data.
[0104] 3-2-5 Video transmission unit 11e Next, the functions of the video transmission unit 11e will be described. The video transmission unit 11e can transmit the video for distribution or the video data generated by the video generation unit 11d to the server 20. As described above, the server 20, upon receiving the video data, can transmit the video data to the user device 10a of the viewing user. The video data is rendered on the server 20 or the user device 10a of the viewing user, and a video corresponding to the video data is generated. If the video processing system 1 employs a video distribution method, the video transmission unit 11e can transmit the video (a series of video frames constituting the video) generated by rendering the video data to the server 20. If the video generation unit 11d generates format data in association with the video data, the format data is transmitted to the server 20 in association with the video data or together with the video data.
[0105] If the video created by the video generation unit 11d is to be streamed in real time (live), the video transmission unit 11e transmits video data representing the video created by the video generation unit 11d to the server 20 in real time. The video created by the video generation unit 11d may also be transmitted to the server 20 after a delay in time.
[0106] Furthermore, the video transmission unit 11e can transmit the formatted video generated by the video generation unit 11d to the server 20. The formatted video may be transmitted to the server 20 together with the format data generated for the formatted video by the video generation unit 11d. The server 20 stores the received formatted video, for example, in storage 25. The server 20 may make the formatted video received from the user device 10a viewable by other users, or it may prevent viewing until instructed by user A.
[0107] The video processing system 1 may have a first distribution mode for distributing formatted videos and a second distribution mode for distributing videos other than formatted videos. In the first distribution mode, in addition to the formatted video, the format data associated with the formatted video is transmitted to the user device of the user who views the formatted video. As a result, the user who views the formatted video can view the formatted video and can also create a new formatted video using the format data associated with the formatted video. In the second distribution mode, the video created by the video generation unit 11d is distributed, but the format data is not distributed. In this specification, the video distributed in the second distribution mode may be referred to as a "normal video". When a normal video is distributed, a formatted video is not distributed in association with the normal video.
[0108] 3-2-6 Video Editing Department 11f Next, the functions of the video editing unit 11f will be described. The video editing unit 11f can edit the formatted video generated by the video generation unit 11d. For example, the video editing unit 11f can edit the video generated by the video generation unit 11d by changing the format data associated with the formatted video generated by the video generation unit 11d. Before making the formatted video viewable, user A can play the formatted video on their user device 10a and check the played formatted video and music. At that time, by changing the selected format data that was selected when the formatted video was generated to a different format data, the formatted video and the music played with the formatted video can be changed. For example, if a formatted video is generated by selecting motion information A, which identifies the movement of an avatar, as the format data, the video editing unit 11f can, based on user A's operation, adopt motion information B instead of motion information A as the format data. As a result, in the formatted image before editing, the avatar performs the motion identified by motion information A, but in the formatted video after editing, it performs the motion identified by motion information B. The video editing unit 11f can store the changed format data in association with the format video if the association between the format video and the format data is changed. For example, if the motion information associated with the format video is changed from motion information A to motion information B, motion information B will be stored in association with the format data. As another example, if song information that identifies song A is associated with the format video as format data, song information that identifies another song selected by the user (e.g., song B) can be associated with the format video as new format data.
[0109] If the association between a formatted video and formatted data is changed by the video editing unit 11f, the changed formatted data is sent to the server 20 along with the formatted video. If the formatted video before editing has already been sent to the server 20, the video editing unit 11f can send an update request to the server 20 to associate the changed formatted data with the said formatted video.
[0110] The format data may be changed after the format video has been made viewable. In this case as well, the video editing unit 11f can send an update request to the server 20 requesting a change (or update) of the format data associated with the format video. The server 20 can store the format video and the changed format data in association and send this changed format data together with the format video to the user device of the viewer. The video editing unit 11f may change any of the format data associated with the video data.
[0111] 4 Server 20 4-1 Server 20 Configuration Next, referring further to Figure 4, the functions of the server 20 and the data stored in the server 20 will be described. The server 20 comprises a processor 21, memory 22, a user interface 23, a communication interface 24, and storage 25. The processor 21, like the processor 11, is a CPU or other type of arithmetic unit. The memory 22 is main memory that the processor 21 can access at high speed. The user interface 23 comprises an input interface that accepts user input for the server 20 and an output interface that outputs various information under the control of the processor 21. The communication interface 24, like the communication interface 14, is a combination of various drivers, software, or a combination thereof for communicating with other devices. The storage 25, like the storage 15, is a storage device capable of storing data, such as a magnetic disk. The processor 21 loads the operating system and various other programs from the storage 25 or other storage into memory and executes the instructions contained in the loaded programs.
[0112] 4-2 Data stored in storage 25 Storage 25 stores format video management information 25a and format data management information 25b. Other data may also be stored in storage 25.
[0113] 4-2-1 Format Video Management Information 25a The format video management information 25a will be explained with reference to Figure 5. The format video management information 25a is a dataset in which data for managing format videos is structurally stored. For each format video, the format video management information 25a includes format video identification information that identifies the format video, associated with creator user identification information that identifies the user who created the format video, the format video itself, and format data.
[0114] Video identification information for a given format video is, for example, a video ID that identifies that format video. The video ID may be assigned to the format video in response to the server 20 receiving the format video from the user device 10. The video ID uniquely identifies the format video.
[0115] The creator identification information for a particular video format is, for example, the user ID of the user who created that video format. For example, if user A generates a video format on user device 10a and sends that video format to server 20, user A's user ID is stored as the creator identification information.
[0116] Format video information for a given format video is, for example, the video frames that make up that format video. As described above, a format video generated in the user device 10a can be transmitted to the server 20 in the form of a two-dimensional video or video frames generated by rendering the video data. In the format video management information 25a, the rendered format video or its video frames may be stored in association with a video ID that identifies the format video. The format video information may also include the video data of the format video (video data before rendering). The video data of the format video is, for example, three-dimensional model data representing a three-dimensional model of the field of view of a virtual space captured by a virtual camera. A two-dimensional video (video frame) is generated by rendering the video data with a rendering engine. The video data may include data other than the three-dimensional model data necessary for rendering. For example, the video data of a format video may include object identification data (object ID) that identifies each of the one or more objects included in the format video, avatar identification data that identifies the avatar included in the format video, avatar motion data that identifies the movement of the avatar included in the format video, and other data necessary to generate the format video.
[0117] Format data for a given format video is the format data associated with that format video. As previously described, the user device 10 of the distribution user transmits the format video together with the format video, or associated with the video data of the format video, the format data associated with that format video. When the server 20 receives the format video together with the format video or the format data associated with the format video, the received format data is stored in association with the video ID of that format video. Specifically, the format data may include various data that define the format of the video data, such as virtual space identification information, coordinate information, and avatar motion data. The same data may be stored separately as video data and format data for the format video. For example, avatar motion data may be included in both the video data and the format data.
[0118] 4-2-2 Format Data Management Information 25b The format data management information 25b will be explained with reference to Figure 6. The format data management information 25b is a dataset in which data for managing format data associated with a format video is structurally stored. For each format data associated with a format video, the format data management information 25b includes format data identification information that identifies the format data, creator user identification information that identifies the user who created the format data, and the format data itself.
[0119] Format data identification information for a given format data is, for example, a format data ID that identifies the format data. The format data ID may be assigned to the format data in response to the server 20 receiving the format data from the user device 10 of the distribution user. As described above, the format data associated with a format video may include the format data ID assigned to the selected format data used when the format video was generated, and format data that identifies a newly created format when the format video was generated (for example, coordinate information, avatar motion data, etc., generated as described above). When the server 20 receives a format video from the user device 10 of the distribution user, a new format data ID may be assigned only to the original format data that identifies a newly created format when the format video was generated, among the format data associated with that format video.
[0120] The user identification information for a given format data is, for example, the user ID of the user who initially (originally) created that format data. When server 20 receives format data from user device 10 that does not have a format data ID associated with a format video, it assigns a format data ID to the format data and stores the user ID of the user device 10 in association with that format data ID. This allows the original format data to be uniquely identified by the format data ID, and the user who created that original format data to be uniquely identified by the user ID associated with that format data ID.
[0121] As described above, when the server 20 receives original format data associated with a format video, a format data ID is assigned to the original format data. Therefore, in the format data management information 25b, the original format data is stored in association with the format data ID (format data identification information). As previously described, the original format data is various data representing a format generated from a format video without using selected format data. The original format data may be at least one selected from the group consisting of, for example, virtual space identification information, coordinate information, region identification information, movement information, avatar direction information, virtual camera setting information, motion information, object information, effect information, filter information, and insertion data information. The format data associated with the format data identification information may be tokenized as an NFT (non-fungible token) or SFT (semi-fungible token). The NFT or SFT conversion of the format data may be performed in response to a request from the distribution user.
[0122] 4-3 Functions of Server 20 Next, the functions performed by the processor 21 of the server 20 will be described. The processor 21 functions as a video data distribution unit 21a and a format data trading unit 21b by executing computer-readable instructions contained in the program recorded in the storage 25.
[0123] In the second distribution mode, the video data distribution unit 21a can distribute the regular video requested in the distribution request to the user device 10 of a viewing user, in response to the distribution request from the user device 10. The server 20 may send a list of available regular videos to the user device 10. The user of the user device 10 may select a regular video they wish to watch from the list of available regular videos. The user device 10 can send a distribution request for the video selected by the user to the server 20. The video data distribution unit 21a may distribute the regular video selected by the user to the user device 10 of that user. In this way, in the second distribution mode, regular videos can be viewed on the user device 10.
[0124] In the second distribution mode, the video data distribution unit 21a may transmit the video data (three-dimensional model data) of a normal video to the user device 10 of the viewing user. In this case, the user device 10 of the viewing user may generate a video for playback (normal video) by rendering the video data. If the video processing system 1 employs a server rendering method or a video distribution method, the video data distribution unit 21a may transmit the rendered normal video to the user device 10a.
[0125] Furthermore, in the first distribution mode, the video data distribution unit 21a may send a format video list, which lists format videos that can be viewed by the user, to the user device 10. The format video list may include format videos stored in the format video management information 25a that the creator has permitted other users to view. Users of the user device 10 may select a format video they wish to view from the format video list. The video data distribution unit 21a may send the format video selected by the user from the format video list to the user device 10 of that user. The format video may be sent to the user device 10 together with the format data associated with the format video. In this way, in the first distribution mode, the format video can be viewed on the user device 10.
[0126] In the first distribution mode, the video data distribution unit 21a may transmit the rendered formatted video to the user device 10 of the viewing user. In this case, the user device 10 of the viewing user can play the formatted video by outputting the formatted video received from the server 20 to a display. Thus, in the first distribution mode, the video may be distributed in a different format than in the second distribution mode. In the video processing system 1, a client rendering method may also be adopted for rendering the formatted video. In this case, the video data (three-dimensional model data) of the formatted video is transmitted from the server 20 to the user device 10 of the viewing user, and the video data is rendered in the user device 10 to generate the formatted video.
[0127] In the services provided by Server 20, the right to sell the original format data can be granted to the user who created the original format data registered in the format data management information 25b. The format data trading unit 21b may create a list of items for sale, including the original format data and the selling price, and make it public to users. For example, if the format data trading unit 21b receives a request to sell the original format data from the user who created it, it can add the original format data to the list of items for sale. If the format data trading unit 21b receives a purchase request from a user who wishes to purchase original format data included in the list of items for sale, it will sell the requested original format data to the user who wishes to purchase it. The format data trading unit 21b may grant the right to use the original format data to other users, either for a fee or free of charge. This right to use the data may be granted exclusively to one user or non-exclusively to multiple users.
[0128] If the format data is in NFT or SFT format, the format data trading unit 21b may provide the functionality of an NFT or SFT exchange. The NFT or SFT format data may be traded on an exchange outside of the video processing system 1.
[0129] 5. Flow for generating formatted videos using formatted data. Next, with reference to Figure 7, a method for generating another format video using format data associated with a format video distributed in the first distribution mode will be described. The processing in each step shown in Figure 7 is performed by processor 11. Processor 11 may, if necessary, perform the processing shown in Figure 7 in cooperation with other processors. In Figure 7, it is assumed that a format video created by user C and made viewable by other users is received by user A's user device 10a. For example, the format video list provided by server 20 includes format video C created by user C. User A obtains the format video list containing format video C from server 20 and selects format video C as the format video to be viewed from the format video list. It is also assumed that multiple format data are associated with the format video created by user C. User A can generate a format video by viewing another user's format video, even if they are not normally distributing videos. If user A is not normally distributing, user A's avatar may not be located in a specific virtual space, but the format video can be filmed using user A's home space. Furthermore, at the time of starting the recording of the formatted video, a selection of candidate virtual spaces may be displayed on user A's user device 10a so that the virtual space to be used for recording the formatted video can be chosen.
[0130] First, in step S11, the format video C created by user C is transmitted from server 20 to user A's user device 10a, and the format video C is received by user device 10a.
[0131] In step S11, the formatted video C received is played back on the user device 10a. For example, the video of formatted video C is output to the display of the user device 10a, and the audio and / or music of formatted video C is output to the speaker of the user device 10a. Formatted video C may also be transmitted from the server 20 in the form of pre-rendered video data, rather than in the form of a video or video frames. In this case, in step S12, the received video data is rendered on the user device 10a, and the formatted video generated by this rendering is played back.
[0132] Figure 8 shows an example of an image corresponding to a formatted video being played on the user device 10a. Image 30, corresponding to the formatted video, is displayed on the user device 10a's display. Image 30 includes the avatar 31 of user C, who created the formatted video, as well as objects 32, 33, 34, and 35. A play button 41 is overlaid on image 30 displayed on the display. When the play button 41 is selected, playback of the formatted video begins.
[0133] Since the formatted video received in step S11 has multiple format data associated with it, the format data is acquired along with the formatted video in step S11. The image 30 displayed on the screen has an overlay icon 42 for using the format data associated with the formatted video received in step S11. Icon 42 is an example of a selection element that accepts selection from the user.
[0134] When user A selects icon 42, in step S13, a selection screen is displayed for user A to select the format data to be used when generating their video from among the format data received along with the format video. Figure 9 is an example of the format data selection screen displayed on the user device 10a's display when icon 42 is selected. As shown in Figure 9, the format data selection screen 50 includes display elements corresponding to each format data acquired along with the format video. Specifically, the format data selection screen 50 includes a display element 51 corresponding to virtual space identification information, a display element 52 corresponding to music information, a display element 53 corresponding to motion information, a display element 54 corresponding to virtual camera setting information, and a display element 55 corresponding to object information. The format data selection screen 50 also includes a shooting button 56 for user A to take a picture to create a new format video. The format data selection screen 50 shown in Figure 9 is an example of a user interface for allowing the user to select the format data to be used to generate the format video. User A may also select the format data to be used in generating the video using a different user interface than the format data selection screen 50.
[0135] Display element 51 includes the name of the virtual space (Virtual Space A) identified by the virtual space identification information, which is one of the format data associated with the received video data; the location (coordinates) of the avatar within Virtual Space A; a thumbnail 51a of Virtual Space A; and a checkbox 51b for selecting Virtual Space A. When generating a format video using the virtual space identification information corresponding to display element A, the placement and orientation of the avatar within the virtual space identified by the virtual space identification information may be specified in the virtual space identification information, or they may be specified by format data different from the virtual space identification information. If the placement and orientation of the avatar within the virtual space are specified by format data different from the virtual space identification information, when selecting to use the virtual space identification information, a prompt may be displayed to select format data for specifying the placement and orientation of the avatar within the virtual space (e.g., coordinate information, avatar orientation information, and / or area identification information). When thumbnail 51a is selected, detailed information of Virtual Space A may be displayed. The detailed information for virtual space A may include the creator of virtual space A, its data capacity, a list of format videos generated using virtual space A to date, and other information about virtual space A. Display element 51 may also display the cost of using virtual space A. By checking checkbox 51b, user A can generate format videos using virtual space A.
[0136] Display element 52 includes the name of the song (Song A) identified by the song information, which is one of the format data associated with the received video data, an icon 52a for Song A, and a checkbox 52b for selecting Song A. When icon 52a is selected, a portion of Song A is output to the speaker of the user device 10a. Therefore, user A can listen to Song A by selecting icon 52a. When icon 52a is selected, detailed information about Song A may be displayed. The detailed information about Song A may include the artist name indicating the name of the artist performing Song A, the data size, a list of videos generated so far using Song A, and other information about Song A. Display element 52 may also display the price for using Song A. By checking checkbox 52b, user A can generate a format video using Song A.
[0137] Display element 53 includes the name of the motion (Motion A) identified by motion information, which is one of the format data associated with the received video data, a video display area 53a, and a checkbox 53b for selecting Motion A. The video display area 53a displays a video of an avatar performing the series of motions identified by Motion A. User A can view the series of motions of the avatar identified by Motion A through the video displayed in the video display area 53a. Detailed information about Motion A is associated with the video display area 53a, and when the video display area 53a is selected, the detailed information about Motion A may be displayed. The detailed information about Motion A may include the username of the user who created Motion A, a list of videos that have been generated so far using Motion A, and other information about Motion A. The username of the user who created Motion A may be obtained from the format data management information 25b. The username of the user who created Motion A may be associated with the received video data and sent from the server 20 to the user device 10a. Display element 53 may also display the price for using Motion A. By checking checkbox 53b, user A can generate a formatted video using motion A. When a formatted video is generated using motion A, user A's avatar can perform the same motion in the generated formatted video as the motion performed by avatar 31 in the received formatted video.
[0138] The display element 54 includes the name of the camera work (camera work A) identified in the virtual camera setting information, which is one of the format data associated with the received video data, a video display area 54a, and a checkbox 54b for selecting camera work A. The video display area 54a displays a video generated by filming a model avatar according to the camera work identified by camera work A. The model avatar is prepared by the operator of the virtual space. User A can understand the camera work identified by camera work A by the video displayed in the video display area 54a. Detailed information about camera work A is associated with the video display area 54a, and when the video display area 54a is selected, the detailed information about camera work A may be displayed. The detailed information about camera work A may include the username of the user who created camera work A, a list of videos generated so far using camera work A, and other information about camera work A. The username of the user who created camera work A may be obtained from the format data management information 25b. The username of the user who created camera work A may be associated with the received video data and sent from the server 20 to the user device 10a. Display element 54 may display the cost of using camera work A. By checking checkbox 54b, user A can generate a formatted video using camera work A. When a video is generated using camera work A, user A's avatar can be filmed in the generated video using the same camera work that is used to film avatar 31 in user C's formatted video data.
[0139] Display element 55 includes a selection button 55a for selecting an avatar item identified by object information, which is one of the format data associated with the received format video. Although checkboxes associated with avatar items are not shown in Figure 9, checkboxes may be provided for display element 55, just like with other format data. If a checkbox is provided for display element 55, when the user checks the checkbox, the avatar items worn by the avatar in the received format video of user C may be directly attached to user A's avatar, and a format video including user A's avatar with these avatar items attached may be generated. In other words, when the checkbox provided for display element 55 is checked, the set of avatar items (coordinate) worn by user C's avatar may be adopted as the set of avatar items (coordinate) worn by user A's avatar.
[0140] When the selection button 55a is selected, an item selection screen for selecting an avatar item is displayed on the user device 10a's display. Figure 10 shows an example of the item selection screen displayed in response to the selection of the selection button 55a. The item selection screen 60 shown in Figure 10 shows avatar items 61 to 65 that can be attached to User A's avatar. The item selection screen 60 may only display avatar items that are attached to avatar 31 included in the received format video. Of these avatar items 61 to 65, avatar item 61 that is attached to avatar 31 included in the received video data is surrounded by a border for emphasis. User A can select avatar item 61 or any other avatar item. The avatar items displayed on the item selection screen 60 may be only a portion of the avatar items that User A can select. By scrolling the area where avatar items 61 to 65 are displayed left or right or up or down, other avatar items may be displayed. The cost of using each avatar item may be displayed in association with each of the avatar items 61 to 65. When any avatar item is selected, the purchase button 66 becomes selectable. When the purchase button 66 is selected, the purchase of the selected avatar item is completed. User A can then equip the purchased avatar item to their own avatar.
[0141] In the example shown in Figure 9, checkboxes 51b, 52b, and 54b are checked. When the shooting button 56 is selected with checkboxes 51b, 52b, and 54b checked, in step S14, a formatted video for user A is generated using virtual space A, the placement location of the avatar set in virtual space A, music A, and camera work A. Specifically, user A's avatar is placed in the placement location in virtual space A, music A is specified as the music for playback, and video data is generated by shooting the field of view in the virtual space as defined by the camera work specified by camera work A. User A's avatar placed in virtual space A may be wearing avatar items purchased on the item selection screen 60.
[0142] In the example shown in Figure 9, since checkbox 53b is not selected, User A's avatar does not perform the motion identified by motion A during shooting in step S14. The motion of User A's avatar may be controlled based on motion data that represents User A's body and facial movements and expressions in chronological order, acquired by the user device 10a or other sensors. Motion data for identifying User A's avatar's motion may be acquired in real time when creating User A's formatted video after selecting the shooting button 56. In other words, after selecting the shooting button 56, changes in User A's movements and expressions may be detected by the user device 10a or sensors, motion data corresponding to these detected changes in User A's movements and expressions may be acquired, and User A's formatted video may be generated using this acquired motion data. User A's avatar, moving according to the motion data acquired in this way, may be displayed on the user device 10a. Therefore, User A can change their own movements and expressions while watching the movements of the avatar displayed on the user device 10a (while receiving feedback from the movements of the avatar displayed on the user device 10a). Since User A's avatar motion is uniquely set by User A, motion information (motion B) that identifies User A's avatar motion is set as format data (original format data for which User A is the creator) associated with the format video created in step S14 when the video data is generated in step S14 (i.e., when the virtual camera is used to shoot to generate the video data). This motion information is uploaded to the server 20 along with the format video in the subsequent step S15.
[0143] If a fee is set for the format data selected by User A, that fee will be charged to User A depending on the selection of the capture button 56. For example, in the example in Figure 9, virtual space A is selected, so selecting the capture button 56 will charge User A the fee for using virtual space A. If virtual space A is user-generated content (UGC) created by the user, part or all of the fee for virtual space A may be paid to the user who created virtual space A. Also, part or all of the fee for using motion A may be paid to the user who created motion A. Also, part or all of the fee for using camera work A may be paid to the user who created camera work A. Payment of the fee may be made in fiat currency, cryptocurrency, coins or tokens that circulate only within the services provided by server 20, and other digital assets.
[0144] The formatted video generated in step S14 is stored in storage in step S15. For example, the generated formatted video is uploaded to server 20 and stored in server 20's storage 15. Server 20 can store the formatted video received by user A in storage 25 as part of the formatted video management information 25a. User A's formatted video may be stored in a storage area allocated to user A within storage 25. User A may make the formatted video uploaded to server 20 public so that other users can view it, or private so that other users cannot view it. Server 20 may switch between public and private states in response to a request from the user. User A's formatted video may be stored in the storage 15 of user device 10a before being uploaded to server 20. User A can upload the formatted video stored in storage 15 to server 20 at a desired time.
[0145] Server 20 makes a formatted video viewable by other users when the formatted video is set to public. Specifically, Server 20 may add a formatted video to the formatted video list when the formatted video is uploaded in a public state, or when the setting of a formatted video is changed from private to public. In this case, when Server 20 receives a viewing request from a user terminal that has obtained the formatted video list, it can send the formatted video to the user device 10 that made the viewing request.
[0146] The formatted video generated in step S14 may be displayed on the display of the user device 10a. User A can view the video displayed on the display to confirm the formatted video. User A can view the video displayed on the display to decide whether or not to edit the formatted video. For example, User A can edit the formatted video by setting effects to be applied to the formatted video. User A can also edit the formatted video by inserting text or graphics into the formatted video. Effect information that identifies the effects set on the formatted video and insertion data information that identifies the insertion data inserted into the video may be associated with the formatted video generated in step S14 and uploaded to the server 20. This allows the effects identified by the effect information and the text and / or graphics identified by the insertion data information to be reflected in the formatted video when it is played on the user device of the viewing user.
[0147] The format corresponding to the format data selected in step S13 can be changed by editing after the format video is generated. For example, if song A is selected as the song to be played with the format video, the song played with the generated format video can be changed by selecting a different song B. Format data other than song information selected in step S13 can also be changed after the format video is generated.
[0148] Editing of formatted videos can be performed at any time after the formatted video has been generated. For example, editing of a formatted video may be performed (1) before uploading the formatted video to server 20, (2) after uploading to server 20 but before making it public (while it is private), or (3) after uploading to server 20 and setting it to public.
[0149] As described above, User A can generate their own formatted video by combining format data associated with the received formatted video. In this case, User A does not need to reproduce all of the multiple format data associated with the received formatted video; they can select the desired format data from among the multiple format data and use it to generate their own formatted video. Therefore, when User A generates their own formatted video, they can use format data associated with other formatted videos (in the example above, the formatted video generated by User C) for some formats (e.g., music), while for other formats (e.g., avatar motion), they can generate a format that reflects their own individuality rather than using the format of other formatted videos. This allows for both ease of video generation and the expression of creativity.
[0150] Furthermore, user A can utilize not only a single format (e.g., music) associated with the received formatted video, but also multiple formats, making it easy to set multiple formats when generating their own formatted videos.
[0151] As described above, when the shooting button 56 is selected, the shooting (generation) of the formatted video begins. In one embodiment, the shooting of the formatted video begins in response to the selection of the shooting button 56, and the resulting formatted video may be live-streamed.
[0152] When the capture button 56 is selected, User A's avatar is placed at a predetermined position in the virtual space specified by the format data, and the field of view of the virtual space, including User A's avatar, is captured. The video of the virtual space thus captured may be displayed on User A's user device 10a. At this point, the generation of the format video has not yet started. The video displayed on User A's user device 10a displays an icon (display element) for executing capture, and by selecting this icon, the capture of the format video may be started. This allows User A to check the video generated using the selected format data on their user device 10a before starting the generation of the format video. If User A decides that the video displayed on their user device 10a is suitable for generation as a format video, they can select the icon for executing capture to start the generation of the format video. When checking the display before the generation of the format video, the format data that was selected before selecting the capture button 56 can be changed to a different format data. After changing the format data, selecting the icon for executing capture allows User A to select a format that suits their preference before executing the capture of the format image.
[0153] The video processing method described above utilizes user-generated format data (original format data), providing users who generated that data with the opportunity to receive economic compensation. This incentivizes users to generate their own format data. Furthermore, increased user-generated format data leads to a richer array of available format data for video generation, making it more convenient to generate videos using that format data.
[0154] In one embodiment, formatted video can be generated without going through the format data selection screen 50. For example, icons 43 to 47 may be overlaid on the image 30 displayed on the display. Icons 43 to 47 are examples of selection elements that accept selection from the user. When icon 43 is selected, a list of formatted videos generated by filming virtual space A is displayed, similar to the formatted video of user C received by the user device 10a. Figure 11 shows an example of a video list screen 70 that includes a list of formatted videos generated by filming virtual space A. The video list screen 70 displays videos 71 to 76 generated by filming virtual space A, and a film button 77. When the film button 77 is selected on the video list screen 70, user A's avatar is placed at a predetermined position in virtual space A, and a formatted video is generated by filming the field of view of this virtual space A with a virtual camera. When generating a formatted video via the video list screen 70, it is not necessary to select format data other than virtual space identification information, so the formatted video can be generated with fewer operations.
[0155] When the icon 44 displayed on the image 30 shown on the display is selected, a list of format videos is displayed on the user device 10a that are set to play music A when the video is played, just like the format video received by the user device 10a. Figure 12 shows a video list screen 80 that includes a list of format videos set to play music A when the video is played. The video list screen 80 displays videos 81-86 that are set to play music A when the video is played, and a capture button 87. When the capture button 87 is selected on the video list screen 80, the generation of the format video begins. When the capture button 87 is selected, the virtual space is not specified, so the avatar may be placed in user A's home space. After the capture button 87 is selected, the user device 10a may display a list of candidate virtual spaces to be used when generating the format video. User A may select one virtual space from these candidate virtual spaces. The format video when the capture button 87 is selected may be generated by capturing the field of view of the virtual space selected in this way. When the format video generated in this way is played, music A is also played. When distributing videos via the video list screen 80, it becomes unnecessary to select format data other than music information, allowing you to start distributing formatted videos with fewer operations.
[0156] When icon 45 is selected, a video list screen is displayed on the user device 10a, which includes a list of format videos that use the same motion A as the received format video, and a capture button. When the capture button is selected on this video list screen, image data is generated that includes the avatar of user A, which is placed in user A's home space, and user A's avatar performing a series of motions specified by motion A in the home space. After the capture button is selected, the user device 10a may display a list of candidate virtual spaces to be used when generating the format video. User A may select one of these candidate virtual spaces. The format video when the capture button is selected may be generated by capturing the field of view of the virtual space selected in this way.
[0157] When icon 46 is selected, a video list screen is displayed on the user device 10a, which includes a list of videos using the same camera work A as the format video received, and a capture button. When the capture button is selected on this video list screen, the avatar is placed in user A's home space, and image data including user A's avatar is generated, captured in the home space using the camera work identified by camera work A. After the capture button is selected, the user device 10a may display a list of candidate virtual spaces to be used when generating the format video. User A may select one of these candidate virtual spaces. The format video when the capture button is selected may be generated by capturing the field of view of the virtual space selected in this way.
[0158] When icon 47 is selected, a video list screen is displayed on the user device 10a, which includes a list of videos containing avatars wearing the same avatar items as avatar 31, and a capture button. When the capture button is selected on this video list screen, the avatar wearing the avatar items is placed in user A's home space, and image data including this avatar is generated. If user A does not own the avatar items worn by avatar 31, the user device 10a may transition to a screen prompting the user to purchase the avatar items worn by avatar 31 when the capture button is selected. After the capture button is selected, the user device 10a may display a list of several virtual space candidates to be used when generating the format video. User A may select one of these virtual space candidates. The format video when the capture button is selected may be generated by capturing the field of view of the virtual space selected in this way.
[0159] In Image 30, multiple icons may be selectable from icons 43 to 47. When a search is performed with multiple icons selected, a video list screen may be displayed that includes a list of format videos in which the format data corresponding to the selected multiple icons is set, along with a shooting button. In this video list screen, for example, a list of format videos generated using the same virtual camera A and camera work A as the format video received by the user device 10a is displayed. In this way, by allowing users to select multiple icons from icons 43 to 47 and search for format videos generated using the format data corresponding to the selected multiple icons, it becomes easier to identify format videos that are close to the user A's preferences.
[0160] Video format uploaded to server 6, file number 20. When a formatted video is uploaded to server 20, server 20 stores the formatted video in storage 25 in either a public or private state. Server 20 may also switch the settings of the formatted video between public and private states in response to instructions from the user who uploaded the formatted video. Server 20 can send a list of formatted videos, including a list of formatted videos set to public, to the user device of a viewing user upon request from the viewing user. Figure 13 shows an example of a formatted video list. Figure 13 shows the formatted video list 90 displayed on the user device 10b of user B, who is a viewing user. The formatted video list 90 displays thumbnails of multiple formatted videos set to public. Of these, thumbnail 91 corresponds to formatted video V1 uploaded by user A according to the flow in Figure 7.
[0161] Each thumbnail contains a playback icon for playing the video in the format corresponding to that thumbnail. In the illustrated example, thumbnail 91 contains a playback icon 91a. User B can play the video in the format corresponding to thumbnail 91 by selecting the playback icon 91a. When the playback icon 91a is selected, the video in the format V1 is sent to user B's user device 10b. The video in the format V1 is played back on user device 10b. If the video in the format V1 is sent to user device 10b in the form of pre-rendered video data, the video data of the received video in the format V1 is processed by the rendering engine on user device 10b. In this case, the video in the format V1 includes all the data necessary for the rendering engine to generate the video.
[0162] When User B selects a portion of the thumbnail 91 other than the playback icon 91a, a list of format data associated with the format video V1 may be sent to User B's user device 10b. Figure 14 shows an example of a list of format data associated with the format video V1. Figure 14 schematically shows the format data list 100 displayed on User B's user device 10b. The format data list 100 displays each of the format data associated with the format video V1. As described above, the format video V1 generated by User A is created by using the virtual space A, music A, and camera work A associated with the format video distributed by User C, and by having User A's avatar perform a motion specified by motion B, which User A created independently. In the format data list 100 shown in Figure 14, User A's icon 101 is displayed to the left of the video name (format video V1) to indicate that the format video V1 was created by User A. Furthermore, to show the usage relationship of virtual space A, music A, motion B, and camera work A associated with format video V1, icons representing the users who used these format data are displayed in front of user A, associated with each of the format data. Since format video V1 uses virtual space A, music A, and camera work A associated with video data distributed by user C, icon 102 representing user C is displayed in the display area corresponding to virtual space A, music A, and camera work A to indicate that these format data were inherited from the format video distributed by user C. Music A was also associated with user C's video, similar to virtual space A and camera work A, but the original creator of music A is a different user from user C, so icon 103 representing the original creator of music A is displayed in the display area corresponding to music A.In other words, user C, represented by icon 102, used song A, created by user 103, when creating the formatted video. Therefore, to indicate this inheritance relationship, icon 103 is displayed in the display area for song A. Thus, the display area for formatted data displays an icon representing the user who initially created the formatted data (e.g., icon 103) and an icon representing the user who subsequently used that formatted data (e.g., icon 102), reflecting their usage relationship.
[0163] The format data list 100 allows viewers to understand the original creator of the format data associated with each video data, as well as the subsequent usage relationships (inheritance relationships). The format data list 100 may also display information indicating the original creator (e.g., an icon) without showing the usage relationships. Viewers can, for example, follow the user who first created their preferred format data, and receive notifications when that user creates new format data. The original creator of the format data can receive notifications when their format data is used by another user to generate format videos. On each user's terminal in the video processing system 1, a list of format videos generated using the format data (original format data) created by each user may be displayed for each original format. For example, if user A creates a first original format and a second original format, a list of format videos generated using the first original format (first usage video list) and a list of format videos generated using the first original format (second usage video list) may be displayed on user A's user device 10a.
[0164] The format data shown in Figure 14 may also be used by other users when creating format videos. By user B selecting the area corresponding to each format data, a list of format videos that use that format data may be generated. For example, if the display area corresponding to motion B is selected, a list of format videos generated using motion B may be generated, and this list of generated format videos may be displayed on user B's user device 10b.
[0165] 7. Setting whether to allow the use of format data. As described above, the format data associated with a formatted video is intended to be used by the user who receives the formatted video as a format for generating their own formatted videos. On the other hand, if the formatted video includes original format data, some users may not want that original format data to be used by other users. For this reason, the video processing system 1 may provide a function to set whether or not to allow other users to use the format data for each user. For example, the server 20 may create a list of original format data associated with users and allow the user to set whether or not to allow other users to use each format data in that list. The permission to use original format data may be set for each original format data. The permission to use original format data may also be set for each formatted video. For example, if user A generates a first formatted video and a second formatted video, and original format data A is associated with both, the use of original format data A associated with the first formatted video may be permitted, while the use of original format data A associated with the second formatted video may be denied.
[0166] Figure 15 shows an example of a list display screen 110 that includes a list of original format data associated with user A. The list display screen 110 is displayed on user A's user device 10a. The list display screen 110 displays four format data, format data A to D, and each has checkboxes 111 to 114. User A can prohibit the use of format data by checking the checkbox for the format data they do not want to use. In the illustrated example, checkbox 113 for format data C is checked, so users other than user A cannot use format data C when generating a video. For example, in a user device that receives a video associated with format data C, the format data corresponding to format data C can be grayed out and made unselectable on the format data selection screen 50. The format data corresponding to format data C does not necessarily have to be displayed on the format data selection screen 50.
[0167] 8. Start recording (generating) video in format. The flow shown in Figure 7 illustrates a scenario in which User A acquires a formatted video generated by another user (User C) and starts generating a formatted video by utilizing the format data associated with that formatted video. The video processing system 1 can start shooting a formatted video in a manner other than that shown in Figure 7. For example, it can store the shooting locations that have been designated as targets for formatted images by any user in each virtual space. When User A switches to shooting mode while exploring the virtual space using their avatar, a list of shooting locations within a predetermined distance from User A's avatar's position at the time of switching to shooting mode is displayed on User A's user device 10a. The list of shooting locations includes, in association with the shooting location, a formatted video shot at that location and the format data associated with that formatted video. By selecting one or more of the shooting locations and the format data associated with the formatted videos shot at those locations from the list, User A can generate a formatted video by using the selected format data to shoot the virtual space so that the selected shooting locations are included in the field of view of the virtual camera. This allows user A to easily identify the shooting location within the virtual space they are using, and furthermore, to easily generate video data by utilizing the format data associated with the format video shot at that location.
[0168] In one embodiment, a user may start recording a formatted video while the user is distributing a regular video. For example, an icon (display element) for starting to record a formatted video may be displayed on the user's device while the user is distributing a regular video. The user may start recording a formatted video while distributing a regular video by selecting the icon for starting to record a formatted video on their user terminal. The formatted video may be generated using the virtual space where the user's avatar is located during the distribution of the regular video. The formatted video thus generated may be generated or distributed in a different format than the regular video. For example, even if the regular video is distributed using a client rendering method, the formatted video may be generated or distributed using a video distribution method or a server rendering method.
[0169] In one embodiment, a user may generate a formatted video and formatted data using the video data used to distribute a regular video after the distribution of the regular video has ended. The formatted video may be a part of the regular video. For example, if a regular video is distributed for 60 minutes, the user may select a portion of the entire regular video (for example, one minute during the distribution) as the formatted video. The formatted data can be generated based on the video data of the selected portion of the regular video. For example, information that identifies the camera work included in the selected portion of the regular video can be set as formatted data (virtual camera setting information). The selection of a portion of the regular video may be made after the start of distribution of the regular video and before the end of distribution.
[0170] Users watching a regular video may use the video data of that regular video to record a video in a format that includes their own avatar, when their own avatar is displayed in that regular video.
[0171] 9. Note The video processing system 1 shown in Figure 1 is an example of a system to which the present invention can be applied, and the system to which the present invention can be applied is not limited to that shown in Figure 1. The video processing system 1 to which the present invention can be applied may not include some of the components shown in the figure. The video processing system 1 may include a cloud environment for distributed processing of processes that would otherwise be performed by the user device 10 or the server 20.
[0172] In the video processing system 1, there are no particular restrictions on the location of data storage. For example, various types of data that can be stored in storage 25 may be stored in storage or a database server that is physically separate from storage 25. In this specification, the data described as being stored in storage 25 may be stored in a single storage or distributed across multiple storages. In this specification and the claims, the term "storage" may refer to either a single storage or a collection of multiple storages, to the extent permitted by the context.
[0173] The embodiments of the present invention are not limited to those described above, and various modifications are possible without departing from the spirit of the invention. For example, some or all of the functions performed by processor 11 may be implemented by processor 21 or other processors not specified herein, as long as they do not depart from the spirit of the invention. Similarly, some or all of the functions performed by processor 21 may be implemented by processor 11 or other processors not specified herein, as long as they do not depart from the spirit of the invention. In Figure 4, processor 11 is shown as a single component, and in Figure 5, processor 21 is shown as a single component, but each of processor 11 and processor 21 may be a collection of multiple physically separate processors.
[0174] In this specification, a program or instructions contained in such a program described as being executed by processor 11 may be executed by a single processor or may be executed in a distributed manner by multiple processors. Furthermore, a program or instructions contained in such a program may be executed by one or more virtual processors. The description of processor 11 in this paragraph also applies to processor 21.
[0175] Programs executed on processor 11 and / or processor 21 may be stored on various types of non-transitory computer-readable media other than those shown in the diagram. Non-transitory computer-readable media include various types of tangible storage media. Examples of non-transitory computer-readable media include magnetic recording media (e.g., flexible disks, magnetic tapes, hard disk drives), magneto-optical recording media (e.g., magneto-optical disks), Compact Disc Read Only Memory (CD-ROM), CD-R, CD-R / W, and semiconductor memory (e.g., mask ROM, programmable ROM (PROM), erasable PROM (EPROM), flash ROM, random access memory (RAM)).
[0176] Even if it is stated that the processes and procedures described herein are performed by a single device, software, component, or module, such processes or procedures may be performed by multiple devices, multiple software programs, multiple components, and / or multiple modules. Similarly, even if it is stated that the data, tables, or databases described herein are stored in a single memory, such data, tables, or databases may be stored in multiple memories on a single device or distributed across multiple devices. Furthermore, the software and hardware elements described herein can also be implemented by integrating them into fewer components or by decomposing them into more components.
[0177] In the processing procedures described herein, particularly those described using flowcharts or sequence diagrams, it is possible to omit some of the steps constituting the processing procedure, add steps not explicitly stated as constituting the processing procedure, and / or change the order of such steps. Processing procedures with such omissions, additions, or changes in order are also included within the scope of the present invention, as long as they do not depart from the spirit of the invention.
[0178] The designations such as "First," "Second," and "Third" in this specification and the claims are used to identify components and do not necessarily limit their number, order, or content. Furthermore, the numbers used to identify components are used context by context, and a number used in one context does not necessarily indicate the same component in another context. Moreover, this does not prevent a component identified by one number from also performing the function of a component identified by another number.
[0179] The functions of the user device 10 and the server 20 may vary depending on the rendering method. For example, if a client rendering method is adopted, the user device 10 must have the function to render video data, but if a server rendering method is adopted, the user device 10 does not necessarily need to have the function to render video data. For example, at least a part of the function of the video playback unit 11b, which was described as a function of the user device 10a, may be implemented as a function of the server 20.
[0180] 10. Addendum This specification also discloses the technologies described in the following sections: [Note 1] A video processing method performed by one or more processors, One or more files used to generate a first video including a first avatar in a virtual space The process of obtaining format elements, One or more selected format elements from the one or more format elements mentioned above A step of generating a second video that includes a second avatar different from the first avatar, using the above method, A video processing method comprising the following features. [Note 2] The one or more format elements include virtual space identification information that identifies the virtual space. nothing, The video processing method according to claim 1. [Note 3] When the virtual space identification information is selected as one or more of the selection format elements In the second video, the second avatar is placed in the virtual space. The video processing method according to claim 2. [Note 4] The first video shows the first avatar captured by a virtual camera in the virtual space. Includes, The one or more format elements are the virtual space of the first avatar at the time of shooting. Includes coordinate information that identifies the position within the interval, The video processing method described in any one of the items from [Appendix 1] to [Appendix 3]. [Note 5] When the coordinate information is selected as one or more selection format elements, In the video, the second avatar is identified by the coordinate information within the virtual space. Placed in the position The video processing method described in [Appendix 4]. [Note 6] The first video shows the first avatar captured by a virtual camera in the virtual space. Includes, The one or more format elements mentioned above are the shooting settings information of the virtual camera at the time of shooting. including, The video processing method described in any one of the items from [Appendix 1] to [Appendix 5]. [Note 7] If the aforementioned shooting setting information is selected as one or more of the selected format elements, The second video is the second avatar filmed in the virtual space according to the aforementioned filming setting information. Generated to include a tar, The video processing method described in [Appendix 6]. [Note 8] The first video shows the first avatar captured by a virtual camera in the virtual space. Includes, The one or more format elements mentioned above are the motion of the first avatar during filming. Includes motion information that identifies The video processing method described in any one of the items from [Appendix 1] to [Appendix 7]. [Note 9] When the motion information is selected as one or more of the selected format elements, The second video is generated to include the second avatar that moves according to the motion information. can be The video processing method described in [Appendix 8]. [Note 10] The first video shows the first avatar captured by a virtual camera in the virtual space. Includes, The one or more format elements are associated with the first avatar at the time of shooting. Includes object information that identifies objects placed in the imaginary space, The video processing method described in any one of the items from [Appendix 1] to [Appendix 9]. [Note 11] When the object information is selected as one or more of the selected format elements In the second video, the second avatar is identified by the object information. Displayed in association with an object, The video processing method described in [Appendix 10]. [Note 12] The one or more format elements are played in the first video in accordance with the video. Includes song information that identifies the song, The video processing method described in any one of the items from [Appendix 1] to [Appendix 11]. [Note 13] When the aforementioned music information is selected as one or more of the selected format elements, 2 When viewing the video, the music information is identified in accordance with the video, including the second avatar. The song will be played. The video processing method described in [Appendix 12]. [Note 14] Along with the first video, one or more selections from the one or more format elements Display one or more selection elements for selecting formatting elements. The video processing method described in any one of the items from [Appendix 1] to [Appendix 13]. [Note 15] The one or more format elements include a first format element, The one or more selection elements are first selection elements for selecting the first format element. Contains the element, In response to detecting the selection of the first selection element, the first format element is used Display video information about the generated video. The video processing method described in [Appendix 14]. [Note 16] The one or more format elements include a first format element, The one or more selection elements are first selection elements for selecting the first format element. Contains the element, In response to detecting the selection of the first selection element, the first format element Corresponding to format information, The video processing method described in [Appendix 14]. [Note 17] The step of generating the second video involves at least one of the selected format elements. This includes the process of changing some parts to different format elements. The video processing method described in any one of the items from [Appendix 1] to [Appendix 16]. [Note 18] At least one of the format elements included in the one or more format elements is: The following will be available for the generation of the second video, subject to payment of consideration. The video processing method described in any one of the items from [Appendix 1] to [Appendix 17]. [Note 19] At least one of the format elements included in the one or more format elements is: It is tokenized as a non-fungible token. The video processing method described in any one of the items from [Appendix 1] to [Appendix 18]. [Note 20] A step of playing the first video generated based on the one or more format elements mentioned above. Furthermore, The video processing method described in any one of the items from [Appendix 1] to [Appendix 19]. [Note 21] An image processing method performed by one or more processors, A process of obtaining multiple format elements used to generate a first video including a first avatar, A step of generating a second video including a second avatar different from the first avatar, based on one or more selected format elements selected from the plurality of format elements, An image processing method comprising: [Note 22] Equipped with one or more processors, The one or more processors execute computer-processable instructions, Obtain one or more format elements used to generate a first video containing a first avatar in a virtual space, Using one or more selected format elements chosen from the aforementioned one or more format elements, a second video is generated that includes a second avatar different from the first avatar. Image processing system. [Note 23] One or more processors, A step of obtaining one or more format elements used to generate a first video including a first avatar in a virtual space, A step of generating a second video including a second avatar different from the first avatar using one or more selected format elements selected from the one or more format elements mentioned above, A video processing program that executes this process. [Explanation of Symbols]
[0181] 1. Video Processing System 10 User devices 11 processors 11a Virtual Space Display Unit 11b Video Playback Section 11c Format Data Selection Section 11d Video Generation Unit 11e Video Transmission Unit 11f Video Editing Department 20 servers 21 processors 21a Video Data Distribution Department 21b Format Data Trading Department 25 storage 25a Video Format Management Information
Claims
1. A video processing method performed by one or more processors, A step of setting one or more of the one or more format data set in the first video generated by capturing a virtual space based on the operation of the first user as the permitted format data, A step of generating a second video different from the first video using one or more selected format data selected from the one or more usage permission format data based on an operation from a second user, A video processing method comprising the following features.
2. The aforementioned one or more format data includes virtual space identification information that identifies a virtual space, The video processing method according to claim 1.
3. The aforementioned one or more format data includes music information that identifies the music played in conjunction with the video in the second video. The video processing method according to claim 1.
4. When the aforementioned music information is selected as one or more of the selected format data, the music identified by the aforementioned music information is played when the second video is viewed. The video processing method according to claim 3.
5. At least one of the one or more format data is the original format data determined by the first user when the first video is generated. The video processing method according to claim 1.
6. One or more of the aforementioned one or more format data are set as unavailable format data. When generating the second video, the use of the unavailable format data is prohibited. The video processing method according to claim 5.
7. Of the one or more format data, the original format data is set as the unavailable format data. The video processing method according to claim 6.
8. The unavailable format data is determined based on the first user's selection from among the one or more format data. The video processing method according to claim 6.
9. The one or more format data mentioned above are classified into original format data determined by the first user at the time of generating the first video and inherited format data inherited from other videos. A format data ID is assigned only to the original format data among the one or more format data mentioned above. The video processing method according to claim 1.
10. The aforementioned original format data is tokenized as non-fungible tokens. The video processing method according to claim 5.
11. The method further includes a step of sending a notification to the first user indicating that the original format data has been used, in response to the generation of the second video using the original format data. The video processing method according to claim 5.
12. The system further includes a step of charging the second user a fee for using the original format data. The video processing method according to claim 5.
13. Equipped with one or more processors, The one or more processors execute computer-processable instructions, One or more of the format data set in the first video generated by capturing a virtual space based on the operation of the first user are set as permitted format data. Based on an operation from a second user, a second video different from the first video is generated using one or more selected format data selected from the one or more permitted format data. Image processing system.
14. One or more processors, A step of setting one or more of the one or more format data set in the first video generated by capturing a virtual space based on the operation of the first user as the permitted format data, A step of generating a second video different from the first video using one or more selected format data selected from the one or more usage permission format data based on an operation from a second user, A video processing program that executes this process.
Citation Information
Patent Citations
Game system, game device, control program, and game control method
JP2017029509A