Information processing system, information processing method, and computer program
The information processing system efficiently projects user characteristics onto 2D animated videos, addressing the challenge of manual control and cost in existing technologies, enabling user immersion and cost-effective production of personalized animation.
Patent Information
- Application Number
- JP2024139236
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-20
- Publication Date
- 2026-03-05
AI Technical Summary
Existing technologies struggle to incorporate user characteristics into 2D animated videos efficiently and cost-effectively, as manual work is required to control characters that reflect user intentions, and the 2D world is not equivalent to the 3D real world, making it difficult to visualize user characters in animation.
An information processing system that acquires user features, projects them onto characters in a virtual space, and converts the scene into a 2D animated image, using a combination of user feature acquisition, scene setting, character projection, and video conversion processes to create a 2D animated video that reflects the user's characteristics and the creator's vision.
Enables easy inclusion of users in 2D animated videos, allowing multiple users to experience and immerse themselves in the animation world, reducing costs and simplifying the production process while maintaining the intended worldview.
Smart Images

Figure 2026036557000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology disclosed in this specification (hereinafter referred to as "the present disclosure") relates to an information processing system, an information processing method, and a computer program that perform processing related to animation or live-action video production. [Background technology]
[0002] By including users in animated videos, it is expected that user engagement with the videos will increase. However, since it requires manual work, it is costly and difficult to provide services to an unspecified number of users. In addition, the 2D animated world is not equivalent to the 3D real world, and it is difficult to control characters that represent users in the video.
[0003] For example, technologies have been proposed that incorporate three-dimensional model data of real objects into three-dimensional virtual spaces (see Patent Documents 1 to 3). Using this type of technology, it is possible to incorporate a user's CG (Computer Graphics) model into a three-dimensional space, manipulate the CG model in the three-dimensional space, and easily animate the movement of the CG model. However, because manipulation of a CG model in a three-dimensional space does not take into consideration manipulation within a 2D animation, it is difficult to visualize a user's character appearing in an animation that reflects the intentions of the user or creator. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2002-58045 [Patent Document 2] Japanese Patent Application Laid-Open No. 2010-92402 [Patent Document 3] Japanese Patent Application Publication No. 10-302085 [Non-patent literature]
[0005] [Non-Patent Document 1] "Lines used to express movement in TV animation" (https: / / www.jstage.jst.gojp / article / tvrsj / 21 / 3 / 21_447 / _pdf) [Non-patent document 2] CCEdit: Creative and Controllable Video Editing via Diffusion Models(arXiv:2309.16496v3 [cs.CV] 7 Apr 2024) Summary of the Invention [Problem to be solved by the invention]
[0006] An object of the present disclosure is to provide an information processing system, an information processing method, and a computer program that perform processing related to the generation of content including a character that projects a user. [Means for solving the problem]
[0007] The present disclosure has been made in consideration of the above-described problems, and a first aspect thereof is an information processing system that performs processing related to content generation, a user feature acquisition unit that acquires features related to the user; a scene setting information acquisition unit that acquires scene setting information related to the first content; a content generation unit that generates second content including a character projected as the user into the first content based on the characteristics of the user and setting information of the scene; The information processing system is provided with the above.
[0008] However, the term "system" used here refers to a logical collection of multiple devices (or functional modules that realize specific functions), regardless of whether each device or functional module is contained within a single housing. In other words, both a single device consisting of multiple parts or functional modules and a collection of multiple devices are considered "systems."
[0009] The first content is video content or animated video, and the second content is video content or animated video including a character that projects the user's characteristics. Also, the first content includes a voice of the character, and the second content is content including a voice of the character that projects the user's characteristics.
[0010] The first content is an original video content, and the scene setting information acquisition unit acquires the setting information based on a second input including a setting of the original work related to at least one of an environment, characters, and sounds in a scene to be visualized.
[0011] The content generation unit includes a character projection processing unit that projects characteristics related to the user onto characters appearing in the original video content, a scene reproduction processing unit that reproduces a scene in a virtual space based on the setting information, a character operation processing unit that moves the character in the virtual space, a video shooting control unit that shoots the virtual space in which the character moves with a virtual camera, and a video conversion processing unit that converts the video shot by the video shooting control unit into the second content consisting of video based on the assumptions of at least one of the creator of the original work or the user.
[0012] A second aspect of the present disclosure is an information processing method for performing processing related to content generation, comprising: a user feature acquisition step of acquiring features related to the user; a scene setting information acquisition step of acquiring scene setting information related to the first content; a content generation step of generating second content including a character projected as the user into the first content based on the characteristics of the user and setting information of the scene; The information processing method has the following features.
[0013] A third aspect of the present disclosure provides a computer program written in a computer-readable format to execute a process related to content generation on a computer, the computer comprising: a user feature acquisition unit that acquires features related to the user; a scene setting information acquisition unit that acquires scene setting information related to the first content; a content generation unit that generates second content including a character projected as the user into the first content based on the characteristics of the user and setting information of the scene; It is a computer program that functions as a
[0014] A computer program according to a third aspect of the present disclosure defines a computer program written in a computer-readable format to perform predetermined processing on a computer. The computer program can be provided to a computer capable of executing various program codes in a computer-readable format via a storage medium or communication medium, such as an optical disk, a magnetic disk, or a semiconductor memory, or via a communication medium such as a network. By installing the computer program according to the third aspect of the present disclosure on a computer via any of these media, a cooperative effect is exerted on the computer, and the same effects as those of the information processing system according to the first aspect of the present disclosure can be obtained. [Brief explanation of the drawings]
[0015] [Figure 1] FIG. 1 is a diagram showing the functional configuration of an information processing system 100 according to the present disclosure. [Figure 2]FIG. 2 is a diagram showing a functional configuration of an information processing system 200 according to the first embodiment of the present disclosure. [Figure 3] FIG. 3 is a diagram showing the operation of the scene setting processing unit 205. [Figure 4] FIG. 4 is a diagram showing an example of the configuration of the video conversion processing unit 208. [Figure 5] FIG. 5 is a diagram showing another example of the configuration of the video conversion processing unit 208. In FIG. [Figure 6] FIG. 6 is a diagram showing yet another example of the configuration of the video conversion processing unit 208. [Figure 7] FIG. 7 is a diagram showing the detailed configuration of the capture processing unit 201 and the character projection processing unit 202. [Figure 8] FIG. 8 is a diagram showing a capture process for handling multiple users. [Figure 9] FIG. 9 is a diagram showing the detailed configuration of the character operation processing unit 203. [Figure 10] FIG. 10 is a diagram showing the functional configuration of an information processing system 200 according to the third embodiment. [Figure 11] FIG. 11 is a diagram showing an example of a configuration for converting an acoustic signal. [Figure 12] FIG. 12 is a diagram showing another example of a configuration for converting an acoustic signal. [Figure 13] FIG. 13 is a diagram showing yet another example of a configuration for converting an acoustic signal. [Figure 14] FIG. 14 is a diagram showing an example of the hardware configuration of the information processing device 2000. [Figure 15] FIG. 15 is a diagram showing an example of a display screen of the video of the second output. [Figure 16] FIG. 16 is a diagram showing an example of a display screen of the video of the third output. [Figure 17] FIG. 17 is a diagram showing an example of a display screen of the video of the fourth output. [Figure 18] FIG. 18 is a diagram showing an example of a method for realizing an animation world that includes a user. [Figure 19] FIG. 19 is a diagram showing an example of a scene (battle scene) depicted in an anime. [Figure 20] FIG. 20 shows an example of a scene depicted in an animation (a scene in which a character stands on the surface of water). [Figure 21] FIG. 21 shows an example of an emotional expression (angry facial expression) of a character drawn in an animation. [Figure 22] FIG. 22 shows examples of emotional expressions (laughing expressions) of characters drawn in animation. [Figure 23] FIG. 23 shows an example of emotional expression (surprised facial expression) of a character drawn in an animation. [Figure 24] FIG. 24 shows an example of a scene depicted in an anime that stimulates the viewer's imagination (a scene in which a bicycle is riding down a slope at high speed). [Figure 25] FIG. 25 shows an example of a scene (a girl eating) that is depicted in an anime in a way that stimulates the viewer's imagination. [Figure 26] FIG. 26 shows an example of a scene (a dunk shot scene) that is depicted in an anime in a way that stimulates the viewer's imagination. DETAILED DESCRIPTION OF THE INVENTION
[0016] Hereinafter, embodiments of the present disclosure will be described in the following order with reference to the drawings.
[0017] A. Overview A-1. System Configuration A-2. Additional elements B. First Example B-1. Overall configuration B-2. Scene setting processing section B-3.Video conversion processing section B-3-1.Video conversion example (1) B-3-2.Video conversion example (2) B-3-3.Video conversion example (3) C. Second Example C-1. Capture processing section C-2. Character operation processing section D. Third Example E. Fourth Example E-1. When producing live-action footage E-2. When creating footage from different eras E-3. Use of generative AI technology F. Fifth Example F-1. Sound conversion example (1) F-2. Sound conversion example (2) F-3. Sound conversion example (3) F-4. Sound conversion example (4) G. System Features and Effects G-1. Main Features G-2. Additional Features (1) G-3. Additional Features (2) H. Configuration of information processing device
[0018] A. Overview By including the user in a specific 2D animated image and visualizing it, the user can experience the world of the 2D anime. For example, by entering the anime world, the user can meet and talk with the characters and participate in the story of the anime. In addition, by saving this as a 2D animated image, the service can be provided to an unspecified number of users.
[0019] FIG. 18 shows a schematic example of a method for realizing an animation world that includes a user.
[0020] First, in the 3D real world, the user's features such as appearance and voice are acquired. Specifically, the user's features are acquired by capturing an image of the user using a camera and recording the user's speech using a microphone.
[0021] Next, a three-dimensional virtual world scene containing the user is recreated. That is, the user is reconstructed in the virtual world and assigned to a character within the virtual world, and the user's characteristics are projected onto the character. The "character" referred to here refers to any character prepared in advance by the animation producer, such as a secondary character, enemy character, extra, or passerby that appears in the animation.
[0022] The scenes from the 3D virtual world thus recreated are then photographed with a certain angle of view and camerawork, and converted into animated images, thereby creating a 2D animated world. Note that in this specification, "filming" refers to the process of photographing a 3D virtual space, combining elements such as characters, backgrounds, and effects, with a virtual camera at a desired angle of view and camerawork, to create a 2D animated image.
[0023] The method of realizing animated images that include the user, as shown in Figure 18, has two main issues: 2D animation expression and 3DCG control, each of which is explained below.
[0024] 2D representation of anime: Overseas animation works, including those in the United States, are mainly produced using 3DCG. Characters and backgrounds are often depicted realistically. 3DCG production techniques include, for example, moving 3D modeled characters by operating a game controller in a 3D scene (virtual world). In Japan, on the other hand, techniques such as the use of 3DCG have been developed while taking advantage of the characteristics of hand-drawn animation. In addition to a cultural background that values the flatness of 2D expression, it can be said that techniques have been developed to allow for efficient production on a limited budget.
[0025] Special methods of expression are sometimes used to express the spatial positional relationships of characters, their depicted size, movements, emotions, etc. in 2D flat representations. For example, there is a method of expressing human movements and emotions using lines (see, for example, Non-Patent Document 1). Known by terms such as speed lines, flow lines, movement lines, and convergence lines, this method gives a sense of dynamism to still images by expressing the speed, direction, and trajectory of character movements with lines.
[0026] In addition, in 2D animation production, the characters and background are often drawn separately and then layered together to create the image.
[0027] In short, moving a character in a 2D world is not as easy as moving a character with a game controller in a 3DCG virtual world.
[0028] Difficulties in controlling 3DCG: The technique of controlling a character in a 3DCG virtual world by operating a game controller has already been realized in video games, etc. However, it is generally difficult to control the character's body freely, manipulate facial expressions and emotional expressions, or use tools skillfully.
[0029] Methods have also been developed that sense the movements of a user's body, limbs, etc., and reflect them in the movements of a character. For example, technologies have been proposed that incorporate a CG model of the user into a three-dimensional space and manipulate the CG model in the three-dimensional space (see Patent Documents 1 to 3). However, it is difficult to freely move a character's body, manipulate facial expressions and emotional expressions, or skillfully use tools.
[0030] The world of 2D animation is not equivalent to the 3D real world, and it is not easy to control and act out characters. Including users in 2D animation requires manual work, which is costly and makes it difficult to provide services to an unspecified number of users. In short, the operation of models in 3DCG is not designed to be performed within a 2D animation, so it is not suitable for visualizing user characters appearing in animation that reflects the intentions of the user or creator.
[0031] Therefore, in this disclosure, a simplified image obtained by shooting a scene in a 3D virtual world with a certain angle of view and camerawork is subjected to image conversion processing so as to be an animated image that preserves a desired worldview. This disclosure is a technology that includes a user in an image of a specific worldview (i.e., a 2D animated image) and visualizes it with simple operations and at low cost.
[0032] A-1. System Configuration 1 schematically shows the functional configuration of an information processing system 100 to which the present disclosure is applied. The illustrated information processing system 100 includes functional components such as a user feature acquisition unit 101, a scene setting information acquisition unit 102, and an animation image generation unit 103. The information processing system 100 can be constructed using one or more computer systems such as personal computers (PCs). Furthermore, each functional component can be realized by executing a computer program, but of course, it may also be configured by a dedicated hardware device.
[0033] The user feature acquisition unit 101 acquires features related to the user. Specifically, the user feature acquisition unit 101 acquires the features related to the user by taking pictures of the user's face and body with a camera and recording the user's voice with a microphone.
[0034] The scene setting information acquisition unit 102 acquires scene setting information for the video to be produced. The scene setting information includes, for example, environmental settings such as location, time, and weather for the video to be produced, settings for the characters such as the number of characters, the roles of each character (protagonist, protagonist's friend, etc.), and each character's face, body, clothing, movement, and voice, and settings for sounds such as the lines of the characters in the scene to be visualized, background sounds such as wind, and music played in the corresponding scene.
[0035] The animation image generation unit 103 generates a 2D animation image including a character projected as the user into the 2D animation to be produced, based on the user characteristics acquired by the user characteristic acquisition unit 101 and the scene setting information acquired by the scene setting information acquisition unit 102.
[0036] In the "character projection process," the animation image generation unit 103 assigns a character from one of the characters appearing in the animation image to be produced to the user, and projects the user's characteristics acquired by the user characteristic acquisition unit 101 onto the character. For example, in accordance with the user's operation using an input device such as a game controller, the animation image generation unit 103 can perform processing such as movements and speech of the character onto which the user's characteristics are projected in the virtual space.
[0037] Furthermore, as a "scene reproduction process," the animation image generation unit 103 constructs a three-dimensional virtual space that reflects the setting information of the scene acquired by the scene setting information acquisition unit 102, and as a "character operation process," it performs processing such as the movement and speech of the character onto which the user's characteristics are projected in the virtual space in accordance with the user's operation using an input device such as a game controller, thereby reproducing a scene in the 3D virtual space that reflects the setting information of the scene acquired by the scene setting information acquisition unit 102 and includes the character onto which the user's characteristics are projected.
[0038] Then, in a "video shooting process," the animation image generation unit 103 shoots the recreated 3D virtual space scene as a 2D animation image using a desired angle of view and camera work. Furthermore, in a "video conversion process," the animation image generation unit 103 converts the shot 2D animation image into an image with the worldview intended by the creator or user. The converted 2D animation image can be saved as data, making it possible to reuse it.
[0039] According to the information processing system 100, it is possible to easily include a user in a specific 2D animated image and visualize it. Therefore, by allowing a user to experience the world of 2D animation and saving it as an image, it is possible to provide a service to an unspecified number of users.
[0040] Furthermore, according to the information processing system 100, users can be involved in the production of video works such as animation and can immediately enjoy the video works that include them, thereby enabling them to enjoy the worldview of the video works more deeply.
[0041] It should be noted that the effects described in this specification are merely examples, and the effects brought about by the present disclosure are not limited thereto. Furthermore, the present disclosure may have additional effects in addition to the effects described above. Further objects, features, and advantages of the present disclosure will become apparent from the following detailed description of the embodiments and the accompanying drawings.
[0042] A-2. Additional elements The information processing system 100 can improve the user's experience value by adding the following elements.
[0043] The user feature acquisition unit 101 is equipped with a multi-view camera and multiple microphones, allowing it to capture multiple people simultaneously. This allows the information processing system 100 to provide a service that targets multiple users at the same time and visualizes the users in a specific 2D animated video.
[0044] The animation image generation unit 103 may be provided with a display device such as a head-mounted display for the "character operation processing." The animation image generation unit 103 may also be provided with means for providing feedback to the user for the "character operation processing," such as a speaker or hubotics.
[0045] In the "character projection process," the animation image generation unit 103 may estimate the user's body movements and emotions and reflect them in the character in the virtual space.
[0046] In the "video shooting process," the animation video generation unit 103 may set up a virtual camera in a virtual space and operate the camera to recreate the scene to be shot in the virtual space, convert the recreated scene into video, and record the video.
[0047] In the "video conversion process," the animation video generation unit 103 may convert the video captured by the video capture control unit 104 into an animation video (or a live-action video, or an 80's-style video).
[0048] In the "image conversion process," the animation image generation unit 103 may convert the image into an animation image based on an original worldview by reflecting the animation characters, background design, story, and the like.
[0049] B. First Example B-1. Overall configuration FIG. 2 shows the functional configuration of an information processing system 200 according to a first embodiment of the present disclosure. The information processing system 200 includes functional components such as a capture processing unit 201, a character projection processing unit 202, a character operation processing unit 203, a video shooting instruction unit 204, a scene setting processing unit 205, a scene reproduction processing unit 206, a video shooting control unit 207, a video conversion processing unit 208, a converted video saving processing unit 209, a scene output unit 210, a captured video output unit 211, and a converted video output unit 212. The information processing system 200 can be constructed using one or more computer systems such as personal computers (PCs). Each functional component can be realized by executing a computer program, but may also be configured as a dedicated hardware device. Each functional component will be described below.
[0050] The capture processing unit 201 receives information about the user, such as the user's face, body, and voice, as a first input, and extracts features about the user by taking a picture with a camera or recording with a microphone.
[0051] The character projection processing unit 202 projects and reflects the user's features extracted by the capture processing unit onto a virtual character in the information processing system 200. The virtual character is a person, animal, or other character configured in a virtual space, and is configured so that the face, body, limbs, facial expression, voice, and the like can be controlled by the character operation processing unit 203. For example, the virtual character is any character prepared in advance by an animation producer, such as a sub-character, enemy character, extra, or passerby that appears in an animation.
[0052] Character operation processing unit 203 performs processing for operating a virtual character projecting the user's features in a virtual space based on a third input from a controller that the user uses to control the virtual character's body movements and voice. For example, the third input can be acquired using a game controller, a motion capture device attached to the user's body, a microphone, a wearable device, or other devices. Character operation processing unit 203 extracts the user's features (facial features, skeletal features, clothing features, voice timbre features, etc.) from the third input and outputs them to scene reproduction processing unit 206 and, if necessary, to video conversion processing unit 208.
[0053] In accordance with the video shooting control unit 207, the video shooting instruction unit 204 outputs a first output to the user as instructions regarding the virtual character's acting, etc. The first output is, for example, audio guidance that conveys instructions to the user such as "Start shooting," "Stop shooting," "Please act," and "Please speak your lines." The first output may also be in a modality other than audio, such as text, images, or haptics. Based on the first output, the user can use a controller or the like to input a third input to the character operation processing unit 203 for controlling the virtual character's body movements and voice.
[0054] The scene setting processing unit 205 performs processing to reflect the second input consisting of scene setting information related to the video to be produced in the 3D virtual space that is to be constructed by the information processing system 200.
[0055] The second input includes information about the scene settings of the original anime work that serves as a reference, which is used to precisely reflect the design, worldview, story, etc. envisioned by the video creator in the virtual space. The second input includes, for example, environmental settings such as the location, time, and weather to be set for the video to be produced, character settings such as the number of characters, the role of each character (protagonist, protagonist's friend, etc.), and each character's face, body, clothing, movements, and voice, as well as sound settings such as the lines of the characters in the scenes to be visualized, background sounds such as wind, and music played in the corresponding scenes.
[0056] Scene reproduction processing unit 206 performs processing to reproduce, in a three-dimensional virtual space, the scenes set by scene setting processing unit 205, such as the movements and lines of the characters and changes in the environment, in accordance with instructions from video shooting control unit 207. Also, scene reproduction processing unit 206 performs processing to reproduce, in a three-dimensional virtual space, the movements and lines of the virtual characters that reflect the user's operations, in accordance with instructions from character operation processing unit 203. Through these processes by scene reproduction processing unit 206, the scenes to be visualized are reproduced in a three-dimensional virtual space.
[0057] The video shooting control unit 207 is a functional component that controls the entire video shooting. The video shooting control unit 207 first controls the scene reproduction processing unit 206 so as to reproduce the scene to be shot in a three-dimensional virtual space, and then visualizes the reproduced scene. To visualize the scene, the video shooting control unit 207 specifically performs processing to create a 2D image (simple image) by shooting the movements of characters and the utterance of lines in the three-dimensional virtual space with a virtual camera at a desired angle of view and camera work. The video shooting control unit 207 also provides instructions on starting, ending, and redoing shooting, acting, etc. In other words, the video shooting control unit 207 gives instructions to the user regarding the operation of the virtual character via the video shooting instruction unit 204.
[0058] The video conversion processing unit 208 inputs video data (simple video) captured by the video shooting control unit 207 and converts it into video data with a different worldview, for example, video data envisioned by at least one of the creator and the user. The video conversion processing unit 208 converts the worldview of the simple video captured by the video shooting control unit 207 so as to correct any deviations from the scenes desired to be visualized as animation. For example, in the case of animation video, the video conversion processing unit 208 performs processing such as adding emotional expressions desired by the user to the character's face, expressing the character's body movements using lines such as speed lines, flow lines, movement lines, and convergence lines, and applying texture to the character's face and body to fit the animation worldview desired by the creator or user.
[0059] The converted video data storage processor 209 stores the video data converted by the video conversion processor 208. There are no particular limitations on the destination where the data is stored. By storing the video data, it becomes possible to reuse it.
[0060] The scene output unit 210 outputs, as a second output, an image of the situation in which the scene is reproduced in the three-dimensional virtual space by the scene reproduction processing unit 206 to a monitor screen. The second output is an image of the scene reproduced in the virtual space (for example, a first-person image of a virtual character), and is expressed, for example, in 3DCG. Note that the monitor is not shown in FIG. 2. The monitor may be a device included in the information processing system 200, or may be an external device of the information processing system 200 (the same applies hereinafter).
[0061] The captured video output unit 211 outputs the 2D video of the scene in the virtual space captured by the video capture control unit 207 to the monitor screen (same as above) as a third output. The third output is expressed as a video of the 3D virtual space captured by a virtual camera with a desired angle of view and camerawork.
[0062] The converted video output unit 212 outputs the 2D video converted into a different worldview by the video conversion processing unit 208 to a monitor screen (same as above) as a fourth output. The fourth output is the same final video as that saved in the converted video saving processing unit 209, and is expressed as, for example, an animated video.
[0063] The second output is, for example, a three-dimensional virtual space viewed from the viewpoint of a virtual character. The third output is a simplified image of the virtual character in the three-dimensional virtual space captured by a virtual camera. The fourth output is an image obtained by converting the worldview of the simplified image, which is the third output, to correct any deviations from the scene to be visualized as an animation. FIG. 15 shows, as an example of the second output, a display screen of an image of a virtual space viewed from above in a 3D scene in which a virtual character as a soccer player is taking a penalty kick. FIG. 16 shows, as an example of the third output, a simplified image of the virtual character (i.e., the soccer player taking the penalty kick) in the virtual space shown in FIG. 15. FIG. 17 shows a 2D image, which is the fourth output, obtained by converting the worldview of the simplified image shown in FIG. 16 to correct any deviations from the scene to be visualized as an animation. The final video shown in Figure 17 adds the user's desired emotional expression to the character's face and expresses the character's body movements using speed lines, etc., resulting in the animated video that the creator and user envisioned.
[0064] While checking the 3D virtual space image of the reproduced scene, which is the second output, the user follows the shooting instructions, which are the first output, and inputs a second input to the information processing system 200 (character operation processing unit 203) using a controller or the like to control the body movements and sound of the virtual character. In addition, the user can check the shot image of the scene in the virtual space from the third output, and can check the final 2D animation image from the fourth output.
[0065] In this way, information processing system 200 allows a user to operate a virtual character in accordance with video shooting instructions while checking a 3DCG scene recreated in a virtual space, a 2D video of the scene in the virtual space, and a 2D video with a converted worldview. Therefore, the user can enter the filming of a video through the virtual character in the virtual space and act, thereby experiencing the worldview of the video.
[0066] Furthermore, according to the information processing system 200, the user can save and reuse the video as a final product that reflects the user's performance through the virtual character. When the information processing system 200 produces 2D animation video, the user can immerse themselves in the worldview of the 2D animation and experience it, and furthermore, can easily save and reuse the video.
[0067] B-2. Scene setting processing section The second input to the scene setting processor 205 is the character settings, costume and accessory settings, background and other art settings that are generally defined in the production of animation. Generally, many people are involved in the production of animation, so the following settings are shared for characters, backgrounds, etc. in order to share the image and worldview of the work. The second input includes the following setting contents:
[0068] Character Settings: Character settings include definitions of the character's front and side profile, height, physique, eyes, nose, mouth, hair style, and facial expressions. Specific details are also defined, such as age, gender, birthplace, place of residence, occupation, hobbies, preferences, special skills, speaking style, personality, experience, family structure, home environment, friendships, lifestyle, and work environment. Also, settings are made for the character's speaking characteristics, such as voice quality, dialect, catchphrases, and pronouns. The character's clothing, including its design and color, is also defined. Furthermore, the design and color of accessories used by the character, such as bags, hats, shoes, watches, glasses, weapons, armor, and vehicles, are also defined.
[0069] Background settings: The background that appears in the work is set as either outdoors or indoors, and if outdoors, images of the location and the surrounding sky, sea, mountains, trees, buildings, etc., design and colors are set; if indoors, images of the room structure, interior, lighting, tableware, home appliances, food, etc. are set.
[0070] In addition to the concept, story, script, scenario, etc. of the anime work, the depiction of each scene is further determined by the characters, their lines, background, surrounding circumstances, environmental sounds, etc.
[0071] Based on the settings described above, the scene setting processor 205 sets the layout for each scene, including the composition, the relative positions of the characters and background, the character's movements, and the angle of view when visualized. Then, based on this layout, a character animation is created using a series of still images representing the start and end of the character's movements and the movements that complement each other. Here, each still image is called an original image. By depicting these original images as continuous movements, a moving image of the animated character is created. Furthermore, by superimposing the corresponding background image on this moving image, an animated video is created.
[0072] FIG. 3 illustrates the operation of the scene setting processor 205. The scene setting processor 205 sets detailed setting information for a scene based on the second input described above. Specifically, the second input includes image information indicating each setting, linguistic information explaining the setting, linguistic information such as the script and dialogue, audio information of the dialogue, background sound information, and music information played in the scene. The second input also includes video data actually created in the animation work. Based on the setting information included in the second input, the scene setting processor 205 creates a 3D model of the background in the virtual space, a 3D model and motion of NPCs (non-plahyer characters), and audio data such as environmental sounds and background music, which are used in the subsequent scene reproduction processor 206. The specific processing means of the scene setting processor 205 can be a deep-learned model ({image, text}-to-{3D model, foley, music} model) that generates 3D models, sound effects, and music from input data such as images and text, or the creator can manually create them.
[0073] B-3.Video conversion processing section In this section B-3, an embodiment of the video conversion processing unit 208 will be described.
[0074] First, let me explain the diverse expressions used in anime. For example, anime often depicts scenes that would be difficult to recreate in the real world, such as a battle scene involving flying through the air (see Figure 19) or a scene in which a character stands on the surface of water (see Figure 20). Anime also expresses emotions such as joy, anger, sadness, and happiness in a variety of ways, using facial expressions, the whole body, and the lines and backgrounds unique to anime (as examples of how characters express emotions in anime, Figure 21 shows an angry expression, Figure 22 a smiling expression, and Figure 23 a surprised expression). Anime also depicts scenes that stimulate the viewer's imagination, such as a scene in which a person rides a bicycle down a hill at high speed (see Figure 24), a scene in which a girl eats with a knife and fork (see Figure 25), and a scene of sports using a ball, such as a dunk shot (see Figure 26). In animation, a method of expressing human movements and emotions using lines such as speed lines, flow lines, movement lines, and convergence lines (for example, see Non-Patent Document 1) is sometimes used.
[0075] On the other hand, video games and other games have made it possible to control characters in 3D computer graphics virtual spaces using controllers. Using a controller, it is possible to control characters by moving around the space, picking up items, and talking to other characters. However, it is difficult to control characters by freely moving their bodies, changing their facial expressions, controlling their emotional expressions, skillfully using vehicles and tools, or skillfully playing sports or eating. While assigning buttons to controllers allows characters to perform techniques and handle weapons, these operations are unusual for the average user, making it difficult to control the character's movements.
[0076] In VR games, controllers have been developed that sense the user's body and hand movements and reflect them in the character's movements. However, as with video games, it is difficult to freely control the character's body, express emotions, or operate tools and vehicles.
[0077] Furthermore, by using dedicated motion capture or glove-type controllers, it is possible to sense the movements of the user's body and hands with high precision, but it is difficult for the user to act out sports scenes that exceed the user's abilities (for example, dunking) in the real world and provide the third input to the character operation processing unit 203.
[0078] In short, when a user controls a 3DCG character using their own body or a controller, it is difficult to operate a character that corresponds to the above-mentioned scenes depicted in the anime (see Figures 19 to 26) due to limitations on the user's body or controller operation.
[0079] This can also be said to be an issue with the character operation processing unit 203. Because this issue remains unresolved, the scene reproduced by the scene reproduction processing unit 206 through the movements, lines, and environmental changes of characters in the virtual space will deviate from the scene desired to be visualized as an animation. In other words, the expression of character movements, facial expressions, cooperative actions with other characters, and the use of tools will be limited in the virtual space. As a result, the video captured in the virtual space by the video capture control unit 207 will include character movements and facial expressions that cannot be expressed due to the limitations of the user's body or the operation of the controller. In other words, the video captured by the video capture control unit 207 will deviate from the scene desired to be visualized as an animation.
[0080] Therefore, in information processing system 200, image conversion processing unit 208 converts the image output from image capture control unit 207 into a desired animated image, and corrects any deviations due to the user's physical limitations or limitations in operating the controller. Therefore, in the image corrected by image conversion processing unit 208, the virtual character operated by the user reflects the relationship between the positions and sizes of the characters, their relationship with the background image, their emotional expressions, their movements, etc.
[0081] B-3-1.Video conversion example (1) In video conversion example (1), the video conversion processing unit 208 converts the features of a character in an existing (e.g., original) animated video into the features of a user. FIG. 4 shows an example configuration of the video conversion processing unit 208. In the illustrated example, the video conversion processing unit 208 is made up of a video-to-video converter 401. The video-to-video converter 401 is a functional component that uses AI technology for generating video, and specifically, is made up of a trained model trained based on a machine learning method such as deep learning.
[0082] The video-to-video converter 401 inputs video (original scenes) of the original 2D animation work as a reference, and also inputs features of the user's face, bone structure, clothing, voice quality, etc. extracted by the capture processing unit 201 and character operation processing unit 203 as auxiliary data.
[0083] However, it is assumed that the video of the original anime work (original scene) that serves as a reference is supplied not from the video shooting control unit 207 but, for example, from the scene reproduction processing unit 206. Furthermore, the capture processing unit 201 extracts the user's features prior to video shooting, and the character operation processing unit 203 extracts the user's features in parallel with video shooting. It is also assumed that the user's features extracted by the capture processing unit 201 will be mixed with the features of a preset character (see FIG. 7, which will be described later), and the features obtained as a result of the mixing will also be referred to here as "user's features."
[0084] Then, the video-to-video converter 401 generates a fourth output consisting of 2D video converted into a different worldview by replacing the partial video related to the character that reflects the user's characteristics with the user's facial features, body features, movement features, voice features, etc.
[0085] B-3-2.Video conversion example (2) In video conversion example (2), video conversion processing unit 208 converts video captured by video capture control unit 207 into video with a different worldview, such as 2D animation video. Fig. 5 shows another configuration example of video conversion processing unit 208. In the illustrated example, video conversion processing unit 208 uses video-to-video converter 501 to convert the captured video output from video capture control unit 207 into video that reflects the characteristics of an animation work.
[0086] The scene reproduction processing unit 206 reproduces a scene in a virtual space, depicting the characters and environment using 3DCG. The video shooting control unit 207 then uses a virtual camera to capture the scene in the virtual space reproduced using 3DCG at a desired angle of view and camerawork, creating a simple video in a 3DCG style. The video conversion processing unit 208 converts the style of the simple video from the video shooting control unit 207 into a high-quality video that reflects the design and worldview envisioned by the creator. The high-quality video referred to here corresponds to, for example, an animated video produced using cel animation.
[0087] The video-to-video converter 501 receives the simplified video from the shooting control unit 207 as input, and also receives, as auxiliary data, feature quantities other than the video obtained during shooting by the video shooting control unit 207, specifically feature quantities such as skeletal feature quantities and emotional feature quantities of characters in the virtual space. The video-to-video converter 501 then converts the texture of the video input from the video shooting control unit 207 (video style conversion) while retaining the feature quantities input as auxiliary data, and also corrects the video in accordance with specifications in the text prompt. Note that extraction of feature quantities other than the video as auxiliary data may be performed by either the video shooting control unit 207 or the video conversion processing unit 208.
[0088] For example, while checking the video converted by the video-to-video converter 501, if the user finds that it deviates from the scene they want to visualize as an animation, they input a text prompt (e.g., "Move your body more") to the video-to-video converter 501 to encourage improvement. However, the prompt given to the video-to-video converter 501 may be in a modality other than text. By repeatedly checking the output video and receiving prompts to improve the video, an animation video containing the expression the user desires can be created.
[0089] B-3-3.Video conversion example (3) In video conversion example (3), video conversion processing unit 208 converts video captured by video capture control unit 207 into video with a different worldview, such as 2D animation video. FIG. 6 shows yet another configuration example of video conversion processing unit 208. In the example shown, video conversion processing unit 208 uses a video-to-video converter 601 to convert video output from video capture control unit 207 into video that reflects the characteristics of an animation work. Video conversion example (3) shown in FIG. 6 can also be said to be a method that combines the above video conversion examples (1) and (2).
[0090] First, feature quantities of a target existing character, such as skeletal feature quantities, are extracted from the video of the original animated work that serves as a reference. This feature extraction process may be performed by feature extraction unit 704 (described below) based on a second input related to scene setting information of the original animated work. Then, scene reproduction processing unit 206 replaces the feature quantities of the virtual character that reflect the user's operation with the feature quantities extracted from the original animated work, and drives a 3D model of the virtual character.
[0091] The video shooting control unit 207 creates a simple 2D video by shooting the scene in the 3DCG virtual space reproduced by the scene reproduction processing unit 206 with a virtual camera at a desired angle of view and camerawork. Then, the video conversion processing unit 208 converts the simple video from the shooting control unit 207 into a high-quality video that reflects the design and worldview envisioned by the creator.
[0092] The video-to-video converter 601 inputs the video shot by the video shooting control unit 207, and also inputs feature amounts extracted from the original animation work that serves as a reference as auxiliary data, and converts the texture of the video while preserving the feature amounts from the original animation work that have been input as auxiliary data. For example, it converts the style to high-quality video equivalent to animation video produced using cel animation.
[0093] According to the technique of video conversion example (3), it is possible to construct an animated video that realizes the movements of a virtual character in a virtual space that cannot be reproduced from a third input using the user's own body or a controller. For example, if an original animation work that serves as a reference has a scene in which a character flies in the sky, by extracting skeletal features of a flying character from the original animation work and driving a 3D model of the user's virtual character based on those skeletal features, it is possible to construct a 3D scene in which a 3D model that reflects the user's features is flying in the sky.
[0094] Although not shown in Fig. 6, the video-to-video converter 601 may further input a prompt to prompt the user to improve the scene to be animated. As in the video conversion example (2) shown in Fig. 5, by repeatedly checking the output video and receiving the prompts to improve the video, it becomes possible to create an animated video that includes the expression desired by the user.
[0095] The above video conversion examples (1) to (3) involve a trade-off between fidelity to the user's movements and fidelity to the desired animation image. The former contributes to the user's sense of ownership and agency, while the latter contributes to a sense of immersion in the world of the work. Therefore, by using video conversion examples (1) to (3) for each individual section of a scene, it may be possible to maximize the user experience value of both agency and immersion. Specifically, the information processing system 200 may operate in such a way that video conversion example (1) is used for sections that deviate significantly from the scene the user wants to visualize, and video conversion example (2) is used for sections that deviate less significantly.
[0096] As a video-to-video conversion method by the video-to-video converter used in each of the above video conversion examples (1) to (3), for example, the technology described in Non-Patent Document 2 can be applied.
[0097] C. Second Example In this section C, a second embodiment will be described, which allows the information processing system 200 to better project the characteristics of the user and to provide services to a plurality of users.
[0098] C-1. Capture processing section and character operation processing section FIG. 7 shows a detailed configuration related to the capture processing unit 201 and the character projection processing unit 202, which are designed to allow the user's characteristics to be more accurately projected onto the virtual character.
[0099] The capture processing unit 201 includes a camera 701, a microphone 702, and a user feature extraction unit 703. The camera 701 captures images of the user's face, body, and the like.
[0100] The camera 701 may be, for example, a multi-view camera that captures images of the user from multiple angles. The microphone 702 records short utterances from the user. The user feature extraction unit 703 extracts feature quantities related to the user's face, bone structure, clothing, voice quality, etc. from the images captured by the camera 701 and the audio recorded by the microphone 702.
[0101] In addition, if the first input includes information about the user other than the user's face, body, or voice, such as the user's biometric information, the capture processing unit 201 may further include a sensor for detecting such information, and the user feature extraction unit 703 may further extract user features based on the biometric information, etc.
[0102] The virtual character is any existing character prepared in advance by an animation producer, etc., such as a sub-character, enemy character, extra, passerby, etc. that appears in the original animation that serves as a reference. The information processing system 200 further includes a feature extraction unit 704 and a feature mixing unit 705 to mix the user's feature amount with settings related to the virtual character's face, body, clothing, movement, voice, etc., included in the scene setting information as the second input, and project the result onto the virtual character.
[0103] The feature extraction unit 704 extracts feature quantities related to the face, bone structure, clothing, voice quality, etc. of the target existing character based on a second input related to the original animation that serves as a reference. The feature extraction unit 704 may be located in, for example, either the scene setting processing unit 205 or the scene reproduction processing unit 206, or may be located independently of the scene setting processing unit 205 and the scene reproduction processing unit 206.
[0104] Then, feature mixing unit 705 mixes, in multiple ratios, each feature of the user extracted by user feature extraction unit 703 and the feature of the existing character extracted by feature extraction unit 704. Feature mixing unit 705 may be located in either capture processing unit 201 or character projection processing unit 202, or may be located independently of capture processing unit 201 and character projection processing unit 202.
[0105] The character projection processing unit 202 includes a 3D model synthesis unit 706 , a voice synthesis unit 707 , a 3D model presentation unit 708 , a voice presentation unit 709 , and a 3D model and voice selection unit 710 .
[0106] The 3D model synthesis unit 706 synthesizes a 3D model of a virtual character for each feature, such as face, bone structure, and clothing, mixed at various ratios by the feature mixing unit 705. Furthermore, the voice synthesis unit 707 synthesizes a voice for the virtual character for each voice timbre feature mixed at various ratios by the feature mixing unit 705. The synthesized voice may be a voice corresponding to a short text. In this way, a 3D model and voice of a virtual character can be generated by mixing the user's features with an existing character at a desired ratio.
[0107] The 3D model presentation unit 708 presents to the user on a monitor screen, etc., a 3D model of the virtual character synthesized for each of the face, bone structure, and clothing features mixed at each ratio. The audio presentation unit 709 presents to the user the voice of the virtual character synthesized for each of the voice timbre features mixed at each ratio.
[0108] The fourth input is a user selection. For example, the fourth input is a user selection using a controller or other input device. The user selects a 3D model and voice that are preferred by the user from among the 3D models of virtual characters for each mixing ratio presented by the 3D model presentation unit 708 and the synthesized voices of virtual characters for each mixing ratio presented by the voice presentation unit 709. Then, based on the fourth input, the 3D model and voice selection unit 710 outputs the 3D model and voice selected by the user to the subsequent character operation processing unit 203 as a virtual character that projects the user's characteristics.
[0109] According to the configuration of the capture processing unit 201 and the character projection processing unit 202 shown in FIG. 7, the information processing system 200 can more effectively project the user's characteristics onto the virtual character.
[0110] By using the above-described process, it is possible to generate a virtual character that combines the user's own characteristics with the characteristics of a character in the original video that serves as a reference, in the proportion desired by the user. This makes it possible to generate virtual characters that meet the needs of various users, such as users who want to project themselves into the worldview of the video they are creating, or users who want to become a character that strongly reflects the worldview of the video work (as set by the creator).
[0111] Next, the capture process for handling multiple users will be described with reference to FIG.
[0112] The capture processing unit 201 is equipped with a multi-view camera as the camera 701 and multiple microphones as the microphones 702, thereby enabling simultaneous capture of multiple users. Then, the user feature extraction unit 703 extracts feature quantities related to the face, bone structure, clothing, voice quality, etc. of each user from the captured images and audio.
[0113] On the other hand, the feature extraction unit 704 extracts feature quantities relating to the face, bone structure, clothing, voice quality, etc. of multiple characters to be the target of feature projection for each user, based on a second input relating to the original animation that serves as a reference.
[0114] Then, the feature mixing unit 705 mixes the user feature amount extracted by the user feature extraction unit 703 and the character feature amount extracted by the feature extraction unit 704 in multiple ratios for each character corresponding to the user.
[0115] For each character corresponding to a user, the 3D model synthesis unit 706 synthesizes a 3D model of a virtual character consisting of features such as face, bone structure, and clothing mixed in various proportions by the feature mixing unit 705. Furthermore, for each character corresponding to a user, the voice synthesis unit 707 synthesizes a voice of the virtual character having voice timbre features mixed in various proportions by the feature mixing unit 705. The synthesized voice may be a voice corresponding to a short text. In this way, a 3D model and voice of a virtual character can be generated for each of multiple users, with the user's features mixed in a desired proportion with an existing character.
[0116] The 3D model presenting unit 708 presents to each of the multiple users a 3D model of a virtual character synthesized using the facial, skeletal, and clothing features mixed at various ratios on a monitor screen, etc. The audio presenting unit 709 presents to each of the multiple users the voice of the virtual character synthesized using the voice timbre features mixed at various ratios.
[0117] The fourth input is a selection by each of the multiple users. Each user selects a 3D model and voice that is preferred by the user from the 3D models of virtual characters for each mixing ratio presented by the 3D model presentation unit 708 and the synthetic voice of the virtual character for each mixing ratio presented by the voice presentation unit 709. Then, based on the fourth input from each of the multiple users, the 3D model and voice selection unit 710 outputs the 3D model and voice selected by each user to the subsequent character operation processing unit 203 as multiple virtual characters onto which the characteristics of the multiple users are respectively projected.
[0118] According to the operations of the capture processing unit 201 and the character projection processing unit 202 shown in FIG. 8, the information processing system 200 can simultaneously create virtual characters onto which the characteristics of each of a plurality of users are projected.
[0119] By using the processing described above, it is possible to provide a service to an unspecified number of users that generates virtual characters that meet the needs of various users, such as users who want to project themselves into the world view of the video they are creating, or users who want to become a character that strongly reflects the world view of the video work (as set by the creator).
[0120] C-2. Character operation processing section In information processing system 200, it is assumed that capture processing unit 201 extracts user features prior to video shooting, and character operation processing unit 203 extracts user features in parallel with video shooting (as described above). Fig. 9 shows a detailed configuration of character operation processing unit 203, which is designed to enable the user's features to be more accurately projected onto the virtual character.
[0121] The character operation processing unit 203 includes a motion capture unit 901 , a multi-view camera 902 , a microphone 903 , a voice enhancement unit 904 , an emotion recognition unit 905 , a 3D model driving unit 906 , and a voice conversion unit 907 .
[0122] The motion capture 901 acquires the feature quantities of the user's position and posture. The motion capture 901 can use, for example, an optical method that uses multiple cameras to track the positions of reflective markers attached to multiple parts of the user's body, or an inertial method that measures body movement by applying acceleration, angular velocity, and direction information obtained from an inertial sensor attached to the user's body to a skeletal model, but of course other methods may also be used.
[0123] The multi-view camera 902 photographs the user from various angles to assist the motion capture 901 in acquiring position and posture feature amounts, and also acquires the user's facial expression feature amounts.
[0124] The microphone 903 records the voice uttered by the user, and may be, for example, a pin microphone attached to the user's clothing or a microphone array, but is not limited to these forms.
[0125] The voice enhancement unit 904 generates a signal in which the voice of the target user is enhanced from the voice recorded by the microphone 903. When a microphone array is adopted for the microphone 903, the voice enhancement unit 904 may generate an enhanced voice signal for each user by utilizing the position feature of each user.
[0126] The emotion recognition unit 905 recognizes the user's emotion from the facial expression feature obtained by the multi-view camera 902 and the audio signal enhanced by the audio enhancement unit 904, and outputs it as an emotion feature.
[0127] The 3D model driving unit 906 drives the 3D model of the user's virtual character so that it has a position / posture, facial expression, and emotion that reflect the user's position / posture feature amounts, facial expression feature amounts, and emotional feature amounts.
[0128] The voice conversion unit 907 converts the user's voice emphasized by the voice enhancement unit 904 into the voice quality of the user's virtual character.
[0129] The character operation processing unit 203 has the configuration shown in FIG. 9, and is therefore able to cause the virtual character to move in a manner that reflects the user's movements, voice, and emotions.
[0130] D. Third Example This section D describes a third embodiment of the information processing system 200 for enhancing the realism of the user's experience of the world. The information processing system 200 according to the third embodiment enhances the realism of the user's experience of the world by performing physical interaction with the user. The following can be given as means for enhancing the realism of the user's experience of the world.
[0131] (1) The device further includes an operating means having a display device, such as a head-mounted display (HMD). (2) Equipped with means for providing feedback to the user, such as speakers, headphones, or haptics. (3) Application of robotics
[0132] Fig. 10 shows the functional configuration of an information processing system 200 according to the third embodiment. The illustrated information processing system 200 further includes a physical interaction processing unit 1000 as a means for allowing a user to physically directly input fine movements of their fingers or the like and for providing physical feedback to the user. Of the functional components of the information processing system 200 shown in Fig. 2, components that are not directly related to the description of the third embodiment are not shown in Fig. 10.
[0133] The physical interaction processing unit 1000 includes a wearable device or robot 1001 , a force / tactile / proximity angle sensor 1002 , and a dynamics calculation unit 1003 .
[0134] The "wearable device or robot" indicated by the reference numeral 1001 should be understood as a component having at least one of a "wearable device" and a "robot." The "wearable device" is, for example, a glove, but may also be a head-mounted display, a smart watch, or other form of device worn by a user. The "robot" is, for example, a humanoid robot or a pet-type robot, but may also be a robot that moves by means other than legs (for example, wheels), a stationary robot without means of movement, or a voice agent that interacts with a user via voice. The wearable device or robot 1001 is used for input of a fifth input and output of a fifth output.
[0135] The fifth input is an input of geometric information from the user to the wearable device or robot 1001. The user can provide the fifth input to the wearable device or robot 1001 by moving their gloved hands or fingers or by moving the robot's body. The fifth input includes, for example, deformation of the glove and joint angles of the robot.
[0136] The geometric information received as the fifth input can be directly used by the character projection processing unit 202. For example, the character projection processing unit 202 directly projects the geometric information of the user's hand acquired by a glove serving as a wearable device onto the character's hand movements. This makes it possible to reproduce fine movements of the hand and the like with higher accuracy than with a camera image. Furthermore, even if the user's hand movements do not fit within the camera's angle of view or if the user's hand movements are hidden by other objects (for example, parts of the user's body other than the hand) and cannot be captured by the camera, the glove can still acquire the fine movements of the hand and the like with high accuracy.
[0137] The force, touch, and proximity angle sensor 1002 is a collection of force sensors, touch sensors, and proximity angle sensors. Of course, it may also include other sensors. The sixth input is an input from the user to the force, touch, and proximity angle sensor 1002. The force sensor acquires the magnitude of the force from the user, the touch sensor acquires minute tactile information from the user, and the proximity angle sensor acquires distance information around the sensor.
[0138] The dynamics calculation unit 1003 performs dynamics calculations in the virtual space reproduced by the scene reproduction unit 206 based on geometric information acquired from the user as a fifth input by the wearable device or robot 1001 and sensor information obtained by sensing a sixth input from the user by the force, touch, and proximity angle sensor 1002. In this way, the user performs physical input into the virtual space.
[0139] The dynamics calculation information calculated by the dynamics calculation unit 1003 based on physical input from the user can be used to operate a character in the character operation processing unit 203, or can be used for interaction with objects and characters other than the user in scene reproduction (physical model of virtual space) in the scene reproduction processing unit 206. By wearing a head-mounted display or headphones and observing this interaction, the user can have a simulated experience of interacting with objects and characters in the virtual space.
[0140] Furthermore, the dynamics calculation unit 1003 calculates force information generated by interactions between the user's character and other objects or characters in the virtual space reproduced by the scene reproduction processing unit 206. Then, the wearable device or robot 1001 feeds back this force information to the user as a fifth output. Therefore, the user can experience interactions in the virtual space with senses other than sight and hearing through the wearable device or robot 1001.
[0141] E. Fourth Example Section B above described, as a first embodiment of the present disclosure, a method for producing animated videos as a final product by using information processing system 200 to produce videos with a specific worldview, such as an anime. Section E describes, as a fourth embodiment of the present disclosure, a method for producing live-action videos other than anime, 80s-style videos, and the like, using information processing system 200. In the fourth embodiment, by improving primarily scene setting processing unit 205 and video conversion processing unit 208 among the functional components of information processing system 200, it becomes possible to produce live-action videos other than anime, videos from a different era (e.g., 80s-style), and the like.
[0142] E-1. When producing live-action footage For example, when producing live-action footage such as a movie or drama, the scene setting processing unit 205 sets, in addition to language information about the characters and scenes, facial images of the performers, full-body images, costume images, background images, and actually shot moving images.
[0143] Meanwhile, the video conversion processing unit 208 receives as input video data relating to a scene in the virtual space captured by the video capture control unit 207 and converts it into live-action video envisioned by the creator. For example, when the scene reproduction processing unit 206 recreates characters and environments in the virtual space using 3DCG, the video captured by the video capture control unit 207 is a simplified video of the characters and environment captured by a virtual camera placed in the three-dimensional virtual space. The video conversion processing unit 208 then converts this simplified video into live-action video, such as a movie or drama, envisioned by the creator. During the video conversion process by the video conversion processing unit 208, the user's facial features, voice characteristics, and the like are reflected. The video conversion processing unit 208 realizes the conversion from the video (3DCG) captured in the virtual space to live-action video.
[0144] E-2. When creating footage from different eras Furthermore, when producing video from a different era, such as the 1980s, the scene setting processing unit 205 sets attribute information related to the video to be produced, such as the era and location. Then, the video conversion processing unit 208 reflects this attribute information in the video conversion process, so that it is possible to convert into video that reflects the intentions of the producer, such as video from a different era, without being limited to animation video or live-action video.
[0145] E-3. Use of generative AI technology Like the video-to-video converter described above, the video conversion processing unit 208 uses AI technology for generating video. The first and third inputs, which are inputs by the user, and the second input, which is input by the producer, are reflected in this generation AI, so that the video after conversion by the video conversion processing unit 208 reflects the intentions of the producer and is created as a result of the user's participation in filming. The third input is input by the user's own body movements or operation of a device such as a controller to control the body movements and voice of a virtual character. There are situations where it is difficult to control 3DCG, such as the virtual character's movements, emotional expression, and use of tools, and there is a concern that the scene may deviate from what the creator or user wants to visualize. In response to this, the video conversion processing unit 208 can convert the world view of the video by using video generation AI technology to correct any deviation from the scene that is desired to be visualized.
[0146] F. Fifth Example Section B-3 above has provided a detailed explanation of an embodiment relating to the video conversion performed by the video conversion processing unit 208. In this section F, processing relating to the conversion of sound by the video conversion processing unit 208 will be explained in detail.
[0147] The video output from the video shooting control unit 207 includes an audio signal. This audio signal consists of the voices of the characters in the virtual space, ambient sounds, music such as background music, narration, etc. The video conversion processing unit 208 also corrects discrepancies in the audio signal in order to convert the simple video output from the video shooting control unit 207 into the desired video (as envisioned by the creator).
[0148] The information contained in the audio signal is broadly divided into three types: linguistic information, paralinguistic information, and non-linguistic information. The video conversion processing unit 208 also performs conversion processing on the linguistic information, paralinguistic information, and non-linguistic information so that the audio signal also matches the worldview intended by the creator.
[0149] Linguistic information: Equivalent to the content of the speech. Paralinguistic information: Information that a speaker intentionally adds to an utterance that is not included in the linguistic information. This corresponds to prosodic information such as intonation and rhythm, and emotional expressions. Non-verbal information: Information given regardless of the speaker's intentions. It corresponds to the speaker's unique voice quality, etc.
[0150] F-1. Sound conversion example (1) 11 shows an example of a configuration for converting an audio signal in the video conversion processing unit 208. In the example shown, the video conversion processing unit 208 is made up of an audio-to-audio converter 1101. The audio-to-audio converter 1101 is a functional component that uses AI technology for generating audio, and specifically, is made up of a trained model that has been trained based on a machine learning method such as deep learning.
[0151] The audio-to-audio converter 1101 inputs the original audio signal contained in the original video work that serves as a reference, and also inputs the user's unique voice quality extracted by the capture processing unit 201 and the character operation processing unit 203 as auxiliary data.
[0152] However, it is assumed that the original audio signal used as a reference is supplied not from the video shooting control unit 207 but, for example, from the scene reproduction processing unit 206. Furthermore, the capture processing unit 201 extracts the user's voice quality prior to video shooting, and the character operation processing unit 203 extracts the user's voice quality in parallel with video shooting. It is also assumed that the user's voice quality extracted by the capture processing unit 201 is mixed with the voice quality of a preset character (see FIG. 7), and the voice quality obtained as a result of the mixing is also referred to here as the "user's voice quality."
[0153] Then, the audio-to-audio converter 1101 converts the character's voice signal contained in the original audio signal into a voice quality unique to the user. As a result, the user's voice quality is reflected in the speech of the character projected by the user in the output audio signal. The output audio signal is included in the fourth output as an audio signal corresponding to the video signal converted by the video conversion processing unit 208.
[0154] The audio-to-audio converter 1101 can be applied not only to converting voice quality, but also to converting intonation, rhythm, and emotion, which are other characteristics of the user's voice.
[0155] Furthermore, in the above description, an example has been described in which the audio-to-audio converter 1101 performs conversion processing on audio signals contained in the input original audio signal, but it is also possible to perform conversion processing on audio signals other than audio. Furthermore, the audio-to-audio converter 1101 can perform conversion processing on each audio source. When the original audio signal is a mixture of audio signals from multiple audio sources (such as character voices, background sounds, BGM, and narration), the signals are separated into each audio source and then input to the audio-to-audio converter 1101, where individual conversion processing is performed on each audio source. The converted audio source signals are then mixed together and an output audio signal (fourth output) is output.
[0156] F-2. Sound conversion example (2) 12 shows another example of the configuration for converting an audio signal in the video conversion processing unit 208. In the example shown, the video conversion processing unit 208 is made up of an audio-to-audio converter 1201. The audio-to-audio converter 1201 uses an original audio signal contained in the original video work that serves as a reference as auxiliary data to convert the audio signal contained in the video shot by the video shooting control unit 207 into an audio signal that reflects the characteristics of the anime work.
[0157] Specifically, ambient sounds, music such as background music, and narration not included in the audio signal of the captured video are extracted from the original audio signal and applied to audio conversion in the audio-to-audio converter 1201. The audio-to-audio converter 1201 also performs audio conversion to match acoustic characteristics such as reverberation to the reverberation of the original audio signal. The audio-to-audio converter 1201 can also convert paralinguistic information of the user's speech in the video captured by the video capture control unit 207, such as speaking style, prosody, and emotional expression, into that of the original audio signal. This allows, for example, the user's inarticulate or flat speech (in other words, low-quality speech by a user who is not a voice actor) to be converted into clear and richly intoned speech that maintains quality compatible with the output video converted by the video conversion processing unit 208 while maintaining the user's unique voice quality. The output audio signal is included in the fourth output as an audio signal corresponding to the video signal converted by the video conversion processing unit 208.
[0158] The audio signals contained in the video captured by the video capture control unit 207 are a mixture of audio signals from multiple sound sources, such as character voices, background sounds, BGM, and narration, so after separating them into individual sound sources, they are input to the audio-to-audio converter 1101, where individual conversion processing is performed for each sound source, and the converted audio source signals are mixed together to output an output audio signal (fourth output).
[0159] F-3. Sound conversion example (3) 13 shows yet another example configuration for converting an audio signal in video conversion processing unit 208. In the example shown, video conversion processing unit 208 is made up of audio-to-audio converter 1301. Audio-to-audio converter 1301 receives as input an audio signal contained in the video shot by video shooting control unit 207 and an original audio signal contained in the original video work that serves as a reference, and converts the input into an audio signal that reflects the characteristics of the animation work, using voice features such as the user's unique voice quality extracted by capture processing unit 201 and character operation processing unit 203 as auxiliary data.
[0160] However, it is assumed that the original audio signal, which is one of the inputs, is supplied not from the video shooting control unit 207 but from, for example, the scene reproduction processing unit 206. Furthermore, the capture processing unit 201 extracts features such as the user's voice quality, intonation, and rhythm prior to video shooting, while the character operation processing unit 203 extracts features of the user in parallel with video shooting. It is also assumed that the user's features extracted by the capture processing unit 201 are mixed with the features of a preset character (see FIG. 7), and the features obtained as a result of the mixing are also referred to here as "user's features."
[0161] By using the audio-to-audio converter 1301 as auxiliary data the characteristics of the user's voice, it is possible to convert the output audio signal into one that better reflects the user's characteristics, such as voice quality, intonation, and rhythm.
[0162] F-4. Sound conversion example (4) The above-described acoustic conversion examples (1) to (3) are mainly cases in which the video conversion processing unit 208 performs post-audio conversion on audio signals contained in video already captured by the video capture control unit 207. In contrast, it is possible to improve the user's experience by performing audio signal conversion processing in real time in parallel with video capture by the video capture control unit 207. In this section F-4, an example in which audio signal conversion processing is performed in real time in parallel with video capture will be described.
[0163] Character operation processing unit 203 extracts user features in parallel with video capture (as described above). Character operation processing unit 203 shown in FIG. 9 is configured to project user features more accurately onto the virtual character and includes voice conversion unit 907. Voice conversion unit 907 has a configuration similar to the audio-to-audio converters shown in FIGS. 11 to 13 above, but applies a real-time conversion method that can convert the user's speech into a desired voice quality and speech style with low latency. In parallel with the conversion of voice quality and speech style, the transfer characteristics from the speaking user to both ears of other users (binaural room impulse response (BRIR) or, more simply, head-related impulse response (HRIR) using the mirror method) are calculated based on the relative positions of the users and their absolute positions in the virtual space. Then, the voice quality conversion unit 907 reflects the transfer characteristics in the voice converted by the audio-to-audio converter, and reproduces the resulting binaural voice signal from headphones attached to a head-mounted display worn by another user.
[0164] By performing processing that reflects these transmission characteristics in the audio with low latency, users other than the speaker can perceive the speaker's direction information and room information in the virtual space acoustically (for example, through reverberation), allowing them to become more immersed in the experience.
[0165] The above-described acoustic conversion examples (1) to (4) make it possible to reflect the features of the user's voice in the voice of the character projecting the user. Furthermore, particularly in scenes where the character speaks emotionally or sings, the user's voice acting or singing skills are often insufficient to achieve the voice or singing quality expected in the fourth output. However, the conversion of the acoustic signal in the video conversion processing unit 208 (or the voice quality conversion unit 907) can correct any discrepancies resulting from the user's insufficient voice acting or singing skills, and can also reflect the user's features in the voice quality, etc.
[0166] G. System Features and Effects In this section G, the features and effects of the information processing system 200 are summarized.
[0167] G-1. Main Features The information processing system 200 is characterized by having a configuration for realizing the following (1) to (4). (1) Capture the user in the real world, project the user's features onto a character in a 3D virtual world, and control the character in the virtual world. (2) The character's experience in the virtual world is virtually captured using a camera in the virtual world, and the captured footage is then converted into footage of another worldview, such as 2D animation footage, and saved. (3) Or, for example, in an existing (original) animated video, convert the characteristics of a character into the characteristics of the user. (4) In the video conversion process of (2) and (3) above, the expression desired by the user is reflected.
[0168] Therefore, with the information processing system 200, it is possible to easily include a user in a video work (for example, to include a character that reflects the user's characteristics in an animated video, and to operate the character within the worldview of the video). Furthermore, it is possible to provide such a service to an unspecified number of users.
[0169] Furthermore, according to the information processing system 200, the user can be involved in the production of a video work, enjoy the video work that includes the user, and enjoy the worldview of the video work more deeply. For example, by the user controlling a character that is included in an animated video and that reflects the user's characteristics, the user can be involved in the production of the video work, enjoy the video work, and enjoy the worldview of the video work more deeply.
[0170] G-2. Additional Features (1) The information processing system 200 is also characterized by a configuration that combines the features of a user with the features of a character that reflects the worldview of a video work, and synthesizes a virtual character that combines the features of both.
[0171] Therefore, according to the information processing system 200, when performing character projection processing, a choice can be made that respects the user's wishes in the trade-off that arises between the degree of self-projection and the sense of immersion in the world of the video work (the sense of becoming a character in the video work).
[0172] G-3. Additional Features (2) The information processing system 200 may further include a physical interaction processing unit as a means for the user to directly physically input fine movements of the fingers or to provide force feedback to the user. The physical interaction processing unit may include a wearable device or a robot.
[0173] By using this physical interaction processing unit, the user can experience interactions in a virtual space using senses other than sight and hearing.
[0174] H. Configuration of information processing device 14 shows an example of the hardware configuration of an information processing device 2000 that can operate as the information processing system 200 according to the present disclosure. The illustrated information processing device 2000 includes a CPU (Central Processing Unit) 2001, a ROM (Read Only Memory) 2002, a RAM (Random Access Memory) 2003, a host bus 2004, a bridge 2005, an expansion bus 2006, an interface unit 2007, an input unit 2008, an output unit 2009, a storage unit 2010, a drive 2011, and a communication unit 2013. The information processing device 2000 is configured, for example, by a personal computer, but some of its functions may be configured by an information terminal such as a tablet or a smartphone.
[0175] The CPU 2001 controls the overall operation of the information processing device 2000 in accordance with various programs. When performing computationally intensive processing such as learning an AI model (described above) on the information processing device 2000, it is desirable that the CPU 2001 be a multi-core CPU (e.g., Apple M1 Max) or that the information processing device 2000 further include a multi-core processor (e.g., NVIDIA R6000) such as a GPU (Graphics Processing Unit) or GPGPU (General-purpose computing on graphics processing units) in addition to the CPU 2001. However, for convenience, these will hereinafter be collectively referred to simply as the CPU 2001.
[0176] ROM 2002 stores in a nonvolatile manner programs (such as a basic input / output system) and calculation parameters used by CPU 2001. RAM 2003 is used to load programs used in the execution of CPU 2001 and to temporarily store parameters such as working data that change as appropriate during program execution. Programs loaded into RAM 2003 and executed by CPU 2001 include, for example, various application programs and an operating system (OS).
[0177] The CPU 2001, ROM 2002, and RAM 2003 are interconnected by a host bus 2004, which includes a CPU bus and other components. The CPU 2001 executes various application programs under an execution environment provided by an OS through cooperative operation of the ROM 2002 and RAM 2003, thereby enabling various functions and services. If the information processing device 2000 is a personal computer, the OS may be, for example, Microsoft Windows (registered trademark), Unix (registered trademark), or a successor OS. For example, programs for operating as each processing unit, such as the capture processing unit 201, character projection processing unit 202, character operation processing unit 203, video capture instruction unit 204, scene setting processing unit 205, scene reproduction processing unit 206, video capture control unit 207, video conversion processing unit 208, and converted video saving processing unit 209 shown in FIG. 2, are executed on the information processing device 2000.
[0178] The host bus 2004 is connected to an expansion bus 2006 via a bridge 2005. The expansion bus 2006 is, for example, a PCI (Peripheral Component Interconnect) bus or PCI Express, and the bridge 2005 is based on the PCI standard. However, the information processing device 2000 does not need to be configured so that the circuit components are separated by the host bus 2004, bridge 2005, and expansion bus 2006, and may be implemented so that almost all circuit components are interconnected by a single bus (not shown).
[0179] The interface unit 2007 connects peripheral devices such as an input unit 2008, an output unit 2009, a storage unit 2010, a drive 2011, and a communication unit 2013 in accordance with the standards of the expansion bus 2006. However, not all of the peripheral devices shown in Fig. 14 are necessarily required, and the information processing device 2000 may further include peripheral devices not shown. Furthermore, the peripheral devices may be built into the main body of the information processing device 2000, or some of the peripheral devices may be externally connected to the main body of the information processing device 2000.
[0180] The input unit 2008 is composed of an input control circuit that generates an input signal based on an input from a user and outputs the signal to the CPU 2001. When the information processing device 2000 is a personal computer, the input unit 2008 may include a keyboard, a mouse, a touch panel, a camera, and a microphone. The output unit 2009 includes display devices such as a liquid crystal display (LCD) device, an organic EL (Electro-Luminescence) display device, and an LED (Light Emitting Diode), as well as an audio output device such as a speaker. The input unit 2008 is used to input a moving image to be processed, and the output unit 2009 is used to display a GUI screen, etc.
[0181] The storage unit 2010 stores files such as programs (applications, OS, etc.) and various data executed by the CPU 2001. The storage unit 2010 is configured with a large-capacity storage device such as an SSD (Solid State Drive) or HDD (Hard Disk Drive), but may also include an external storage device.
[0182] The removable storage medium 2012 is a storage medium configured as a cartridge, such as a microSD card. The drive 2011 performs read and write operations on the loaded removable storage medium 113. The drive 2011 outputs data read from the removable storage medium 2012 to the RAM 2003 or the storage unit 2010, and writes data on the RAM 2003 or the storage unit 2010 to the removable storage medium 2012.
[0183] The communication unit 2013 is a device that performs wireless communication such as Wi-Fi (registered trademark), Bluetooth (registered trademark), or a cellular communication network such as 4G or 5G. The communication unit 2013 may also include terminals such as a Universal Serial Bus (USB) and a High-Definition Multimedia Interface (HDMI) (registered trademark), and may further include a function for performing HDMI (registered trademark) communication with USB devices such as scanners and printers, displays, etc. Programs executed on the information processing device 2000 are installed from the outside via the communication unit 2013, for example. [Industrial Applicability]
[0184] The present disclosure has been described in detail above with reference to specific embodiments. However, the present disclosure should not be construed as being limited to the above-described embodiments, and it is obvious that those skilled in the art can modify or substitute the embodiments without departing from the spirit of the present disclosure. Furthermore, the effects described in this specification are merely examples, and the effects brought about by the present disclosure are not limited thereto, and additional effects not described in this specification may exist.
[0185] Although the present specification has mainly described an embodiment in which the present disclosure is applied to the production of animated videos, the gist of the present disclosure is not limited thereto. According to the present disclosure, by further converting a simple video of a scene in a virtual space reproduced by 3DCG, captured by a virtual camera, it is possible to produce various forms of video works envisioned by the producer, such as 2D animated videos, live-action videos like movies and dramas, and videos from different eras.
[0186] In short, the present disclosure has been described in the form of examples, and the contents of the specification should not be interpreted as limiting. To determine the gist of the present disclosure, the claims should be taken into consideration.
[0187] The series of processes described in this specification can be executed by hardware, software, or a configuration that combines hardware and software. When executing processes by software, a program recording a processing sequence related to realizing the present disclosure is installed in memory in a computer incorporated in dedicated hardware and executed. It is also possible to install the program in a general-purpose computer capable of executing various processes and execute the processes related to realizing the present disclosure.
[0188] The program can be stored in advance on a recording medium installed in the computer, such as a HDD, SSD, or ROM. Alternatively, the program can be temporarily or permanently stored on a removable recording medium, such as a flexible disk, a CD-ROM (Compact Disc Read Only Memory), an MO (Magneto optical) disk, a DVD (Digital Versatile Disc), a BD (Blu-Ray Disc (registered trademark)), a magnetic disk, or a USB (Universal Serial Bus) memory. Using such a removable recording medium, a program related to realizing the present disclosure can be provided as so-called package software.
[0189] The program may also be transferred wirelessly or via a wired connection from a download site to a computer via a network such as a wide area network (WAN) (e.g., a cellular network), a local area network (LAN), or the Internet. The computer can receive the program transferred in this manner and install it on a large-capacity storage device such as a hard disk drive (HDD) or solid state drive (SSD) within the computer.
[0190] The present disclosure may also be configured as follows.
[0191] (1) An information processing system that performs processing related to content generation, a user feature acquisition unit that acquires features related to the user; a scene setting information acquisition unit that acquires scene setting information related to the first content; a content generation unit that generates second content including a character projected as the user into the first content based on the characteristics of the user and setting information of the scene; An information processing system comprising:
[0192] (2) The first content is video content, and the second content is video content including a character that reflects the characteristics of the user. The information processing system according to (1) above.
[0193] (3) The first content is an animated video, and the second content is an animated video including a character that reflects the characteristics of the user. The information processing system according to any one of (1) and (2) above.
[0194] (4) The first content includes a voice of a character, and the second content includes a voice of a character that projects the characteristics of the user. The information processing system according to any one of (1) to (3) above.
[0195] (5) The first content is an original video content; the scene setting information acquisition unit acquires the setting information based on a second input including a setting of the original work related to at least one of an environment, characters, and sounds in a scene to be visualized; The information processing system according to any one of (1) to (4) above.
[0196] (6) The content generation unit includes a character projection processing unit that projects characteristics related to the user onto a character appearing in the original video content, a scene reproduction processing unit that reproduces a scene in a virtual space based on the setting information, a character operation processing unit that moves the character in the virtual space, a video shooting control unit that shoots the virtual space in which the character moves with a virtual camera, and a video conversion processing unit that converts the video shot by the video shooting control unit into the second content consisting of video based on the assumptions of at least one of the creator of the original work or the user. The information processing system according to (5) above.
[0197] (7) The user feature acquisition unit acquires features related to the user based on a first input including information about at least one of the user's face, body, and voice. The information processing system according to (6) above.
[0198] (8) The character operation processing unit performs processing for moving the character in the virtual space based on a third input from the user. An information processing system according to any one of (6) or (7) above.
[0199] (9) A video shooting instruction unit that outputs a first output regarding an instruction to the user to operate the character, The information processing system according to any one of (6) to (8) above.
[0200] (10) The video shooting control unit shoots the scene in the virtual space with a desired angle of view and camera work. The information processing system according to any one of (6) to (9) above.
[0201] (11) outputting at least one of a second output relating to the image of the scene in the virtual space reproduced by the scene reproduction processing unit, a third output relating to the image captured by the shooting control unit, and a fourth output relating to the image after conversion by the image conversion processing unit; The information processing system according to any one of (6) to (10) above.
[0202] (12) The scene setting information acquisition unit acquires at least one of a 3D model of a background in the virtual space, a 3D model and motion of a character in the virtual space, and sound data in the virtual space based on the second input. The information processing system according to any one of (6) to (11) above.
[0203] (13) The video conversion processing unit replaces the feature amount of the user with a partial video related to the character in the original video content. The information processing system according to any one of (6) to (12) above.
[0204] (13-1) Further comprising an audio conversion unit that converts the voice of a character in an audio signal included in the original video content into the voice quality of the user. The information processing system according to (13) above.
[0205] (14) The video conversion processing unit uses the feature amount of the character in the video captured by the video capture control unit as an auxiliary input, and converts the style of the video captured by the video capture control unit in accordance with a prompt from the user. The information processing system according to any one of (6) to (12) above.
[0206] (14-1) The video conversion processing unit converts the style of the captured video while preserving the auxiliary input feature amount. The information processing system according to (14) above.
[0207] (14-2) An audio conversion unit is further provided that converts the audio signal included in the captured video by applying at least one of ambient sound, music such as background music, narration, and reverberation extracted from the audio signal included in the original video content. The information processing system according to (14) above.
[0208] (15) The video shooting control unit shoots the virtual space in which the character moves, the character being replaced with a feature extracted from the character in the original video content; the video conversion processing unit converts the style of the video captured by the video capture control unit using the feature amount extracted from the character in the original video content as an auxiliary input; The information processing system according to any one of (6) to (12) above.
[0209] (15-1) The video conversion processing unit converts the style of the captured video while preserving the auxiliary input feature amount. The information processing system according to (15) above.
[0210] (15-2) The audio system further includes an audio conversion unit that receives an audio signal included in the captured video and an audio signal included in the original video content, and converts the audio signal included in the captured video by using characteristics of the user's voice as retention information. The information processing system according to (15) above.
[0211] (16) The character projection processing unit projects onto the character a feature that is a combination of the feature related to the user acquired by the user feature acquisition unit and the feature of the original character based on the second input. The information processing system according to any one of (6) to (15) above.
[0212] (17) The character operation processing unit extracts a feature of the user based on the third input. The information processing system according to (8) above.
[0213] (17-1) The character operation processing unit extracts features of the user in parallel with the video shooting control unit shooting. The information processing system according to (17) above.
[0214] (17-2) The character operation processing unit includes a motion capture device, a multi-view camera, and a microphone; The character is made to move in the virtual space based on the position and posture feature amounts of the user acquired by the motion capture and the multi-view camera, the facial expression feature amounts of the user acquired by the multi-view camera, and the emotion feature amounts of the user recognized based on the facial expression feature amounts and the voice of the user picked up by the microphone. The information processing system according to (17) above.
[0215] (17-3) The character operation processing unit further includes a voice conversion unit that converts the user's voice into the character's voice quality. The information processing system according to (17-2) above.
[0216] (18) Further comprising a physical interaction processing unit that performs physical interaction with the user. The information processing system according to any one of (6) to (17) above.
[0217] (18-1) The physical interaction processing unit a wearable device or robot that receives a fifth input relating to geometric information from the user and outputs a fifth output for providing feedback to the user via a sense other than vision and hearing; a force, touch, and proximity angle sensor that detects a sixth input from the user relating to force, touch, and proximity angle; a dynamics calculation unit that performs a dynamics calculation based on the fifth input and the sixth input; The information processing system according to (18) above, comprising: The physical interaction processing unit
[0218] (18-2) The character projection processing unit projects the fifth input acquired by the wearable device or the robot onto the movement of the character. The information processing system according to (18-1) above.
[0219] (18-3) The character operation processing unit operates the character based on the dynamics calculation information calculated by the dynamics calculation unit. The information processing system according to (18-1) above.
[0220] (18-4) The scene reproduction processing unit reproduces an interaction between the user and an object or character other than the user in the virtual space based on the dynamics calculation information calculated by the dynamics calculation unit. The information processing system according to (18-1) above.
[0221] (18-5) The dynamics calculation unit calculates force information generated by interactions between objects and characters in the virtual space reproduced by the scene reproduction processing unit, The wearable device or the robot outputs the fifth output based on the force information. The information processing system according to (18-1) above.
[0222] (19) An information processing method for processing content generation, a user feature acquisition step of acquiring features related to the user; a scene setting information acquisition step of acquiring scene setting information related to the first content; a content generation step of generating second content including a character projected as the user into the first content based on the characteristics of the user and setting information of the scene; An information processing method comprising:
[0223] (20) A computer program written in a computer-readable format to execute a process related to content generation on a computer, the computer comprising: a user feature acquisition unit that acquires features related to the user; a scene setting information acquisition unit that acquires scene setting information related to the first content; a content generation unit that generates second content including a character projected as the user into the first content based on the characteristics of the user and setting information of the scene; A computer program that functions as a [Explanation of symbols]
[0224] 100...information processing system, 101...user characteristic acquisition unit 102...Scene setting information acquisition unit, 103...Animation image generation unit 200...information processing system, 201...capture processing unit 202...character projection processing unit, 203...character operation processing unit 204...video shooting instruction unit, 205...scene setting processing unit 206...scene reproduction processing unit, 207...video shooting control unit 208...video conversion processing unit, 209...converted video storage processing unit 210...scene output unit, 211 shot image output unit 212...Converted video output unit 401, 501, 601...video-to-video converter 701...camera, 702...microphone, 703...user feature extraction unit 704...Feature extraction unit, 705...Feature mixing unit 706...3D model synthesis unit, 707...voice synthesis unit 708...3D model presentation unit, 709...audio presentation unit 710...3D model and audio selection section 901...Motion capture, 902...Multi-view camera 903...microphone, 904...voice enhancement unit, 905...emotion recognition unit 906...3D model driving unit, 907...Voice conversion unit 1001...Wearable devices or robots 1002...mechanical calculation unit, 1003...force, tactile, and proximity angle sensor 2000...information processing device, 2001...CPU, 2002...ROM 2003...RAM, 2004...Host Bus, 2005...Bridge 2006... Expansion bus, 2007... Interface section 2008...input section, 2009...output section, 2010...storage section 2011...Drive, 2012...Removable Recording Media 2013…Communications Department
Claims
1. An information processing system that performs processing related to content generation, a user feature acquisition unit that acquires features related to the user; a scene setting information acquisition unit that acquires scene setting information related to the first content; a content generation unit that generates second content including a character projected as the user into the first content based on characteristics related to the user and setting information of the scene; An information processing system comprising:
2. The first content is video content, and the second content is video content including a character that projects characteristics of the user. The information processing system according to claim 1 .
3. The first content is an animated video, and the second content is an animated video including a character that projects characteristics of the user. The information processing system according to claim 1 .
4. The first content includes a voice of a character, and the second content is a content including a voice of a character that projects characteristics of the user. The information processing system according to claim 1 .
5. the first content is an original video content; the scene setting information acquisition unit acquires the setting information based on a second input including a setting of the original work related to at least one of an environment, characters, and sounds in a scene to be visualized; The information processing system according to claim 1 .
6. The content generation unit includes a character projection processing unit that projects characteristics related to the user onto a character appearing in the original video content, a scene reproduction processing unit that reproduces a scene in a virtual space based on the setting information, a character operation processing unit that causes the character to operate in the virtual space, a video shooting control unit that shoots the virtual space in which the character operates with a virtual camera, and a video conversion processing unit that converts the video shot by the video shooting control unit into the second content consisting of video based on an assumption by at least one of the creator of the original work or the user. The information processing system according to claim 5 .
7. the user feature acquisition unit acquires features related to the user based on a first input including information about at least one of a face, a body, and a voice of the user; The information processing system according to claim 6.
8. the character operation processing unit performs processing for moving the character in the virtual space based on a third input from the user. The information processing system according to claim 6.
9. a video shooting instruction unit that outputs a first output regarding an instruction to the user to operate the character; The information processing system according to claim 6.
10. the video shooting control unit shoots the scene in the virtual space with a desired angle of view and camera work; The information processing system according to claim 6.
11. outputting at least one of a second output relating to the image of the scene in the virtual space reproduced by the scene reproduction processing unit, a third output relating to the image captured by the shooting control unit, and a fourth output relating to the image after conversion by the image conversion processing unit; The information processing system according to claim 6.
12. the scene setting information acquisition unit acquires, based on the second input, at least one of a 3D model of a background in the virtual space, a 3D model and motion of a character in the virtual space, and sound data in the virtual space; The information processing system according to claim 6.
13. the video conversion processing unit replaces a partial video related to the character in the original video content with the feature amount of the user; The information processing system according to claim 6.
14. the video conversion processing unit uses a feature amount of the character in the video captured by the video capture control unit as an auxiliary input, and converts the style of the video captured by the video capture control unit in accordance with a prompt from the user; The information processing system according to claim 6.
15. the video shooting control unit shoots the virtual space in which the character, whose movements have been replaced with features extracted from the character in the original video content, is moving; the video conversion processing unit converts the style of the video captured by the video capture control unit using the feature amount extracted from the character in the original video content as an auxiliary input; The information processing system according to claim 6.
16. the character projection processing unit projects onto the character a feature that is a combination of the feature related to the user acquired by the user feature acquisition unit and the feature of the original character based on the second input; The information processing system according to claim 6.
17. the character operation processing unit extracts a feature of the user based on the third input; The information processing system according to claim 8 .
18. Further comprising a physical interaction processing unit that performs physical interaction with the user. The information processing system according to claim 6.
19. An information processing method for performing processing related to content generation, a user feature acquisition step of acquiring features related to the user; a scene setting information acquisition step of acquiring scene setting information related to the first content; a content generation step of generating second content including a character projected as the user into the first content based on the characteristics of the user and setting information of the scene; An information processing method comprising:
20. A computer program written in a computer-readable format to execute a process related to content generation on a computer, the computer comprising: a user feature acquisition unit that acquires features related to the user; a scene setting information acquisition unit that acquires scene setting information related to the first content; a content generation unit that generates second content including a character projected as the user into the first content based on characteristics related to the user and setting information of the scene; A computer program that functions as a
Citation Information
Patent Citations
Operation recording system for cg model
JP1998302085A
System and method for entering real object into virtual three-dimensional space
JP2002058045A
Simple animation creation apparatus
JP2010092402A