Content generation system, method, program, and storage medium
The content generation system facilitates the creation of content with user-intended character expressions and motions by utilizing a user interface with operation areas, improving the customization and creativity of content generation.
Patent Information
- Application Number
- JP2025018186
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2026-01-07
- Estimated Expiration
- 2043-11-17
AI Technical Summary
Existing methods lack the ability to easily generate content comprising frames with user-intended facial expressions for characters.
A content generation system that includes a user interface with operation areas for determining character expressions, line of sight, motion, and background settings, along with voice and text generation capabilities, allowing users to create content with intended character expressions through intuitive operations.
Enables users to easily generate content with intended character expressions, motions, and backgrounds, enhancing the creativity and customization of content creation processes.
Smart Images

Figure 0007795660000001 
Figure 0007795660000002 
Figure 0007795660000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a content generation system, a method, a program, and a storage medium. [Background technology]
[0002] A method for controlling a virtual character (or "avatar") using a multimodal model is known (see Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Special Publication No. 2022-534708 Summary of the Invention [Problem to be solved by the invention]
[0004] However, there has been no proposal so far for generating content made up of frames containing characters with expressions intended by the user.
[0005] The problem to be solved by the present invention is to easily generate content made up of frames containing characters with facial expressions that the user intends. [Means for solving the problem]
[0006] [1] A content generation system according to one aspect includes: A content generation system that generates content consisting of one or more frames including moving or still images of a character, a display control unit that displays, on a display unit, a user interface including a first two-dimensional operation area and a preview image of a frame to be edited by a user; an expression determination unit that determines an expression of the character on the preview image in accordance with the position of a pointer on the first two-dimensional operation area; a content generation unit that generates content consisting of one or more frames including the frame edited by the user; Equipped with.
[0007] [2] A content generation system according to one aspect is the content generation system according to [1] above, The user interface includes a second two-dimensional operation area, and further includes a line of sight determination unit that determines the line of sight of the character on the preview image according to the position of a pointer on the second two-dimensional operation area.
[0008] [3] A content generation system according to one aspect is the content generation system according to [1] or [2] above, The game device further includes a motion determination unit that determines a motion of the character on the preview image in response to an operation of the user interface by the user.
[0009] [4] A content generation system according to one aspect is the content generation system according to [3] above, the user interface includes a one-dimensional operation area; The motion determination unit determines a reproduction range of the motion according to the position of the knob on the one-dimensional operation area.
[0010] [5] A content generation system according to one aspect is the content generation system according to any one of [1] to [4] above, The device further includes a motion determination unit that determines a motion of the character on the preview image in accordance with the position of the pointer on the first two-dimensional operation area.
[0011] [6] A content creation system according to one aspect is the content creation system according to any one of [1] to [5] above, The image display device further includes a character determination unit that determines a character to be displayed on the preview image.
[0012] [7] A content creation system according to one aspect is the content creation system according to any one of [1] to [6] above, The image processing device further includes a background image determination unit that determines a background image for the preview image.
[0013] [8] A content creation system according to one aspect is the content creation system according to any one of [1] to [7] above, The display device further includes a text information determination unit that determines at least one of the text to be displayed in the preview image, the type of speech bubble in which the text is displayed, the color of the text, and the size of the text in response to an operation of the user interface by the user.
[0014] [9] A content creation system according to one aspect is the content creation system according to any one of [1] to [8] above, The first two-dimensional operation area is a user interface for setting the emotion of the character.
[0015]
[10] A content creation system according to one aspect is the content creation system according to any one of [1] to [9] above, The system further includes a voice generation unit that generates voice data for a character by inputting image data of the character selected by the user into a voice model that uses a character image and voice data associated with the character image as training data.
[0016]
[11] A content creation system according to one aspect is the content creation system according to any one of [1] to
[10] above, The system further includes a speech data generation unit that generates speech data for a character that responds to the user's speech data by inputting the user's speech data into a language model trained using the character's image data and / or voice data as part of the training data.
[0017]
[12] A content generation system according to one aspect is the content generation system according to any one of [1] to
[11] above, The image processing device further includes a character image generation unit that generates a character image by inputting the image data input by the user into a trained image processing model.
[0018]
[13] A content generation system according to one aspect is the content generation system according to any one of [1] to
[12] above, The image processing device further includes a background image generation unit that generates a background image by inputting the image data input by the user into a trained image processing model.
[0019]
[14] A content generation system according to one aspect is a content generation system according to any one of [1] to
[13] above. In the content generation system described above, The image data generated by the character image generation unit and the background image generation unit is three-dimensional image data and is placed in the metaverse space.
[0020]
[15] A content generation system according to one aspect is the content generation system according to any one of [1] to
[14] above, The device further includes an NFT issuing unit that issues NFTs related to the content generated by the content generating unit and / or image data of characters within the content.
[0021]
[16] In one embodiment, the method comprises: A method for generating content consisting of one or more frames including a moving or still image of a character, comprising: a display control step of displaying on a display unit a user interface including a first two-dimensional operation area and a preview image of a frame to be edited by the user; an expression determination step of determining an expression of the character on the preview image according to the position of the pointer on the first two-dimensional operation area; a content generation step of generating content consisting of one or more frames including the frame edited by the user; It has.
[0022]
[17] A program according to one aspect includes: A program for causing a computer to execute a method for generating content consisting of one or more frames including a moving image or a still image of a character, the method comprising: a display control step of displaying on a display unit a user interface including a first two-dimensional operation area and a preview image of a frame to be edited by the user; an expression determination step of determining an expression of the character on the preview image according to the position of the pointer on the first two-dimensional operation area; a content generation step of generating content consisting of one or more frames including the frame edited by the user; It has.
[0023]
[18] A storage medium according to one aspect includes: A computer-readable storage medium storing the program described in
[17] above.
[0024]
[19] In one aspect, a content generation system includes: A content generation system that generates content consisting of one or more frames including moving or still images of a character, a display control unit that displays a user interface and a preview image of a frame being edited by a user on a display unit; a character determination unit that determines a character to be displayed on the preview image in response to an operation of the user interface by the user; an expression determination unit that determines an expression of the character in response to an operation of the user interface by the user; a line of sight determination unit that determines a line of sight of the character in response to an operation of the user interface by the user; a motion determination unit that determines a motion of the character in response to an operation of the user interface by the user; a background determination unit that determines a background image of the preview image in response to an operation of the user interface by the user; a text information determination unit that determines at least one of text to be displayed in the preview image, a type of speech bubble displaying the text, a color of the text, and a size of the text in response to an operation of the user interface by the user; a content generation unit that generates content consisting of one or more frames including the frame edited by the user; Equipped with. [Effects of the Invention]
[0025] According to the present invention, it is possible to easily generate content made up of frames containing characters with expressions intended by the user. [Brief explanation of the drawings]
[0026] [Figure 1] 1 is a diagram showing an example of a system configuration of a content generation system 1 according to the present embodiment. [Figure 2] 1 is a block diagram showing an example of a hardware configuration of a content generation system 1 according to the present embodiment. [Figure 3] 1 is a block diagram showing an example of a functional configuration of a content generation system 1 according to the present embodiment. [Figure 4] 10 is a flowchart showing an example of the operation of the content creation system 1 according to the present embodiment. [Figure 5] FIG. 10 is a diagram showing an example of a content creation screen according to the embodiment. [Figure 6] FIG. 10 is a diagram showing an example of a content creation screen according to the present embodiment, showing a state after a character has been determined. [Figure 7] FIG. 10 is a diagram showing an example of a content creation screen according to the present embodiment, showing a state after a character's facial expression has been determined. [Figure 8] FIG. 10 is a diagram showing an example of a content creation screen according to the present embodiment, showing a state after a character's line of sight has been determined. [Figure 9] FIG. 10 is a diagram showing an example of a content creation screen according to the present embodiment, showing a state in which a character motion is being selected. [Figure 10] FIG. 10 is a diagram showing an example of a content creation screen according to the present embodiment, showing a state after a character's motion has been determined. [Figure 11] FIG. 10 is a diagram showing an example of a content creation screen according to the present embodiment, showing a state after the placement of characters has been determined. [Figure 12] FIG. 10 is a diagram showing an example of a content creation screen according to the present embodiment, showing a state after a background image has been determined. [Figure 13] FIG. 10 is a diagram showing an example of a content creation screen according to the present embodiment, showing a state after text information has been determined. [Figure 14] FIG. 4 is a diagram illustrating an example of a facial expression setting unit according to the embodiment. [Figure 15] FIG. 4 is a diagram illustrating an example of a line-of-sight setting unit according to the present embodiment. [Figure 16] FIG. 4 is a diagram illustrating an example of a motion setting unit according to the embodiment. [Figure 17] FIG. 10 is a diagram illustrating another example of the facial expression setting unit according to the embodiment. [Figure 18] 10A and 10B are diagrams illustrating another example of the line of sight setting unit according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0027] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. In each drawing, components having equivalent functions are designated by the same reference numerals, and detailed description of the components having the same reference numerals will not be repeated.
[0028] (Content Generation System Overview) The content generation system according to this embodiment allows a user to create one or more content such as manga or anime. It is used to generate content made up of frames. The content generation system according to this embodiment will be described in detail below.
[0029] (Content generation system configuration) 1 is a diagram showing a schematic configuration of a content creation system 1 according to this embodiment. As shown in FIG. 1, the content creation system 1 according to this embodiment includes a user terminal 2 (terminal device) used by a user, and a server 3.
[0030] The user terminal 2 and the server 3 are connected to each other so that they can communicate with each other via a network 4 such as the Internet. The network 4 may be either a wired line or a wireless line, and the type and form of the line do not matter. At least a part of the user terminal 2 and the server 3 is realized by a computer (information processing device). The user terminal 2 is, for example, a terminal device such as a personal computer, a smartphone, or a tablet terminal. There are assumed to be a large number of user terminals 2, depending on the number of users who use the content generation system 1.
[0031] (Hardware configuration) Next, a description will be given of the hardware configuration of the content generation system 1 according to this embodiment. Fig. 2 is a block diagram showing an example of the hardware configuration of the user terminal 2 and the server 3 included in the content generation system 1 according to this embodiment.
[0032] In the user terminal 2, the CPU 201 is a processing device that controls the overall operation of the user terminal 2. The ROM 202 is a non-volatile memory that stores control programs executed by the CPU 201 and various data. The RAM 203 is a volatile memory used as a load area and work area for programs executed by the CPU 201. The storage device 204 is a storage means for storing various types of information, and may be built into the user terminal 2 itself or may have a removable storage medium. The input device 205 is a device through which a user of the user terminal 2 inputs information, and may be, for example, a keyboard, mouse, touch panel, microphone, etc. The display 206 is a display device that displays various types of information (user interface, etc.). The image sensor 207 is a photoelectric conversion element that captures an image of a subject. The communication I / F (interface) 208 is an interface for connecting to the network 4. The bus 209 is a bus line that interconnects the above components.
[0033] In the server 3, the CPU 301 is a processing device that controls the overall operation of the server 3. The ROM 302 is a non-volatile memory that stores the control programs executed by the CPU 301 and various data. The RAM 303 is a volatile memory that is used as a load area and work area for the programs executed by the CPU 301. The storage device 304 is a storage means for storing various information, and may be built into the server 3 main body or may have a removable storage medium. The communication I / F (interface) 305 is an interface for connecting to the network 4. The bus 306 is a bus line that connects the above components to each other.
[0034] (Functional configuration) Next, a functional configuration of the content generation system 1 according to this embodiment will be described. Fig. 3 is a diagram showing an example of the functional configuration of the content generation system 1 according to this embodiment.
[0035] First, we will explain the functional configuration of the user terminal 2. As shown in Fig. 3, the user terminal 2 has a communication unit 21, a control unit 22 that controls the overall operation of the user terminal 2, an input unit 23 through which the user inputs various information, an output unit 24 that outputs various information, an imaging unit 25 that captures an image of a subject, and a storage unit 26 that stores various information.
[0036] The communication unit 21 is a communication interface between the user terminal 2 and the network 4. The communication unit 21 transmits and receives information between the user terminal 2 and the server 3 via the network 4.
[0037] The control unit 22 has a display control unit 22a, a frame determination unit 22b, a character determination unit 22c, a facial expression determination unit 22d, a line of sight determination unit 22e, a motion determination unit 22f, a character placement determination unit 22g, a background image determination unit 22h, and a text information determination unit 22i.
[0038] The display control unit 22a displays a content creation screen (operation screen) of a content creation application (program) related to the content creation system 1 on an output unit 24 (display unit) described later.
[0039] 5 is a diagram showing an example of a content creation screen 50 (operation screen) displayed by display control unit 22a on output unit 24 (display unit). As shown in FIG. 5, content creation screen 50 displays a plurality of GUIs (Graphical User Interfaces) for editing frames that constitute content. Specifically, content creation screen 50 includes tabs 51 for switching editing items, a frame list area 52 in which thumbnail images of frames that constitute content are displayed as a list and where a frame to be edited is selected, a preview display area 53 in which a preview image of the frame being edited is displayed, a background image area 53a which is at least a portion of the preview display area and in which a background image is displayed, a character selection area 54 in which a character is selected, a character facial expression setting area 55, a character line of sight setting area 56, a character motion setting area 57, and a cursor C which the user uses to perform a selection operation on each GUI.
[0040] Returning to Fig. 3, frame determination unit 22b determines which frame the user will edit from among one or more frames that make up the content. Specifically, frame determination unit 22b determines which frame the user will edit in response to the user's operation of using cursor C to select a thumbnail image of a frame displayed in frame list area 52 (see Fig. 5). Fig. 5 shows a state in which frame 1 is selected by the user from thumbnail images of frames 1 to 4 displayed in the frame list area, and frame determination unit 22b determines frame 1 as the frame to be edited.
[0041] Note that a frame can be added to the frame list area 52 by selecting the "+" button 52a displayed within the frame list area 52. Furthermore, the selected frame (frame 1 in the example shown in FIG. 5) can be deleted from the frame list area 52 by selecting the "-" button 52b. Furthermore, by selecting the play button 52c, content made up of one or more frames (frames 1 to 4 in the example shown in FIG. 5) displayed in the frame list area 52 is generated by the content generation unit 32c (described below), and the generated content is displayed in the preview display area. For example, in the example shown in FIG. 5, frames 1 to 4 displayed as thumbnails in the frame list area 52 are displayed in the preview display area 53, switching continuously at predetermined time intervals, like a slideshow.
[0042] Returning to FIG. 3, the character determination unit 22c determines the character to be displayed in the frame being edited by the user. Specifically, the character determination unit 22c determines the character selected by the user using cursor C from among the characters displayed in the character selection area 54 (see FIG. 5) as the frame to be displayed in the frame being edited by the user. The image of the character determined by the character determination unit 22c is superimposed on the preview image of the frame displayed in the preview display area 53. As in the example shown in FIG. 5, when the "Hide" button is selected in the character selection area 54, no character is displayed in the preview display area 53, but when any character is selected from the character selection area 54, the character selection is determined by the character determination unit 22c, and the character is displayed as in the example shown in FIG. As shown in the figure, an image of character A selected by the user is superimposed on the preview image of the frame displayed in preview display area 53. Character A displayed in preview display area 53 can be moved or enlarged or reduced in size within preview display area 53 using cursor C. For example, if a mouse is used as input device 205 of user terminal 2, character A may be moved to any position by operating cursor C with the mouse and dragging character A within preview display area 53. Character A may also be enlarged or reduced in size by operating the mouse wheel during a drag operation.
[0043] The image of character A displayed in preview display area 53 may be either a moving image or a still image, but if the motion determination unit described below determines that the character's motion (action) is "do nothing," the image of the character displayed in preview display area 53 will be a still image.
[0044] The facial expression determination unit 22d determines the "facial expression" of the character A displayed in the preview display area 53 in response to an operation on the facial expression setting unit 55 by the user.
[0045] FIG. 14 is a diagram illustrating an example of the facial expression setting unit 55. As shown in FIG. 14, the facial expression setting unit 55 includes a two-dimensional operation area 55a (first two-dimensional operation area) and a pointer 55b that is movable on the two-dimensional operation area 55a. The facial expression setting unit 55 can set a facial expression (joy, anger, sadness, or happiness) of the character A displayed in the preview display area 53 according to the position of the pointer 55b on the two-dimensional operation area 55a. Specifically, the more the pointer 55b is moved toward the "happy" side (the +X direction in FIG. 14) on the two-dimensional operation area 55a of the facial expression setting unit 55, the more the facial expression of the character A displayed in the preview display area 53 changes to an expression that expresses "joy." Conversely, the more the pointer 55b is moved toward the "sad" side (the -X direction in FIG. 14) on the two-dimensional operation area 55a of the facial expression setting unit 55, the more the facial expression of the character A displayed in the preview display area 53 changes to an expression that expresses "sadness." Furthermore, the more the pointer 55b is moved toward the "joy" side (the +Y direction in FIG. 14) on the two-dimensional operation area 55a of the expression setting unit 55, the more the expression of the character A displayed in the preview display area 53 changes to an expression that expresses "joy." Conversely, the more the pointer 55b is moved toward the "anger" side (the -Y direction in FIG. 14) on the two-dimensional operation area 55a of the expression setting unit 55, the more the expression of the character A displayed in the preview display area 53 changes to an expression that expresses "anger." In this way, the expression determination unit 22d determines the "expression" of the character A displayed in the preview display area 53 according to the position of the pointer 55b on the two-dimensional operation area 55a of the expression setting unit 55. This allows the user of the content creation system 1 to set the "expression" of the character in the content by intuitive operations.
[0046] 7 shows the state after the user has decided on the "expression" of character A by operating expression setting unit 55 using cursor C. As shown in FIG. 7, pointer 55b on two-dimensional operation area 55a of expression setting unit 55 has been moved toward the "joy" side, and therefore the expression of character A displayed in preview display area 53 has changed to an expression expressing "joy."
[0047] Returning to FIG. 3, the line of sight determination unit 22e determines the "line of sight" of the character A displayed in the preview display area 53 in response to an operation on the line of sight setting unit 56 by the user.
[0048] FIG. 15 is a diagram showing an example of the line of sight setting unit 56. As shown in FIG. 15, the line of sight setting unit 56 includes a two-dimensional operation area 56a (second two-dimensional operation area) and a pointer 56b that can move on the two-dimensional operation area 56a. The line of sight setting unit 56 can set the line of sight (direction of line of sight) of the character A displayed in the preview display area 53 according to the position of the pointer 56b on the two-dimensional operation area 56a. Specifically, the two-dimensional operation area of the line of sight setting unit 56 15) on the two-dimensional operation area 56a of the eye line setting unit 56, the more the pointer 56b is moved to the "left" side (the -X direction in FIG. 15) on the two-dimensional operation area 56a of the eye line setting unit 56, the more the eye line of the character A displayed in the preview display area 53 will be directed to the right (the character's eyes will move to the right as you face the screen). Conversely, the more the pointer 56b is moved to the "left" side (the -X direction in FIG. 15) on the two-dimensional operation area 56a of the eye line setting unit 56, the more the eye line of the character A displayed in the preview display area 53 will be directed to the left (the character's eyes will move to the left as you face the screen). Furthermore, the more the pointer 56b is moved to the "up" side (the +Y direction in FIG. 15) on the two-dimensional operation area 56a of the eye line setting unit 56, the more the eye line of the character A displayed in the preview display area 53 will be directed upward. Conversely, the more the pointer 56b is moved to the "down" side (the -Y direction in FIG. 15) on the two-dimensional operation area 56a of the eye line setting unit 56, the more the eye line of the character A displayed in the preview display area 53 will be directed downward. In this way, the line of sight determination unit 22e determines the "line of sight" of character A displayed in the preview display area 53 according to the position of the pointer 56b on the two-dimensional operation region 56a of the line of sight setting unit 56. This allows the user of the content creation system 1 to set the "line of sight" of the character in the content by an intuitive operation.
[0049] 8 shows the state after the user has determined the "line of sight" of character A by operating line of sight setting unit 56 using cursor C. As shown in Fig. 8, pointer 56b on two-dimensional operation area 56a of line of sight setting unit 56 has been moved to the "right" side, so the line of sight of character A displayed in preview display area 53 has changed to face right (to the right as you face the screen).
[0050] Returning to FIG. 3, the motion determination unit 22f determines the motion (movement) of the character A displayed in the preview display area 53 in response to an operation on the motion setting unit 57 by the user.
[0051] Fig. 16 is a diagram showing an example of the motion setting unit 57. As shown in Fig. 16, the motion setting unit 57 includes a pull-down menu 571 for selecting a motion type of a character, and a playback range setting unit 572 for setting a playback range for the selected motion type. The playback range setting unit 572 includes a first shaft 572a, a first knob 572b that can move on the first shaft 572a, a second shaft 572c, and a second knob 572d that can move on the second shaft 572c.
[0052] In the motion setting unit 57, when a pull-down menu 571 is selected with a cursor C (not shown in FIG. 16), a list of settable motion types is displayed. Then, by selecting a desired motion type from the motion types displayed in the list, the motion of the character is determined by the motion determination unit 22f.
[0053] In addition, the shafts (first shaft 572a and second shaft 572c) of the playback range setting unit 572 represent the time axis of a preset motion, and by moving the knobs (first knob 572b and second knob 572d) on the shafts, it is possible to set a motion within a desired playback range.
[0054] Fig. 9 is a diagram showing a state in which a list of motion types is displayed after selecting pull-down menu 571 of motion setting unit 57 with cursor C. For example, when "wave" is selected from the list of motion types shown in Fig. 9, motion determination unit 22f determines the character's motion to be "wave", and as shown in Fig. 10, character A displayed in preview display area 53 dynamically changes to show a "wave" motion (animation).
[0055] Returning to FIG. 3, the character placement determination unit 22g determines the character placement setting by the user. The position of character A displayed in preview display area 53 is determined in accordance with an operation on character position setting unit 58 (see FIG. 11). Character position setting unit 58 includes, for example, setting items for "camera," "character position," "rotation," and "zoom out / in," as shown in FIG. 11. The "camera" setting item allows the setting of the camera's angle of view, such as the character's "full body," "upper body," or "head." The "character position" setting item allows the display position of the character in preview display area 53 to be set. The "rotation" setting item allows character A in preview display area 53 to be rotated in the yaw direction or pitch direction at predetermined angular intervals (e.g., 90-degree intervals). Furthermore, the "zoom out / in" setting item allows character A displayed in preview display area 53 to be enlarged or reduced by moving a knob on a shaft.
[0056] Returning to FIG. 3, background image determination unit 22h determines the background image to be displayed in preview display area 53 in response to a user's operation on background image setting unit 59 (see FIG. 12).
[0057] 12 is a diagram showing a state in which a background image is set by a user operation on background image setting unit 59 and the background image is determined by background image determination unit 22h. As shown in Fig. 12, background image setting unit 59 includes setting items such as "background," "background position," "background size," and "lighting."
[0058] In the "Background" setting item, it is possible to select from multiple pre-prepared background images (two-dimensional or three-dimensional images, still or moving images), as well as "monochrome" (no background image). Also, by selecting the "Add" button, the user can add any background image. Furthermore, by selecting the "Auto-generate" button, it is possible to automatically generate a background image using a trained image processing model. Specifically, in response to the "Auto-generate" button being selected, a background image generation unit (not shown) of the control unit 32 inputs the character selected by the user, the character's facial expression, line of sight, motion, text information (described later), etc. as input data into the trained image processing model, thereby generating a background image. This image processing model is trained using, for example, generative adversarial networks (GAN). It should be noted that still images such as landscape photos prepared in advance by the user can also be generated. A still image may be input as input data into an image processing model to generate a background image. This allows a background desired by the user to be automatically generated. Here, if the background image is generated as three-dimensional data (three-dimensional image), the three-dimensional data of the background image may be placed in the metaverse space together with the three-dimensional data of the character generated by the character image generation unit 32a.
[0059] The "Background Position" setting item allows the position of the background image area 53a on the preview display area 53 to be set. Specifically, the position of the background image area 53a on the preview display area 53 can be set by moving the pointer on the two-dimensional operation area. The "Background Size" setting item allows the size of the background image area 53a on the preview display area 53 to be set. Specifically, the size of the background image area 53a on the preview display area 53 can be enlarged or reduced by moving the knob on the shaft. The "Lighting" setting item allows the selection of "Daytime," "Evening," "Nighttime," etc. from a pull-down menu, allowing the depiction of the background image to be set according to the time of day.
[0060] Returning to FIG. 3, the text information determination unit 22i determines the text information to be displayed in the preview display area 53 in response to a user's operation on the text information setting unit 60 (see FIG. 13).
[0061] FIG. 13 is a diagram illustrating a state in which text information is set by a user operation on the text information setting unit 60 and the text information is determined by the text information determination unit 22i. As shown in FIG. 13, the text information setting unit 60 includes setting items such as “speech bubble,” “font,” “text,” “text color,” and “character size.” The “speech bubble” setting item allows a user to select from multiple pre-prepared speech bubble images, as well as “hide” (no text information) and “text only” (no speech bubble image). The “font” setting item allows a user to select the font type of the text to be displayed in the preview display area 53 from a pull-down menu. The “text” setting item allows a user to input text to be displayed in the preview display area 53 through a user input operation via the input unit 23 (described later). The “text color” setting item allows a user to set the color of the text to be displayed in the preview display area 53. The “character size” setting item allows a user to set the size of the text to be displayed in the preview display area 53. When the text information is set by the text information setting unit 60, the text information determination unit 22i determines the text information, and the text information is displayed in the preview display area 53. In the example shown in FIG. 13, a speech bubble image 61 is displayed in preview display area 53, in which the text "Hello" input by the user in the "Text" setting item of text information setting section 60 is combined with a speech bubble image selected by the user in the "Speech Bubble" setting item of text information setting section 60. This speech bubble image 61 can be moved within preview display area 53 by a user operation (e.g., a drag operation) using cursor C. In addition, the size of speech bubble image 61 can be changed (enlarged or reduced) by operating (e.g., a drag operation) a zoom icon D displayed near speech bubble image 61. When an area outside preview display area 53 is selected with cursor C, the position and size of speech bubble image 61 in preview display area 53 are fixed, and zoom icon D disappears.
[0062] Returning to FIG. 3, the input unit 23 is an element for the user of the user terminal 2 to input information, and is, for example, a keyboard, a mouse, a touch panel, a microphone, a gesture input device, or the like.
[0063] The output unit 24 is an interface that outputs various information (images and sounds) from the user terminal 2 to the user, and is, for example, a video display device (display unit) such as a liquid crystal display, or a speaker. When the output unit 24 is configured as a display unit, a GUI for receiving operations from the user is displayed on this display unit by the display control unit 22a.
[0064] The imaging unit 25 captures an image of a subject using a camera module including an imaging element of the user terminal 2. The image captured by the imaging unit 25 is output to the control unit 22, where it is subjected to various image processing, and is displayed on the display unit (output unit 24) by the display control unit 22a, or transmitted to the server 3 via the network 4 by the communication unit 21.
[0065] The storage unit 26 is, for example, a data storage such as an internal memory or an external memory (such as an SD memory card). The storage unit 26 stores various data handled by the control unit 22, various information downloaded by the communication unit 21 from the server 3 via the network 4, images captured by the imaging unit 25, etc. Note that the storage unit 26 does not necessarily have to be provided within the user terminal 2, and part or all of the storage unit 26 may be provided in another device that is communicably connected to the user terminal 2 via the network 4.
[0066] Next, a description will be given of the functional configuration of the server 3. As shown in FIG.
[0067] The communication unit 31 is a communication interface between the server 3 and the network 4. 31 transmits and receives information between the server 3 and the user terminal 2 via the network 4 .
[0068] The control unit 32 controls the overall operation of the server 3. The control unit 32 also has a character image generation unit 32a, a voice generation unit 32b, a content generation unit 32c, a speech data generation unit 32d, and an NFT issuing unit 32e.
[0069] The character image generation unit 32a generates image data (two-dimensional or three-dimensional image data) of a character based on an image received from the user terminal 2 (for example, a face image or a full-body image of the user captured using the imaging unit 25, or a character image prepared in advance). For example, the character image generation unit 32a may generate a character image by using an image generation AI. Specifically, the character image generation unit 32a may generate a character image by using, for example, a generative adversarial network (GAN) for learning. A character image may be generated by inputting image data input by a user via the user terminal 2 into the generated image processing model. The generated character image may be stored in a character image DB 33a of the storage unit 33, which will be described later, or may be transmitted to the user terminal 2 via the network 4 by the communication unit 31.
[0070] The voice generation unit 32b generates voice data corresponding to the character image (two-dimensional image or three-dimensional image) generated by the character image generation unit 32a. Specifically, the voice generation unit 32b inputs the character image generated by the character image generation unit 32a into a trained voice synthesis model and outputs voice data. This makes it possible to automatically generate voice data with a voice tone that matches the impression of the character's appearance, atmosphere, and personality from the character image. Note that the voice model is trained using, for example, a large number of character images and voice data associated with the character images (such as lines spoken by animated characters) as training data.
[0071] Furthermore, the voice generation unit 32b performs text-to-speech synthesis to generate voice based on the voice data generated by the voice generation unit 32b, the text determined by the text information determination unit 22i, and the text generated by the speech data generation unit 32d (described later). The voice generated by the voice generation unit 32b through text-to-speech synthesis is output as a voice file.
[0072] The content generation unit 32c generates content (manga, animation, etc.) consisting of one or more frames edited by a user through various interfaces displayed on the display unit (output unit 24) of the user terminal 2. For example, the content generation unit 32c generates content data in response to a selection operation of a play button 52c in a frame list area 52 of a content creation screen 50 shown in FIG. 5. Here, the content data refers to data in which one or more frames edited by a user are encapsulated into a single file, such as a video file such as MP4 or AVI, or a still image file such as a High Efficiency Image File Format (HEIF) or AVIF (AV1 Image File Format) that can store multiple images. If an audio file is generated by the audio generation unit 32b, the audio file is also included. The generated content data is stored in a content DB 33e of the storage unit 33 (described later) and transmitted to the user terminal 2 via the communication unit 31 via the network 4.
[0073] The utterance data generation unit 32d generates utterance data of a character that responds to the user's utterance data (voice data or text data) input via the input unit 23 of the user terminal 2. This allows the character generated by the character image generation unit 32a to function as a conversational bot. Specifically, the utterance data generation unit 32d converts the user's utterance data input via the input unit 23 of the user terminal 2 into text data and inputs it into a language model (large-scale language model), generating utterance data of the character that responds to the user's utterance. The speech data is generated by outputting the character's image data and / or voice data. Here, the language model may be trained using the character's image data and / or voice data as part of the training data. By inputting the character's image and voice data into the language model along with the user's speech data, speech data that matches the impression of the character's appearance, atmosphere, and personality can be generated. The character's speech data generated by the speech data generation unit 32d is transmitted to the user terminal 2 via the network 4 by the communication unit 31, and is output as sound from the output unit 24 (speaker) of the user terminal 2.
[0074] The NFT issuing unit 32e issues an NFT (Non-Fungible Token) related to the content data generated by the content generating unit 32c or image data of a character within the content (data of a character image (two-dimensional image or three-dimensional image) generated by the character image generating unit 32a). That is, the NFT issuing unit 32e converts the content generated by the content generating unit 32c and / or the image data of a character within the content into an NFT. Specifically, for example, when the NFT issuing unit 32e receives an instruction to convert the content data or the character image data into an NFT in response to an operation on an interface (not shown) displayed on the display unit (output unit 24) of the user terminal 2, the NFT issuing unit 32e generates an NFT ID linked to the content data or the character image data. Then, the NFT issuing unit 32e generates a code having metadata including the generated NFT ID and transmits it to a blockchain network (not shown). This allows the content data or the character image data within the content to be converted into an NFT. In addition, users will be able to buy and sell NFT content data and character image data on NFT exchanges, or exchange them for other NFT data.
[0075] The storage unit 33 is, for example, a data storage such as an internal memory or an external memory (such as an SD memory card). The storage unit 33 stores various data handled by the control unit 32, various information received by the communication unit 31 from the user terminal 2 via the network 4, various databases (DBs), and the like. The databases include a character image DB 33a, a background image DB 33b, a speech bubble image DB 33c, and a content DB 33d.
[0076] The character image DB 33a stores image data (including 2D or 3D model data) of multiple characters prepared in advance. The character image DB 33a also stores character images generated by the character image generation unit 32a of the control unit 32. The character images stored in the character image DB 33a may be associated with multiple motion patterns. In response to a request from the user terminal 2, the character images stored in the character image DB 33a are transmitted to the user terminal 2 by the communication unit 31 via the network 4, and are displayed on the display unit (output unit 24) by the display control unit 22a of the user terminal 2.
[0077] The background image DB 33b stores image data of a plurality of background images prepared in advance. In response to a request from the user terminal 2, the background images stored in the background image DB 33b are transmitted to the user terminal 2 via the network 4 by the communication unit 31, and are displayed on the display unit (output unit 24) by the display control unit 22a of the user terminal 2.
[0078] The balloon image DB 33c stores image data of a plurality of balloon images prepared in advance. In response to a request from the user terminal 2, the balloon images stored in the balloon image DB 33c are transmitted to the user terminal 2 via the network 4 by the communication unit 31, and are displayed on the display unit (output unit 24) by the display control unit 22a of the user terminal 2.
[0079] The content DB 33d stores data of the content generated by the content generating unit 32c. The content data stored in the content DB 33d is stored in the user terminal In response to the request from user 2, the image data is transmitted to the user terminal 2 via the network 4 by the communication unit 31, and is displayed on the display unit (output unit 24) by the display control unit 22a of the user terminal 2.
[0080] In addition, the memory unit 33 does not necessarily have to be provided within the server 3, and part or all of the memory unit 33 may be provided within another device that is communicatively connected to the server 3 via the network 4.
[0081] (Example of operation) Next, we will explain an example of the operation of the content creation system 1. Fig. 4 is a flowchart showing the operation of a user of the content creation system 1 to create content made up of one or more frames using a user interface displayed on the display unit (output unit 24) of the user terminal 2.
[0082] First, in the user terminal 2, the display control unit 22a displays a user interface on the display unit (output unit 24) (step S1).
[0083] Next, in the user terminal 2, the frame determining unit 22b determines a frame to be edited in response to a user operation on the user interface displayed on the display unit (output unit 24) (step S2).
[0084] Next, in the user terminal 2, the character determination unit 22c determines a character in response to a user operation on the user interface displayed on the display unit (output unit 24) (step S3).
[0085] Next, in the user terminal 2, the facial expression determination unit 22d determines the facial expression of the character in response to a user operation on the user interface (facial expression setting unit 55) displayed on the display unit (output unit 24) (step S4).
[0086] Next, in the user terminal 2, the line of sight determination unit 22e determines the line of sight of the character in response to a user operation on the user interface (line of sight setting unit 56) displayed on the display unit (output unit 24) (step S5).
[0087] Next, in the user terminal 2, the motion determination unit 22f determines the motion of the character in response to a user operation on the user interface (motion setting unit 57) displayed on the display unit (output unit 24) (step S6).
[0088] Next, in the user terminal 2, the character placement determination unit 22g determines the placement of the characters in response to a user operation on the user interface (character placement setting unit 58) displayed on the display unit (output unit 24) (step S7).
[0089] Next, in the user terminal 2, the background image determination unit 22h determines a background image in response to a user operation on the user interface (background image setting unit 59) displayed on the display unit (output unit 24) (step S8).
[0090] Next, in the user terminal 2, the text information determination unit 22i determines text information in response to a user operation on the user interface (text information setting unit 60) displayed on the display unit (output unit 24) (step S9).
[0091] Next, in the server 3, the content generating unit 32c generates content including one or more frames (step S10).
[0092] As described above, according to this embodiment, the content generation system 1 is a content generation system 1 that generates content consisting of one or more frames including moving or still images of a character, and is equipped with a user interface (facial expression setting unit 55) including a two-dimensional operation area 55a, a display control unit 22a that displays a preview image (preview display area 53) of the frame being edited by the user on the display unit (output unit 24), an facial expression determination unit 22d that determines the facial expression of the character on the preview image depending on the position of pointer 55b on the two-dimensional operation area 55a, and a content generation unit 32c that generates content consisting of one or more frames including the frame edited by the user, so that content can be easily generated that is made up of frames including characters with facial expressions intended by the user.
[0093] In this embodiment, an example will be described in which facial expressions corresponding to four emotions, "joy," "anger," "sorrow," and "happiness," are set by the facial expression setting unit 55, but facial expressions corresponding to four or more emotions may be set. For example, facial expressions corresponding to eight emotions (Plutchik model), "anger," "fear," "expectation," "surprise," "joy," "sadness," "trust," and "disgust," may be set.
[0094] Furthermore, in this embodiment, as an example, an example of setting a character's motion using the pull-down menu 571 has been described, but it is also possible to associate motions corresponding to facial expressions that can be set by the facial expression setting unit 55 in advance, so that the motion corresponding to the facial expression set by the facial expression setting unit 55 is automatically set.
[0095] Furthermore, in the present embodiment, as an example, the facial expression setting unit 55 and the motion setting unit 57 are provided separately, and the "facial expression" and "motion" are set by each setting unit. However, the "facial expression" and "motion" may be set by a single setting unit. For example, a single emotion setting unit having a two-dimensional operation area in which multiple emotions can be set as a user interface may be provided, and the "facial expression" and "motion" of a character corresponding to the emotions set by this emotion setting unit may be associated in advance, so that the "facial expression" and "motion" corresponding to the emotions set by the emotion setting unit may be automatically set. Here, the types of emotions set by the emotion setting unit may be, for example, the four emotions of "joy," "anger," "sadness," and "pleasure," or the eight emotions (Plutchik model) of "anger," "fear," "expectation," "surprise," "joy," "sadness," "trust," and "disgust." Alternatively, more types of emotions may be set.
[0096] In addition, in this embodiment, an example has been described in which the facial expression setting unit 55 has a rectangular two-dimensional operation area 55a, but the shape of the two-dimensional operation area is not limited to this. For example, as shown in Fig. 17, the facial expression setting unit 55 may have a circular two-dimensional operation area 55c, and each facial expression such as joy, anger, sadness, or happiness may be set by moving a dot-shaped pointer 55d within this circular two-dimensional operation area 55c, or the shape of the two-dimensional operation area may be a shape other than rectangular or circular, and each facial expression may be set by moving a pointer within the shape.
[0097] In addition, in this embodiment, an example has been described in which the gaze setting unit 56 has a rectangular two-dimensional operation area 56a, but the shape of the two-dimensional operation area is not limited to this. For example, as shown in Fig. 18, the gaze setting unit 56 may have a circular two-dimensional operation area 56c, and the gaze in each direction may be set by moving a dot-shaped pointer 56d within this circular two-dimensional operation area 56c, or the shape of the two-dimensional operation area may be a shape other than rectangular or circular, and the gaze in each direction may be set by moving the pointer within that shape.
[0098] Any part or all of the functional units described in this specification may be realized by a program. The software may be distributed by being non-temporarily recorded on a computer, or via a communication line (including wireless communication) such as the Internet, or may be distributed in a state where it is installed on any terminal.
[0099] Based on the above description, a person skilled in the art may be able to conceive additional effects and various modifications of the present invention, but the aspects of the present invention are not limited to the individual embodiments described above. Various additions, modifications, and partial deletions are possible within the scope of the conceptual idea and spirit of the present invention, which is derived from the content defined in the claims and their equivalents.
[0100] For example, what is described herein as a single device (or component, the same applies hereinafter) (including what is depicted as a single device in the drawings) may be realized by multiple devices. Conversely, what is described herein as multiple devices (including what is depicted as multiple devices in the drawings) may be realized by a single device. Alternatively, some or all of the means and functions included in a certain device (e.g., a server) may be included in another device (e.g., a user terminal).
[0101] Furthermore, not all of the features described in this specification are essential requirements. In particular, features described in this specification but not included in the claims can be considered optional additional features.
[0102] It should be noted that the applicant is merely aware of the inventions disclosed in the documents listed in the "Prior Art Documents" section of this specification, and the present invention does not necessarily aim to solve the problems of the disclosed inventions. The problem that the present invention aims to solve should be determined by taking into consideration the entire specification. For example, if this specification states that a specific configuration achieves a certain effect, it can also be said that the present invention solves a problem that is the reverse of that effect. However, it is not necessarily intended that such a specific configuration be an essential requirement. [Explanation of symbols]
[0103] 1 Content Generation System 2. User terminal 21 Communications Department 22 Control Unit 22a Display control unit 22b Piece determination section 22c Character Selection Section 22d Facial expression determination section 22e Eye Line Determination Department 22f Motion determination section 22g Character placement determination section 22h Background image determination section 22i Text information determination section 23 Imaging unit 24 Input section 25 Output section (display section) 26 Memory section 3 Server 31 Communications Department 32 Control section 32a Character image generation unit 32b Voice generation unit 32c Content Generation Unit 32d Speech data generation unit 32e NFT Issuance Department 33 Storage section 33a Character Image DB 33b Motion pattern DB 33c Background image DB 33d speech bubble DB 33e Content DB
Claims
1. A content generation system that generates content consisting of one or more frames including moving or still images of a character, a display control unit that displays, on a display unit, a user interface including a first two-dimensional operation area and a preview image of a frame to be edited by a user; an expression determination unit that determines an expression of the character on the preview image in accordance with the position of a pointer on the first two-dimensional operation area; a motion determination unit that determines a motion of the character on the preview image in accordance with a position of a pointer on the first two-dimensional operation area; a content generation unit that generates content consisting of one or more frames including the frame edited by the user; A content generation system comprising:
2. The content generation system according to claim 1 , further comprising a motion determination unit that determines a motion of the character on the preview image in response to an operation of the user interface by the user.
3. the user interface includes a one-dimensional operation area; The content generation system according to claim 2 , wherein the motion determination unit determines a reproduction range of the motion depending on a position of the knob on the one-dimensional operation area.
4. The content generation system according to claim 1 , further comprising a character determination unit that determines a character to be displayed on the preview image.
5. The content generation system of claim 1 , further comprising a background image determination unit that determines a background image for the preview image.
6. 2. The content generation system according to claim 1, further comprising a text information determination unit that determines at least one of the text to be displayed in the preview image, the type of speech bubble in which the text is displayed, the color of the text, and the size of the text in response to an operation of the user interface by the user.
7. The content generation system according to claim 1 , wherein the first two-dimensional operation area is a user interface for setting an emotion of the character.
8. The content generation system of claim 1, further comprising a voice generation unit that generates voice data for a character by inputting image data of the character selected by the user into a voice model that uses a character image and voice data associated with the character image as training data.
9. The content generation system of claim 1 further comprises a speech data generation unit that generates speech data for a character that responds to the user's speech data by inputting the user's speech data into a language model trained using the character's image data and / or voice data as part of the training data.
10. The content generation system according to claim 1 , further comprising a character image generation unit that generates a character image by inputting the image data input by the user into a trained image processing model.
11. The content generation system according to claim 10 , further comprising a background image generation unit that generates a background image by inputting the image data input by the user into a trained image processing model.
12. The content generation system according to claim 11 , wherein the image data generated by the character image generation unit and the background image generation unit is three-dimensional image data and is arranged in a metaverse space.
13. The content generation system according to claim 10 , further comprising an NFT issuing unit that issues an NFT related to image data of the content generated by the content generation unit or a character in the content.
14. A method for generating content consisting of one or more frames including a moving or still image of a character, comprising: a display control step of displaying on a display unit a user interface including a first two-dimensional operation area and a preview image of a frame to be edited by the user; an expression determination step of determining an expression of the character on the preview image in accordance with the position of a pointer on the first two-dimensional operation area; a motion determination step of determining a motion of the character on the preview image according to a position of a pointer on the first two-dimensional operation area; a content generation step of generating content consisting of one or more frames including the frame edited by the user; A method comprising:
15. A program for causing a computer to execute a method for generating content consisting of one or more frames including a moving image or a still image of a character, the method comprising: a display control step of displaying on a display unit a user interface including a first two-dimensional operation area and a preview image of a frame to be edited by the user; an expression determination step of determining an expression of the character on the preview image in accordance with the position of a pointer on the first two-dimensional operation area; a motion determination step of determining a motion of the character on the preview image according to a position of a pointer on the first two-dimensional operation area; a content generation step of generating content consisting of one or more frames including the frame edited by the user; A program having:
16. A computer-readable storage medium storing the program according to claim 15.
17. A content generation system that generates content consisting of one or more frames including moving or still images of a character, a display control unit that displays, on a display unit, a user interface including a first two-dimensional operation area and a preview image of a frame to be edited by a user; a character determination unit that determines a character to be displayed on the preview image in response to an operation of the user interface by the user; an expression determination unit that determines an expression of the character in accordance with a position of a pointer on the first two-dimensional operation area; a line of sight determination unit that determines a line of sight of the character in response to an operation of the user interface by the user; a motion determination unit that determines a motion of the character in accordance with a position of a pointer on the first two-dimensional operation area; a background determination unit that determines a background image of the preview image in response to an operation of the user interface by the user; a text information determination unit that determines at least one of text to be displayed in the preview image, a type of speech bubble displaying the text, a color of the text, and a size of the text in response to an operation of the user interface by the user; a content generation unit that generates content consisting of one or more frames including the frame edited by the user; A content generation system comprising:
Citation Information
Patent Citations
Multimodal Model for Dynamically Responsive Virtual Characters
JP2022534708A
JPP7633357B