Information processing method, information processing system, and information processing program
The information processing system addresses the challenge of maintaining virtual actor consistency by using work and persona data to generate content aligned with the actor's traits, ensuring appropriate utilization and dialogue, thereby enhancing content quality.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SONY GROUP CORP
- Filing Date
- 2025-10-24
- Publication Date
- 2026-05-07
AI Technical Summary
Existing technologies struggle to maintain the consistency of virtual actors' appearances, voices, and personalities across different works, leading to inappropriate actions and limited utilization in content generation, and lack the ability for virtual actors to comment on their performances or discuss their roles.
An information processing system that acquires work information and virtual actor persona data to generate content suitable for the actor, incorporating prohibited actions and allowing user interaction to correct deviations, using a machine learning model to produce content that aligns with the actor's persona.
Ensures that generated content adheres to the virtual actor's characteristics, preventing inappropriate actions and enabling meaningful dialogue, thus effectively utilizing virtual actors while maintaining their intended image.
Smart Images

Figure JP2025037491_07052026_PF_FP_ABST
Abstract
Description
Information Processing Method, Information Processing System, and Information Processing Program
[0001] The present disclosure relates to an information processing method, an information processing system, and an information processing program.
[0002] Techniques for generating content by using a machine learning model are known. For example, there is a known technique for generating a chatbot in which a character responds to a user's utterance by inputting a prompt based on the character of a chatbot, which is an example of content, and the user's utterance into a large language model, which is an example of a machine learning model (for example, Patent Document 1).
[0003] Japanese Patent Application Laid-Open No. 2022-180282
[0004] In the prior art, with respect to content such as a chatbot, the image of the character can be maintained by modifying the chatbot so as to conform to the image of the character.
[0005] By the way, when generating a character with a virtual personality using a large language model as in the prior art, there may be a case where a character with a common appearance and voice is to appear in a plurality of works. This means that, so to speak, a virtual actor (virtual performer) appears while changing the roles played in a plurality of works. When the recognition of such a virtual performer increases, thereafter, a moving image (for example, a concert video) using the virtual performer can be generated and its value can be enhanced.
[0006] At this time, it is desirable for the creator side of the virtual performer to define not only the appearance and voice of the virtual performer but also the personality and virtual background (place of origin, favorite food, etc.) and maintain the image of the virtual performer. That is, even when the virtual performer appears in the generated moving image as described above, an act that deviates from the image of the virtual performer should be avoided.
[0007] However, even if an action is not unnatural for a character within the content, if it is inappropriate for a virtual actor, it has traditionally been difficult to prevent the creation of such footage. Furthermore, virtual actors were unable to comment on their own performances when participating in fan events or other discussions about the anime works in which they appeared, nor were they able to discuss their acting with directors or co-stars, nor could these discussions be reflected in the content. In other words, traditionally, it has been difficult to appropriately utilize virtual actors while simultaneously restricting their deviant behavior.
[0008] Therefore, this disclosure aims to propose an information processing method, information processing system, and information processing program that can appropriately utilize virtual actors.
[0009] The information processing method relating to this disclosure includes an acquisition step of acquiring work information relating to content and virtual actor persona information, which is information of a virtual actor who plays a character appearing in said content, and a content generation step of generating said content, which includes said character played by said virtual actor, based on the acquired work information and virtual actor persona information.
[0010] This is a diagram illustrating the outline of the information processing system according to the first embodiment. This is a diagram (1) showing an example of work information. This is a diagram (1) showing an example of virtual actor persona information. This is a diagram (1) showing an example of input information. This is a diagram showing an example of dialogue history. This is a diagram (2) showing an example of input information. This is a block diagram showing an example of the configuration of the information processing system according to the first embodiment. This is a diagram (1) illustrating an example of input information generation. This is a diagram (2) illustrating an example of input information generation. This is a diagram illustrating an example of work information prompt generation. This is a diagram illustrating an example of character information prompt generation. This is a diagram illustrating an example of dialogue input generation. This is a diagram (3) illustrating an example of input information generation. This is a flowchart illustrating an example of the information processing flow according to the first embodiment. This is a diagram (2) showing an example of work information. This is a diagram (2) showing an example of virtual actor persona information. This is a diagram illustrating an example of work information generation. This is a block diagram illustrating an example of the configuration of the information processing system according to the second embodiment. This is a diagram (4) illustrating an example of input information generation. This is a diagram (5) illustrating an example of input information generation. This is a diagram illustrating an example of dialogue history prompt generation. This is a diagram (6) illustrating an example of input information generation. This is a flowchart illustrating an example of the information processing flow according to the second embodiment. This is a block diagram illustrating an example of the configuration of an information processing system according to the third embodiment. This is Figure (7) illustrating an example of input information generation. This is a hardware configuration diagram showing an example of a computer that realizes the functions of an information processing device.
[0011] The embodiments of this disclosure will be described in detail below with reference to the drawings. In the following embodiments, the same parts will be denoted by the same reference numerals, and redundant descriptions will be omitted.
[0012] The embodiments of this disclosure will be described below in the following order: 1. First Embodiment 1-1. Overview of the Information Processing System According to the First Embodiment 1-2. Configuration of the Information Processing System According to the First Embodiment 1-3. Flow of Information Processing According to the First Embodiment 2. Modifications 2-1. First Modification 2-2. Second Modification 2-3. Third Modification 2-4. Fourth Modification 2-5. Fifth Modification 3. Second Embodiment 3-1. Configuration of the Information Processing System According to the Second Embodiment 3-2. Flow of Information Processing According to the Second Embodiment 4. Third Embodiment 5. Other Embodiments 6. Effects of the Information Processing System According to this Disclosure 7. Hardware Configuration 8. Notes
[0013] (1. First Embodiment) (1-1. Overview of the Information Processing System According to the First Embodiment) An overview of the information processing system 10 according to the first embodiment will be explained using Figure 1. Figure 1 is a diagram for explaining the overview of the information processing system according to the first embodiment.
[0014] The information processing system 10 includes an information processing device 100 and a terminal device 200. The information processing device 100 is, for example, a computer such as a server. The information processing device 100 is a content generation device that generates content 300 such as animated videos and dramas featuring virtual characters. For example, the information processing device 100 generates animated videos and dramas as content 300 by inputting input information 400, which is information for generating content 300 such as the content of an animated video or a drama, into a machine learning model such as a large-scale language model.
[0015] In this case, the information processing device 100 first generates input information 400 based on information for generating input information 400. For example, the information processing device 100 generates input information 400 such as prompts based on information for generating input information 400, such as work information 500 related to content 300 and virtual actor persona information 600 related to virtual characters appearing in content 300. Subsequently, the information processing device 100 generates content 300 by inputting the input information 400 into a machine learning model.
[0016] A machine learning model is, for example, a large-scale language model that outputs content 300 in response to input information 400. Specifically, a machine learning model is a Multi-Modal Large Language Model (MMLLM) that is tuned to output content 300 such as moving images or videos, such as anime or dramas featuring virtual characters.
[0017] Content 300 is, for example, a video or image featuring a virtual human as a virtual character. One example of Content 300 is an anime or drama featuring an actor (virtual actor) as a virtual human. Another example of Content 300 is a video or image of a concert or fan event featuring a virtual actor.
[0018] In other words, Content 300 in this disclosure refers to video footage, etc., consisting of a virtual actor with a virtual personality playing a virtual character. Hereinafter, for the sake of distinction, the virtual character as a character (cast) in the content will be referred to as the "character," and the actor with a virtual personality who plays that character will be referred to as the "virtual actor."
[0019] The input information 400 is, for example, a prompt indicating the prerequisites and context of the content 300. Specifically, the input information 400 is natural language text, etc., which includes at least the work information 500 and the virtual actor persona information 600.
[0020] The work information 500 consists of materials related to the content 300, such as the scenario, storyboards, and images. Below, an example of work information 500 will be explained using Figure 2. Figure 2 is Figure (1), which shows an example of work information.
[0021] In Figure 2, the work information 500 is text that shows the scenario for Scene 1, which is part of the video footage that is an example of the content 300. The work information 500 may also include setting information for character A, who appears in Scene 1. Furthermore, the work information 500 may also include information that specifies the character's actions, such as in a smoking scene 510 which states that "the character has a cigarette in their mouth."
[0022] Let's return to the explanation of Figure 1. The virtual actor persona information 600 is, for example, the characteristics of the personality (persona) set for the virtual actor. Specifically, the virtual actor persona information 600 is information that shows the virtual actor's profile, such as their appearance, voice, personality, and virtual background (place of origin, favorite food).
[0023] Below, an example of virtual actor persona information 600 will be explained using Figure 3. Figure 3 is Figure (1) showing an example of virtual actor persona information 600. In Figure 3, virtual actor persona information 600 is text that includes the appearance of the virtual actor and prohibition information 610 regarding prohibited actions and statements for the virtual character, such as "smoking scenes are prohibited."
[0024] Returning to the explanation of Figure 1, the terminal device 200 is a general term for smartphones and other devices used to perform roles such as directing content 300, including moving images. The terminal device 200 transmits various information, such as work information 500 and virtual actor persona information 600, to the information processing device 100. The terminal device 200 also receives various information, such as content 300, from the information processing device 100. Note that there may be multiple terminal devices 200 communicating with the information processing device 100.
[0025] As described above, the information processing device 100 generates content 300 by inputting input information 400, which is generated based on the work information 500 and the virtual actor persona information 600, into a machine learning model.
[0026] However, in technologies that generate content 300 using machine learning models, there is a risk that content 300 that does not fit the virtual actor's character may be generated. For example, if an instruction to generate content 300 that includes a smoking scene 510 is input to the machine learning model, even if the virtual actor has a personality that does not smoke, there is a risk that content 300 in which the character smokes will be generated.
[0027] The information processing device 100 acquires work information 500 and virtual actor persona information 600 in order to solve the problem of appropriately utilizing virtual actors. Subsequently, the information processing device 100 generates input information 400 for a machine learning model that requests content 300 suitable for the virtual actor, based on the acquired work information 500 and virtual actor persona information 600. Subsequently, the information processing device 100 generates content 300 including the character played by the virtual actor by inputting the generated input information 400 into the machine learning model.
[0028] As a result, even if the work information 500 includes a smoking scene 510, the information processing device 100 generates content 300 that is suitable for the virtual actor's settings based on the virtual actor persona information 600. Consequently, the information processing device 100 can generate content 300 in which the character does not smoke. Therefore, the information processing device 100 can make appropriate use of the virtual actor.
[0029] The following describes an example of the above-described process in which the information processing device 100 performs the processes shown in steps S1 to S7 of Figure 1. That is, the information processing method according to this embodiment, which consists of the processes shown in steps S1 to S7, is executed by a computer such as the information processing device 100.
[0030] The terminal device 200 transmits the work information 500 and the virtual actor persona information 600 to the information processing device 100 (step S1). The information processing device 100 receives the work information 500 and the virtual actor persona information 600 from the terminal device 200 and acquires the work information 500 and the virtual actor persona information 600. For example, as shown in Figure 3, the information processing device 100 acquires the virtual actor persona information 600 which includes the prohibited information 610.
[0031] Next, the information processing device 100 generates input information 400X for a machine learning model that requests content 300 suitable for a virtual actor, based on the acquired work information 500 and virtual actor persona information 600 (step S2).
[0032] Content 300 suitable for a virtual actor is, for example, content 300 that satisfies the characteristics of the character indicated by the virtual actor persona information 600. Specifically, content 300 suitable for a virtual actor is content 300 that satisfies the profile of the virtual actor, such as appearance, as indicated by the virtual actor persona information 600, while not containing any prohibited words or actions corresponding to the prohibited information 610.
[0033] An example of generating input information 400X will be explained using Figure 4. Figure 4 is Figure (1) showing an example of input information. The information processing device 100 generates input information 400X that includes work information 500 and virtual actor persona information 600, etc., by adding work information 500 and virtual actor persona information 600, etc., to the template that forms the basis of the input information 400X.
[0034] For example, the information processing device 100 inserts the work information 500 under a prompt 410 that requests content 300 based on the work information 500. Similarly, the information processing device 100 inserts the virtual actor persona information 600 under a prompt 420 that requests content 300 based on the virtual actor persona information 600.
[0035] Furthermore, the information processing device 100 generates input information 400X based on the virtual actor persona information 600 which includes the prohibited information 610. For example, if the smoking scene 510 corresponding to the prohibited behavior indicated by the prohibited information 610 is included in the work information 500, the information processing device 100 adds a prompt 430 to the input information 400X which includes a command to prevent the generation of a video related to the prohibited behavior.
[0036] Returning to the explanation of Figure 1, the information processing device 100 can interact with the user via the terminal device 200, in addition to the generated input information 400X (step S3).
[0037] Since such dialogue takes place between a dialogue model containing the persona information of the virtual actor and the user, it can also be described as a dialogue between the virtual actor and the user (who can also be called the content director). If it is anticipated that the content 300 generated based on the input information 400X will not be in line with the virtual actor's settings, such dialogue may include a response 700 to request a response from the user regarding a proposed revision.
[0038] An example of such a dialogue will be explained using Figure 5. Figure 5 is a diagram showing an example of a dialogue history. For example, the information processing device 100 receives input from the user, "Create Scene 1. Character A should have a cigarette in his mouth." The information processing device 100 then refers to information such as the virtual actor persona information 600 and outputs a response such as, "Smoking scenes are prohibited for actor X. Please consider other staging options." In response to this, the user inputs instruction information 800, "Then, create Scene 1 with Character A chewing gum."
[0039] Returning to the explanation of Figure 1, as described above, when the terminal device 200 receives instruction information 800 from the user (step S4), it transmits the instruction information 800 to the information processing device 100 via dialogue (step S5). The information processing device 100 acquires the instruction information 800 by receiving it from the terminal device 200.
[0040] Subsequently, the information processing apparatus 100 generates input information 400 for generating content 300 in which prohibited actions are corrected, based on the acquired instruction information 800, with the input information 400X having dialogue information added thereto as the input information 400 (step S6). Hereinafter, an example of the generation of the input information 400 will be described using FIG. 6. FIG. 6 is a diagram (2) showing an example of the input information.
[0041] In FIG. 6, the information processing apparatus 100 generates the input information 400 by further adding the instruction information 800 to the input information 400X. For example, the information processing apparatus 100 generates the input information 400 by inserting the response 700X and the instruction information 800X after the prompt 440 indicating the dialogue history.
[0042] The response 700X is a response 700 such as a message requesting the user for instruction information 800 regarding an amendment for correcting a smoking scene 510 or the like corresponding to the prohibited action, with the text "Output (Actor X):" added to the beginning. The instruction information 800X is the instruction information 800 with the text "Input (User):" added to the beginning.
[0043] Returning to the description of FIG. 1, the information processing apparatus 100 generates content 300 in which prohibited actions are corrected by inputting the generated input information 400 into the machine learning model. For example, the information processing apparatus 100 generates a moving image in which the smoking scene 510 corresponding to the prohibited action is corrected to a scene of chewing gum as the content 300.
[0044] Subsequently, the information processing apparatus 100 transmits the generated content 300 to the terminal device 200 (step S7). The terminal device 200 receives the content 300 from the information processing apparatus 100. Subsequently, the terminal device 200 displays the received content 300.
[0045] As described above, the information processing apparatus 100 generates content 300 including a character played by the virtual actor based on the acquired work information 500 and the virtual actor persona information 600.
[0046] As a result, for example, even if the work information 500 includes a smoking scene 510, the information processing apparatus 100 generates the content 300 that conforms to the setting in which the virtual actor does not smoke based on the virtual actor persona information 600. As a result, the information processing apparatus 100 can generate the content 300 in which the character does not smoke. Therefore, the information processing apparatus 100 can appropriately utilize the virtual actor.
[0047] (1-2. Configuration of the information processing system according to the first embodiment) Next, an example of the configuration of the information processing system 10 according to the first embodiment will be described using FIG. 7. FIG. 7 is a block diagram showing an example of the configuration of the information processing system according to the first embodiment. The information processing system 10 includes an information processing apparatus 100 and a terminal device 200.
[0048] (Configuration of the information processing apparatus) The information processing apparatus 100 includes a communication unit 110, a storage unit 120, a reception unit 130, and a control unit 140.
[0049] (Communication unit) The communication unit 110 is realized by, for example, a NIC (Network Interface Card) or a network interface controller (Network Interface Controller). The communication unit 110 is connected to the network N by wire or wirelessly, and transmits and receives information to and from the terminal device 200 via the network N. The network N is realized by, for example, a wireless communication standard or method such as Bluetooth (registered trademark), the Internet, Wi-Fi (registered trademark), UWB (Ultra-Wide Band), LPWA (Low Power Wide Area).
[0050] For example, the communication unit 110 receives the work information 500, the virtual actor persona information 600, the instruction information 800, the casting information 900 regarding the character, and the dialogue history 950 between the virtual actor and the user from the terminal device 200. Further, the communication unit 110 transmits the response 700 and the content 300 to the terminal device 200.
[0051] Casting information 900 is text that identifies which virtual human will play which character in content 300, for example, "Character A will be played by actor X." Note that casting information 900 may also be included within the work information 500.
[0052] Dialogue history 950 is a record of the conversations that took place between the virtual actor and the user.
[0053] (Memory Unit) The memory unit 120 is implemented by semiconductor memory elements such as RAM (Random Access Memory) and flash memory, or by storage devices such as hard disks, SSDs (Solid State Drives), and optical discs. For example, the memory unit 120 stores work information 500, virtual actor persona information 600, instruction information 800, casting information 900, and dialogue history 950.
[0054] (Reception Unit) The reception unit 130 is a user interface that receives various types of information from users of the information processing device 100. For example, the reception unit 130 is a user interface that can only be accessed by specific users of the information processing device 100, such as a virtual actor's talent agency.
[0055] (Control Unit) The control unit 140 is implemented by a CPU (Central Processing Unit), an MPU (Micro Processing Unit), a GPU (Graphics Processing Unit), etc., which executes a program stored inside the information processing device 100 (for example, the information processing program according to this disclosure) using RAM or the like as a working area. The control unit 140 is also a controller and may be implemented by an integrated circuit such as an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or an MCU (Micro Controller Unit).
[0056] The control unit 140 includes an acquisition unit 141, a determination unit 142, an input information generation unit 143, and a content generation unit 144.
[0057] (Acquisition Unit) The acquisition unit 141 acquires various types of information. For example, the acquisition unit 141 acquires work information 500 and virtual actor persona information 600 received from the terminal device 200 by the communication unit 110. Specifically, the acquisition unit 141 acquires prohibited information 610 as virtual actor persona information 600. The acquisition unit 141 also acquires information about virtual actors who play characters appearing in content 300 as virtual actor persona information 600.
[0058] Furthermore, the acquisition unit 141 acquires at least one of the instruction information 800, casting information 900, and dialogue history 950 received from the terminal device 200 by the communication unit 110.
[0059] Furthermore, the acquisition unit 141 acquires instruction information 800 when access from a specific user is accepted by the reception unit 130, which is accessible only to specific users of the information processing device 100. Specifically, when access from a virtual actor's talent agency is accepted by the reception unit 130 as a specific user, the acquisition unit 141 acquires instruction information 800 from the video director.
[0060] (Determination Unit) The determination unit 142 determines whether the dialogue between the virtual actor and the user has ended. For example, if the reception unit 210, described later, such as the GUI (Graphical User Interface), receives an explicit operation to end the dialogue from the user of the terminal device 200, the determination unit 142 determines that the dialogue has ended. Also, if the reception unit 210 does not receive any operation from the user of the terminal device 200 for a predetermined period of time or longer, the determination unit 142 determines that the dialogue has ended.
[0061] Furthermore, the determination unit 142 determines whether or not the dialogue between the virtual actor and the user has ended, based on at least one of the instruction information 800 and the dialogue history 950 acquired by the acquisition unit 141.
[0062] For example, the determination unit 142 functions as a classifier to determine whether a dialogue has ended if the instruction information 800 or dialogue history 950 contains an utterance by the user of the terminal device 200, such as "I've finished the operation for now." In this case, the determination unit 142 determines that the dialogue has ended if, for example, keywords such as "operation," "dialogue," or "end" are included in the instruction information 800 or dialogue history 950.
[0063] (Input Information Generation Unit) The input information generation unit 143 generates input information 400 for a machine learning model that requests content 300 suitable for a virtual actor, based on the acquired work information 500 and virtual actor persona information 600. Below, an example of the generation of input information 400 by the input information generation unit 143 will be explained using Figure 8. Figure 8 is Figure (1) for illustrating an example of input information generation.
[0064] The input information generation unit 143 generates input information 400 for the machine learning model based on the work information 500, virtual actor persona information 600, instruction information 800, casting information 900, and dialogue history 950. For example, if information about a virtual actor is acquired as virtual actor persona information 600, the input information generation unit 143 generates input information 400 based on the acquired information.
[0065] Furthermore, the input information generation unit 143 generates input information 400 for the large-scale language model as input information 400 for the machine learning model. Specifically, the input information generation unit 143 generates an input token sequence, which is a sequence of tokens to be input into MMLLM, as input information 400 for the large-scale language model.
[0066] Furthermore, the input information generation unit 143 generates input information 400 based on the prohibited information 610 of the acquired virtual actor persona information 600. For example, if prohibited words or actions are included in the acquired work information 500, the input information generation unit 143 generates input information 400 that includes prompts to restrict the generation of videos containing prohibited words or actions, and dialogue history related to alternative performances.
[0067] As an example, the input information generation unit 143 generates input information 400X in which the smoking scene 510 corresponding to the prohibited behavior indicated by the prohibited behavior information 610 is included in the work information 500. Subsequently, the content generation unit 144, described later, inputs the input information 400X into a machine learning model and outputs a response 700 regarding the prohibited behavior, such as "Smoking scenes by actor X are prohibited. Please consider alternative staging," to the user as a dialogue.
[0068] Furthermore, response 700, which states, "Smoking scenes involving actor X are prohibited. Please consider alternative staging," also serves as a request for instructions from the user.
[0069] Next, the input information generation unit 143 receives instruction information 800 as a response from the user to response 700. Subsequently, the input information generation unit 143 generates input information 400 which includes response 700X with "Output (Actor X):" added to the beginning of response 700, and instruction information 800X with "Input (User):" added to the beginning of instruction information 800. In other words, the input information generation unit 143 generates input information 400 which includes a dialogue history consisting of response 700X and instruction information 800X added to input information 400X.
[0070] As another example, let's consider a case where the instruction to smoke, corresponding to the prohibited behavior indicated by the prohibited information 610, is included in the instruction information from the user. In this case as well, the input information generation unit 143 generates input information 400X, etc., to request a response from the user regarding the prohibited behavior.
[0071] Below, an example of the generation of input information 400Y by the input information generation unit 143 will be explained using Figure 9. The input information 400Y is generated based on the work information 500, the virtual actor persona information 600, the instruction information 800, as well as the casting information 900 and the dialogue history 950.
[0072] Figure 9 is a diagram (2) illustrating an example of input information generation. In Figure 9, the input information generation unit 143 comprises a work information prompt generation unit 1431, a character information prompt generation unit 1432, a dialogue input generation unit 1433, and a prompt coupling unit 1434.
[0073] (Work Information Prompt Generation Unit) The work information prompt generation unit 1431 generates work information prompts, which are prompts related to the work, based on the work information 500 and a portion of the input information 400Y. An example of the generation of work information prompts by the work information prompt generation unit 1431 will be explained using Figure 10. Figure 10 is a diagram illustrating an example of the generation of work information prompts.
[0074] The work information prompt generation unit 1431 generates the work information prompt 401 by adding the work information 500 to the template 401X which serves as the basis for the work information prompt 401. For example, the work information prompt generation unit 1431 generates the work information prompt 401 by inserting the work information 500 into the blank space 411 below the prompt 410 in the template 401X.
[0075] (Character Information Prompt Generation Unit) Returning to the explanation of Figure 9, the character information prompt generation unit 1432 generates character information prompts, which are prompts related to the virtual character, from a portion of the input information 400Y, based on the casting information 900 and the virtual actor persona information 600.
[0076] Below, an example of character information prompt generation by the character information prompt generation unit 1432 will be explained using Figure 11. Figure 11 is a diagram illustrating an example of character information prompt generation.
[0077] The character information prompt generation unit 1432 generates the character information prompt 402 by adding casting information 900 and virtual actor persona information 600 to the template 402X which serves as the basis for the character information prompt 402.
[0078] For example, the character information prompt generation unit 1432 inserts the casting information 900 into the blank space 441 below the prompt 440 that requests content 300 based on the virtual actor persona information 600 in template 402X. The character information prompt generation unit 1432 also inserts the virtual actor persona information 600 into the blank space 421 below the prompt 420X in template 402X. As a result, the character information prompt generation unit 1432 generates the character information prompt 402.
[0079] While prompt 420X is similar to prompt 420 in that it requests content 300 based on virtual actor persona information 600, it differs from prompt 420 in that it also requests content 300 based on casting information 900.
[0080] (Dialogue Input Generation Unit) Returning to the explanation of Figure 9, the dialogue input generation unit 1433 generates dialogue input, which is a prompt related to the dialogue, from a portion of the input information 400Y, based on the dialogue history 950 and the instruction information 800. Below, an example of dialogue input generation by the dialogue input generation unit 1433 will be explained using Figure 12. Figure 12 is a diagram illustrating an example of dialogue input generation.
[0081] The dialogue input generation unit 1433 generates the dialogue input 403 by adding the dialogue history 950 and instruction information 800 to the template 403X which serves as the basis for the dialogue input 403.
[0082] For example, the dialogue input generation unit 1433 inserts the dialogue history 950 into the blank space 451 below the prompt 450 that requests content 300 based on the dialogue history 950 in the template 403X. Specifically, the dialogue input generation unit 1433 inserts the dialogue history 950 into the blank space 451, which includes instruction information 800Y with the text "Input (User):" added to the beginning of past instruction information 800, and response 700X with the text "Output (Actor X):" added to the beginning of past response 700.
[0083] Furthermore, the dialogue input generation unit 1433 inserts instruction information 800X, which has the text "Input (User):" added to the beginning of the most recent instruction information 800, into the blank space 461 below the prompt 460 that requests content 300 based on instruction information 800 in template 403X. As a result, the dialogue input generation unit 1433 generates the dialogue input 403.
[0084] (Prompt coupling unit) Returning to the explanation of Figure 9, the prompt coupling unit 1434 generates input information 400Y by combining the generated work information prompt 401, the character information prompt 402, and the dialogue input 403. Below, an example of the generation of input information 400Y by the prompt coupling unit 1434 will be explained using Figure 13. Figure 13 is a diagram (3) illustrating an example of input information generation.
[0085] The prompt merging unit 1434 generates input information 400Y by merging these prompts so that, from top to bottom, they become the work information prompt 401, the character information prompt 402, and the dialogue input 403.
[0086] Furthermore, in Figure 13, the smoking scene 510 corresponding to the prohibited behavior indicated by the prohibited information 610 is included in the work information 500. In this case, the input information generation unit 143 generates modified input information as input information 400Y, requesting content 300 in which the smoking scene 510 corresponding to the prohibited behavior has been modified.
[0087] For example, the input information 400Y may include the response 700X generated by the content generation unit 144 and the user's instruction information 800Y in response to it. That is, the input information generation unit 143 can also generate a suggested revision input information requesting a suggested revision of the content 300 as the input information 400Y, based on the acquired dialogue history 950.
[0088] (Content Generation Unit) Returning to the explanation of Figure 7, the content generation unit 144 generates content 300 that includes the character played by the virtual actor, based on the acquired work information 500 and virtual actor persona information 600. For example, if the virtual actor does not smoke according to the settings in the virtual actor persona information 600, even if the work information 500 includes a smoking scene 510, the content generation unit 144 generates content 300 that conforms to the setting that the virtual actor does not smoke.
[0089] Furthermore, the content generation unit 144 generates content 300 by inputting the generated input information 400 into a machine learning model. For example, the content generation unit 144 generates moving images such as anime or dramas featuring virtual actors as content 300 by inputting the generated input information 400 into MMLLM.
[0090] Furthermore, if the content generation unit 144 receives input information containing prohibited speech or actions, it generates a response 700 to request correction instructions from the user. For example, the content generation unit 144 generates a response 700 regarding a suggested correction (another scene), such as, "Smoking scenes by actor X are prohibited. Please consider alternative staging."
[0091] Furthermore, if the content generation unit 144 receives correction input information as input information 400Y, which requests content 300 in which prohibited behavior has been corrected, it inputs the correction input information into a machine learning model to generate content 300 in which the prohibited behavior has been corrected. For example, the content generation unit 144 generates video content 300 such as an anime or drama in which a smoking scene 510 corresponding to the prohibited behavior indicated by the prohibition information 610 has been corrected to a scene of chewing gum.
[0092] Furthermore, if alternative input information requesting alternative content is generated as input information 400Y based on the dialogue history 950, the content generation unit 144 generates an alternative content as content 300 by inputting the alternative input information into the machine learning model.
[0093] For example, if the content generation unit 144 makes a suggestion such as, "Smoking scenes with actor X are prohibited. Please consider alternative staging," and the user requests, "Then try a different action," the content generation unit 144 will generate content 300 based on the alternative suggestion.
[0094] For example, the content generation unit 144 generates an alternative content in which the smoking scene 510 is modified to a scene of chewing gum by inputting alternative input information into a machine learning model. The content generation unit 144 may also automatically generate the alternative without explicit instructions from the user. In this case, the content generation unit 144 responds to the user with a message such as, "Smoking scenes with actor X are prohibited. We will switch to another scene," and then generates content 300 including the modified scene.
[0095] (Configuration of the terminal device) The terminal device 200 includes a reception unit 210, a communication unit 220, a display unit 230, and a control unit 240.
[0096] (Reception Unit) The reception unit 210 is a user interface, etc., that receives various information from the user of the terminal device 200. For example, the reception unit 210 is a touch panel, etc.
[0097] (Communication Unit) The communication unit 220 is implemented by, for example, a NIC or a network interface controller. The communication unit 220 is connected to the network N by wire or wireless connection and transmits and receives information with the information processing device 100 via the network N. The network N is implemented by, for example, a wireless communication standard or method such as Bluetooth®, the Internet, Wi-Fi®, UWB, or LPWA.
[0098] For example, the communication unit 220 transmits to the information processing device 100 work information 500, virtual actor persona information 600, instruction information 800, casting information 900, and the dialogue history 950 between the virtual actor and the user. The communication unit 220 also receives responses 700 and content 300 from the information processing device 100.
[0099] (Display Unit) The display unit 230 is a desktop or the like that displays various information. For example, the display unit 230 is a touch panel or the like that displays response 700 and content 300.
[0100] (Control Unit) The control unit 240 is implemented by, for example, a CPU, MPU, or GPU, which executes a program stored inside the terminal device 200 using RAM or the like as a working area. The control unit 240 is also a controller and may be implemented by, for example, an integrated circuit such as an ASIC, FPGA, or MCU.
[0101] (1-3. Information Processing Flow According to the First Embodiment) An example of the information processing flow by the information processing device 100 according to the first embodiment will be explained using Figure 14. Figure 14 is a flowchart showing an example of the information processing flow according to the first embodiment.
[0102] First, the determination unit 142 determines whether or not user input has been received (step S11).
[0103] The following describes an example of how user input is received. In this example, the acquisition unit 141 receives user input by acquiring the work information 500, virtual actor persona information 600, instruction information 800, casting information 900, and dialogue history 950 received from the user by the receiving unit 210 of the terminal device 200. Specifically, the acquisition unit 141 indirectly receives user input by acquiring this information received by the receiving unit 210 of the terminal device 200 from the terminal device 200.
[0104] In this case, the determination unit 142 determines that it has received user input (step S11; Yes), and then determines whether the dialogue between the virtual actor and the user has ended (step S12). On the other hand, if the determination unit 142 determines that there has been no user input for a certain period of time, such as when no operation from the user has been performed to the reception unit 210 for a predetermined period of time (step S11; No), it terminates the information processing due to a timeout.
[0105] Furthermore, in step S12, if the acquisition unit 141 acquires an explicit operation to end the conversation from the user of the terminal device 200 through the reception unit 210, the determination unit 142 determines that the conversation has ended. If the determination unit 142 determines that the conversation has ended (step S12; Yes), it terminates the information processing.
[0106] If the determination unit 142 determines that the dialogue has not ended (step S12; No), the input information generation unit 143 generates input information 400 for a machine learning model that requests content 300 suitable for the virtual actor (step S13). For example, the input information generation unit 143 generates input information 400X, which is a sequence of input tokens requesting a response 700, by combining work information 500, virtual actor persona information 600, instruction information 800, casting information 900, and dialogue history 950.
[0107] Next, the content generation unit 144 generates content 300 by inputting the generated input information 400 into a machine learning model (step S14). For example, if the content generation unit 144 generates input information 400X, which is a sequence of input tokens requesting a response 700, as input information 400, it generates the response 700 by inputting the input information 400X into a large-scale language model.
[0108] Next, the input information generation unit 143 stores the instruction information 800 and the generated response 700 in the storage unit 120 as a dialogue history 950 (step S15). For example, the input information generation unit 143 stores a dialogue history 950 in which the instruction information 800 and the response 700 are combined. After the completion of step S15, the information processing device 100 returns to the processing of step S11.
[0109] (2. Modified Versions) (2-1. First Modified Version) In the above example, the case in which various types of information acquired by the acquisition unit 141, such as work information 500 and virtual actor persona information 600, are in text form was explained. However, such various types of information may also be images, audio, etc.
[0110] For example, the input information generation unit 143 generates at least one of text, images, and audio to be input to a large-scale language model such as MMLLM as input information 400Y to a machine learning model. In this case, the input information generation unit 143 generates the input information 400Y by inserting images such as moving images or links to images, or audio or links to audio, into the blanks 411, 421, 441, 451, and 461 of the templates 401X to 403X that serve as the basis for the input information 400Y.
[0111] Below, an example of artwork information 500 will be explained using Figure 15. Figure 15 is Figure (2) showing an example of artwork information. Artwork information 500X in Figure 15 is information that corresponds to the following text in artwork information 500, but is represented as an image instead of text: • It is getting dark in the evening. • It is raining. • Character A is waiting for a bus at a bus stop on an unpaved road cut through a forest.
[0112] Next, an example of virtual actor persona information 600 will be explained using Figure 16. Figure 16 is Figure (2) showing an example of virtual actor persona information. Virtual actor persona information 600X and virtual actor persona information 600Y in Figure 16 are information where the following text from virtual actor persona information 600 is represented as an image instead of text. Also, virtual actor persona information 600Z is information where the following text from virtual actor persona information 600 is represented as audio instead of text. Persona information of actor X: Height 175cm, Weight 60kg, Black hair, Brown eyes
[0113] (2-2. Second Modification) In the above example, the case in which the input information generation unit 143 generates the input information 400Y by filling in the blanks 411, 421, 441, 451 and 461 of the templates 401X to 403X that form the basis of the input information 400Y using a fill-in-the-blank method. However, the input information generation unit 143 can also generate the input information 400Y by using a model method in which it inputs the acquired work information 500 and virtual actor persona information 600 and other various information into a machine learning model different from the machine learning model used by the content generation unit 144.
[0114] Even in this model method, the input information generation unit 143 generates the input information 400Y by filling in various pieces of information into the blanks 411, 421, 441, 451, and 461 of the templates 401X to 403X that form the basis of the input information 400Y, which is the same as the fill-in-the-blank method. However, in the model method, the input information generation unit 143 generates the various pieces of information to be filled into the blanks 411, 421, 441, 451, and 461 by using individual machine learning models to fill in each of these blanks. These individual machine learning models are, for example, large-scale language models such as MMLLM.
[0115] As an example, the input information generation unit 143 generates various information to be inserted into blanks 411, 421, 441, 451, and 461 by converting various information such as work information 500 and virtual actor persona information 600 into embedding vectors using individual machine learning models.
[0116] For example, the artwork information prompt generation unit 1431 of the input information generation unit 143 encodes the artwork information 500 by inputting it into a machine learning model for generating the artwork information prompt 401. As a result, the artwork information prompt generation unit 1431 converts the artwork information 500 into an embedding vector to be placed in the blank space 411 in Figure 10.
[0117] Furthermore, the character information prompt generation unit 1432 of the input information generation unit 143 encodes the casting information 900 and the virtual actor persona information 600 by inputting them into a machine learning model for generating the character information prompt 402. As a result, the character information prompt generation unit 1432 converts the casting information 900 and the virtual actor persona information 600 into embedding vectors to be placed in the blanks 441 and 421 in Figure 11.
[0118] Furthermore, the dialogue input generation unit 1433 of the input information generation unit 143 encodes the dialogue history 950 and instruction information 800 by inputting them into a machine learning model for generating dialogue input 403. As a result, the dialogue input generation unit 1433 converts the dialogue history 950 and instruction information 800 into embedding vectors to be placed in the blanks 451 and 461 in Figure 12.
[0119] Furthermore, the prompt coupling unit 1434 of the input information generation unit 143 converts the three aforementioned embedding vectors into a single embedding vector by multiplication, weighted averaging, or the like. Subsequently, the input information generation unit 143 inputs the converted single embedding vector and the template for input information 400Y into a machine learning model for generating input information 400Y, thereby generating input information 400Y in which various information is filled into blanks 411, 421, 441, 451, and 461.
[0120] The input information generation unit 143 may also send the embedded vectors directly to the content generation unit 144. Alternatively, the input information generation unit 143 may directly create an input token sequence from the embedded vectors to the content generation unit 144.
[0121] As another example, the input information generation unit 143 generates various information to be inserted into blanks 411, 421, 441, 451, and 461 by inputting the user's utterance from the terminal device 200 into a machine learning model. Below, using Figure 17, an example of generating artwork information 500Y to be inserted into blank 411 in Figure 10, using the model method of the artwork information prompt generation unit 1431 of the input information generation unit 143, will be explained. Figure 17 is a diagram illustrating an example of artwork information generation.
[0122] For example, if the user of terminal device 200 speaks the following during a video call, the work information prompt generation unit 1431 generates work information 500Y by inputting the content of the user's speech into a machine learning model for generating work information prompts 401. • "The scene we're about to shoot, I've been thinking about it since yesterday, and I think it would be better if it were raining, and the time of day would be around evening when it starts to get dark, and I want to shoot a close-up of the main character with this kind of expression (the user makes a face as they say this), and then sighs like 'Phew'."
[0123] Specifically, the work information prompt generation unit 1431 extracts text work information 500Y1 from the user's speech content on the terminal device 200. The work information prompt generation unit 1431 also extracts an image of the user's frown from the video call footage as work information 500Y2. Furthermore, the work information prompt generation unit 1431 extracts the user's sigh, "phew," from the audio of the video call as work information 500Y3. As a result, the work information prompt generation unit 1431 generates work information 500Y.
[0124] (2-3. Third variation) In the above example, we described a case where the prompt 430 requesting a response 700 regarding prohibited conduct corresponds to prohibited conduct of the virtual actor, and is a prompt requesting a response 700 notifying the prohibited conduct. However, the prompt 430 may also be a prompt requesting a response regarding a proposed amendment as the response 700.
[0125] In this case, the character information prompt generation unit 1432 of the input information generation unit 143 generates a character information prompt 402 in which text requesting a response regarding a proposed revision is added to prompt 430. Specifically, the character information prompt generation unit 1432 generates a character information prompt 402 that includes a prompt with the text, "If the generated video does not match the actor's image, please suggest a change," added to the end of prompt 430.
[0126] Furthermore, when the content generation unit 144 determines whether the content 300 matches the image of the virtual actor, it also determines whether the virtual actor persona information 600 contains information about the aspects of the virtual actor's image that are important to the unit. For example, if the virtual actor persona information 600 contains the statement, "Black hair. Brown eyes. These are the important aspects of the actor's image," the content generation unit 144 will determine whether or not to suggest changes to the content 300 depending on whether the virtual actor fits that image.
[0127] (2-4. Fourth Modification) The above example described a case where the number of virtual characters, each given a virtual personality such as a virtual actor, is one. However, the information processing device 100 can also be applied when there are multiple virtual actors.
[0128] In this case, the acquisition unit 141 of the information processing device 100 acquires information about multiple virtual actors as virtual actor persona information 600. The input information generation unit 143 of the information processing device 100 generates information requesting content 300 suitable for multiple virtual actors as input information 400, based on the acquired information about multiple virtual actors.
[0129] (2-5. Fifth Modification) In the above example, an example was described in which the content generation unit 144 of the information processing device 100 generates moving images such as anime or dramas featuring virtual actors as content 300. However, for example, the virtual actors may be given a virtual or real avatar (robot) body.
[0130] In this case, the content generation unit 144 generates, for example, content 300 in which a virtual actor appears in real space as a robot. Specifically, the content generation unit 144 generates, as content 300, content for a virtual actor, who has been given a body as a robot, to interact with the user of the terminal device 200 in the real world, such as through dialogue.
[0131] (3. Second Embodiment) The information processing device 100 can also store past dialogue history and generate alternative input information that requests alternative content as input information 400 based on past dialogue history. Past dialogue history includes dialogue history with a user other than the user of the terminal device 200 currently interacting with the virtual actor, and dialogue history with the user currently interacting with the virtual actor after a predetermined period of time has passed.
[0132] (3-1. Configuration of the Information Processing System According to the Second Embodiment) Hereinafter, an example of the configuration of the information processing device 100A included in the information processing system 10A according to the second embodiment will be described with reference to Figure 18. Figure 18 is a block diagram showing an example of the configuration of the information processing system according to the second embodiment. The information processing device 100A includes a storage unit 120A and a control unit 140A instead of the storage unit 120 and control unit 140 shown in the information processing device 100.
[0133] (Memory Unit) The memory unit 120A further stores past dialogue history 950X. The memory unit 120A also stores at least one of the work information 500 and the instruction information 800 in association with at least one of the dialogue history 950 and past dialogue history 950X.
[0134] (Control Unit) The control unit 140A includes a determination unit 142A and an input information generation unit 143A, instead of the determination unit 142 and input information generation unit 143 shown in the control unit 140.
[0135] (Determination Unit) The determination unit 142A determines whether or not there is a past dialogue history 950X with the user of the terminal device 200. For example, if the determination unit 142A searches for the past dialogue history 950X with the user of the terminal device 200 and determines that there is a past dialogue history 950X, it inputs the past dialogue history 950X to the input information generation unit 143A.
[0136] Furthermore, if the determination unit 142A determines that the dialogue between the virtual actor and the user of the terminal device 200 has ended, it stores the dialogue as past dialogue history 950X in the storage unit 120A.
[0137] Furthermore, the determination unit 142A determines the consistency of the user characteristics of the terminal device 200, which is the conversation partner with the virtual actor, based on the conversation history 950 and past conversation history 950X. These user characteristics are, for example, a persona that represents the profile of the terminal device 200, such as the personality and background (place of origin, favorite food) of the user of the terminal device 200.
[0138] The determination unit 142A determines consistency by, for example, whether the dialogue history 950 being compared and the past dialogue history 950X are contradictory, whether they are independent pieces of information, or whether they are identical.
[0139] As an example, the determination unit 142A determines whether the keywords indicating the personality and background of the user of the terminal device 200, as shown in the dialogue history 950 and the past dialogue history 950X, match. Specifically, the determination unit 142A determines whether the characteristics of the user of the terminal device 200, such as the absence of periods at the end of utterances or the absence of honorific language, match, as shown in the dialogue history 950 and the past dialogue history 950X.
[0140] The consistency determination by the determination unit 142A may be made by a machine learning model related to language processing that performs sentence semantic determination, or by a rule-based determination. An example of a mechanism for determining contradictions, independent information, and information consistency between the dialogue history 950 and past dialogue history 950X is contradiction detection using natural language inference.
[0141] Natural language inference is a technique for inferring whether two texts are contradictory, implicative, or unrelated. For example, natural language inference regarding personas is illustrated in "Generating Persona Consistent Dialogues by Exploiting Natural Language Inference (AAAI-20)". In addition to the above-mentioned method, the determination unit 142A can also determine consistency based on various known techniques.
[0142] (Input Information Generation Unit) The input information generation unit 143A generates input information 400 based on the past dialogue history 950X.
[0143] For example, if the determination unit 142A determines the consistency of the user characteristics of the terminal device 200 based on the dialogue history 950 and past dialogue history 950X, the input information generation unit 143A generates input information 400 based on the determined consistency. Specifically, if the input information generation unit 143A determines that the user characteristics are consistent, such as the user who interacted with the virtual actor being the same person in the dialogue history 950 and past dialogue history 950X, it generates input information 400.
[0144] Below, an example of the generation of input information 400 by the input information generation unit 143A will be explained using Figure 19. Figure 19 is a diagram (4) illustrating an example of input information generation.
[0145] The input information generation unit 143A generates alternative input information requesting alternative content, as input information 400, based on the past dialogue history 950X. For example, the input information generation unit 143A generates alternative input information by further inserting the past dialogue history 950X stored in the storage unit 120A into the template of the input information 400.
[0146] Furthermore, the input information generation unit 143A generates alternative input information based on the work information 500. For example, the input information generation unit 143A generates alternative input information based on the work information 500 linked to past dialogue history 950X. Specifically, the input information generation unit 143A generates alternative input information that proposes homages to past content 300 or other works that appeared in past dialogue history 950X as alternatives.
[0147] Below, an example of the generation of input information 400Z by the input information generation unit 143A will be explained using Figure 20. Input information 400Z differs from input information 400Y in that it is generated based on past dialogue history 950X in addition to the work information 500, virtual actor persona information 600, instruction information 800, casting information 900, and dialogue history 950.
[0148] Figure 20 is a diagram (5) illustrating an example of input information generation. In Figure 20, the input information generation unit 143A further comprises a dialogue history prompt generation unit 1435 in addition to the input information generation unit 143. Furthermore, the input information generation unit 143A includes a prompt coupling unit 1434A instead of the prompt coupling unit 1434.
[0149] (Dialogue History Prompt Generation Unit) The dialogue history prompt generation unit 1435 generates dialogue history prompts, which are a part of the prompts in the input information 400Z, based on past dialogue history 950X. Below, an example of dialogue history prompt generation by the dialogue history prompt generation unit 1435 will be explained using Figure 21. Figure 21 is a diagram illustrating an example of dialogue history prompt generation.
[0150] The dialogue history prompt generation unit 1435 generates the dialogue history prompt 404 by adding past dialogue history 950X to the template 404X which serves as the basis for the dialogue history prompt 404.
[0151] For example, the dialogue history prompt generation unit 1435 generates a dialogue history prompt 404 by inserting the past dialogue history 950X into the blank space 471 below the prompt 470 which requests content 300 based on the past dialogue history 950X of template 404X. Specifically, the dialogue history prompt generation unit 1435 generates the dialogue history prompt 404 by inserting the past dialogue history 950X of a video director, who is a different user from the user of the terminal device 200 currently interacting with the virtual actor, into the blank space 471.
[0152] For example, if the user of terminal device 200 is supervisor L and another user is supervisor M, the dialogue history prompt generation unit 1435 inserts past dialogue history 950X containing names such as "Input (Supervisor L)" and "Input (Supervisor M)" into the blank 471 in the "Input (User)" section. In this way, when the dialogue history prompt generation unit 1435 uses past dialogue history 950X of different users, it inserts past dialogue history 950X containing different usernames into the blank 471. The terminal device 200 also registers usernames such as supervisor L and supervisor M.
[0153] Furthermore, if the work information 500 is linked to past dialogue history 950X, the dialogue history prompt generation unit 1435 can also add text to the end of the prompt 470, etc., to generate alternative suggestion input information based on the work information 500. In this case, the dialogue history prompt generation unit 1435 adds text such as, "If you can also suggest an homage to past works, please suggest it if you think it is a very good suggestion."
[0154] (Prompt coupling unit) Returning to the explanation of Figure 20, the prompt coupling unit 1434A generates input information 400Z by further coupling the generated dialogue history prompt 404 with the work information prompt 401, the character information prompt 402, and the dialogue input 403. Below, an example of input information generation by the prompt coupling unit 1434A will be explained using Figure 22. Figure 22 is a diagram (6) illustrating an example of input information generation.
[0155] The prompt merging unit 1434A generates input information 400Z by merging the prompts so that, from top to bottom, they are: work information prompt 401, character information prompt 402, dialogue input 403, and dialogue history prompt 404. The order of these prompts is not particularly limited. For example, the prompt merging unit 1434A can also merge these prompts so that the dialogue history prompt 404 is at the top.
[0156] (3-2. Information Processing Flow According to the Second Embodiment) An example of the information processing flow by the information processing device 100A according to the second embodiment will be explained using Figure 23. Figure 23 is a flowchart of an example of the information processing flow according to the second embodiment. Steps S24, S25, S28, and S29 are the same as steps S11, S12, S14, and S15, so their explanation will be omitted.
[0157] First, the determination unit 142A searches for past dialogue history 950X between the virtual actor and the user of the terminal device 200 (step S21). For example, the determination unit 142A searches for past dialogue history 950X from the storage unit 120A.
[0158] If the determination unit 142A determines that past dialogue history 950X exists (step S22; Yes), it inputs the past dialogue history 950X to the input information generation unit 143A (step S23). For example, if the determination unit 142A determines that past dialogue history 950X exists in the storage unit 120A, it retrieves the past dialogue history 950X from the storage unit 120A and inputs it to the input information generation unit 143A. If the determination unit 142A determines that past dialogue history 950X does not exist (step S22; No), it proceeds to step S24.
[0159] If the determination unit 142A determines that the dialogue between the virtual actor and the user of the terminal device 200 has ended (step S25; Yes), it stores the dialogue with the user of the terminal device 200 as past dialogue history 950X in the storage unit 120A (step S26), and terminates the information processing.
[0160] If the determination unit 142A determines that the interaction between the virtual actor and the user of the terminal device 200 has not ended (step S25; No), the input information generation unit 143A generates input information 400 based on the past interaction history 950X (step S27). For example, the input information generation unit 143A generates input information 400X requesting a response 700, which is a combination of the work information 500, the virtual actor persona information 600, the instruction information 800, the casting information 900, the interaction history 950, and the past interaction history 950X.
[0161] (4. Third Embodiment) The information processing device 100A can also generate information requesting content 300 relating to dialogue between multiple virtual actors as input information 400. An example of the configuration of the information processing system 10B according to the third embodiment will be described below with reference to Figure 24.
[0162] Figure 24 is a block diagram showing an example of the configuration of an information processing system according to the third embodiment. In the information processing system 10B, a plurality of information processing devices 100B are interconnected via a network N. Furthermore, the information processing device 100B includes a storage unit 120B and a control unit 140B instead of a storage unit 120A and a control unit 140A compared to the information processing device 100A. Note that the information processing system 10B can also be implemented with a single information processing device 100B.
[0163] (Storage Unit) The storage unit 120B stores instruction information 800B and dialogue history 950B instead of the instruction information 800 and dialogue history 950 shown in the storage unit 120A.
[0164] The instruction information 800B is information that further includes at least one of the instructions and utterances of the virtual actor of the other information processing device 100B, in addition to the instruction information 800. For example, the instruction information 800B is information that further includes a response 700B regarding the utterance of the virtual actor output from the other information processing device 100B.
[0165] The dialogue history 950B is information that, in addition to the dialogue history 950, further includes the history of the dialogue between the virtual character of one information processing device 100B and the virtual actor of the other information processing device 100B.
[0166] (Control Unit) Compared to the control unit 140A, the control unit 140B includes an acquisition unit 141B, a determination unit 142B, an input information generation unit 143B, and a content generation unit 144B, instead of an acquisition unit 141, a determination unit 142A, an input information generation unit 143A, and a content generation unit 144.
[0167] (Acquisition Unit) The acquisition unit 141B further acquires instruction information 800B and dialogue history 950B. For example, the acquisition unit 141B acquires the response 700B related to the virtual actor's speech as instruction information 800B. The acquisition unit 141B also acquires the dialogue history 950B, which is the history of the dialogue between the virtual actor of one information processing device 100B and the virtual actor of the other information processing device 100B.
[0168] (Determination Unit) The determination unit 142B further determines whether the instruction information 800B and the dialogue history 950B have been acquired by the acquisition unit 141B. For example, if the acquisition unit 141B has acquired the response 700B related to the utterance of the virtual actor, the determination unit 142B determines that the instruction information 800B has been acquired by the acquisition unit 141B. Also, if the acquisition unit 141B has acquired the history of the dialogue between the virtual actor of one information processing device 100B and the virtual character of the other information processing device 100B, the determination unit 142B determines that the dialogue history 950B has been acquired.
[0169] Furthermore, if the determination unit 142B determines that the instruction information 800B and the dialogue history 950B have been acquired by the acquisition unit 141B, it stores the acquired instruction information 800B and dialogue history 950B in the storage unit 120B.
[0170] (Input Information Generation Unit) The input information generation unit 143B generates input information 400B as input information 400, which requests content 300 related to dialogue between multiple virtual actors. Below, an example of the generation of input information 400B by the input information generation unit 143B will be explained using Figure 25. Figure 25 is a diagram (7) for illustrating an example of input information generation.
[0171] The input information generation unit 143B generates input information 400B requesting content 300 relating to a dialogue between multiple virtual actors, based on the instruction information 800B and the dialogue history 950B, instead of the instruction information 800 and the dialogue history 950. For example, the input information generation unit 143B generates input information 400B requesting a response 700B relating to an utterance by one of the virtual actors, which forms the basis of the dialogue between the virtual actors.
[0172] Specifically, the input information generation unit 143B generates input information 400B that requests a response 700B regarding actor X's speech to actor Y by adding the following prompts to the template that forms the basis of the input information 400B: • "Ask the other person about their likes and dislikes" • "Discuss with each other the difficulties you faced in production H" • "Tell us about actor Y"
[0173] (Content Generation Unit) Returning to the explanation of Figure 24, the content generation unit 144B generates content 300 relating to a dialogue between multiple virtual actors based on the input information 400B.
[0174] For example, the content generation unit 144B generates content 300 related to dialogue between multiple virtual actors, such as utterances by actor X to actor Y, and responses 700B related to actor Y's utterances to actor X. In this way, the content generation unit 144B generates content 300 of a fan event in which virtual characters such as virtual actors interact with other virtual actors who are co-stars in the anime, as well as fans, about acting and the like.
[0175] Furthermore, the content generation unit 144B causes the communication unit 110 to transmit the generated response 700B to the other information processing device 100B, so that the other information processing device 100B acquires the response 700B as instruction information 800B, as shown in Figure 25. The content generation unit 144B also causes the other information processing device 100B to acquire the instruction information 800B corresponding to the response 700B as dialogue history 950B.
[0176] In the third embodiment, one of the information processing devices 100B comprising multiple virtual actors may act as a supervisor, creating a simulated dialogue between the virtual actors and generating content based on that dialogue. In this case, the information processing device 100B does not necessarily need to interact with the user, and can generate content requested by the virtual actors or content that conforms to the settings of the virtual actors by creating a simulated dialogue between the virtual actors.
[0177] (5. Other Embodiments) Of the processes described in the embodiments of this disclosure described above, all or part of the processes described as being performed automatically may be performed manually, or all or part of the processes described as being performed manually may be performed automatically by known methods. In addition, the processing procedures, specific names, and information including various data and parameters shown in the above documents and drawings may be changed at will unless otherwise specified. For example, the various information shown in each figure is not limited to the information shown.
[0178] Furthermore, the components of each illustrated device are functionally conceptual and do not necessarily need to be physically configured as shown. In other words, the specific forms of distribution and integration of each device are not limited to those shown, and all or part of them can be functionally or physically distributed and integrated in any unit according to various loads and usage conditions. For example, an information processing system 10 may be an integrated information processing device 100 and a terminal device 200.
[0179] Furthermore, the embodiments of this disclosure described above can be combined as appropriate in areas that do not contradict the processing content. Also, the order of each step shown in the sequence diagram or flowchart of this embodiment can be changed as appropriate. For example, each step may be processed chronologically, repeatedly, or partially in parallel.
[0180] (6. Effects of the Information Processing System Related to This Disclosure) The information processing method related to this disclosure is, for example, an information processing method that generates content (content 300 in this embodiment) using a computer. The information processing method includes an acquisition step (step S11 in this embodiment) and a content generation step (step S14 in this embodiment).
[0181] The acquisition step acquires work information (work information 500 in the embodiment) related to the content and virtual actor persona information (virtual actor persona information 600 in the embodiment) related to the virtual actor who plays the character appearing in the content. The generation step generates content including the character played by the virtual actor based on the acquired work information and virtual actor persona information.
[0182] This means that, for example, even if the work information includes smoking scenes, the information processing method generates content that conforms to the virtual actor's non-smoking setting based on the virtual actor's persona information. As a result, the information processing method can generate content in which the virtual actor does not smoke. Therefore, the information processing method can be used appropriately.
[0183] Furthermore, the information processing method further includes an input information generation step (step S13 in an embodiment) which generates input information (in an embodiment, input information 400) for a machine learning model that requests content suitable for a virtual actor, based on acquired work information and virtual actor persona information, and the content generation step generates content by inputting the generated input information into the machine learning model.
[0184] This allows the information processing method to generate input information, such as natural language text, that requests content suitable for the virtual actor's settings, and then input this generated input information into a machine learning model, thereby enabling the generation of content suitable for the virtual actor's settings from natural language and other sources.
[0185] Furthermore, the acquisition step acquires prohibited information (prohibited information 610 in the embodiment) regarding prohibited words and actions of the virtual actor as virtual actor persona information, and the input information generation step generates input information based on the acquired prohibited information.
[0186] As a result, for example, even if the work information includes smoking scenes corresponding to prohibited information, the information processing method generates input information that requests content that conforms to the virtual actor's non-smoking setting based on the prohibited information. Therefore, the information processing method can generate content in which the virtual actor does not smoke. Therefore, the information processing method can generate content that is even more suitable to the image of the virtual actor.
[0187] Furthermore, if the input information generation step includes prohibited behavior in the acquired work information, it generates response input information as input information requesting a response (response 700 in the embodiment) regarding the prohibited behavior, and the content generation step generates a response by inputting the generated response input information into a machine learning model.
[0188] As a result, even if the work information includes prohibited actions such as smoking scenes, the information processing method generates response input information requesting a response regarding the prohibited actions, such as, "Smoking scenes by actor X are prohibited. Please consider alternative staging." Therefore, the information processing method can generate content in which the virtual actor does not smoke, thus preventing deviation from the virtual actor's image.
[0189] Furthermore, the information processing method could also involve a specific virtual actor, extracted from work information such as the script of an anime, responding to the user and the anime character, thereby enabling interaction between the user and the anime character.
[0190] Furthermore, the input information generation step generates modification input information as response input information, requesting a response regarding proposed modifications to the prohibited behavior, and the content generation step generates a response regarding the modifications by inputting the generated modification input information into a machine learning model.
[0191] As a result, the information processing method generates input information for proposed revisions, such as, "Smoking scenes with actor X are prohibited. Please consider alternative staging." This allows the information processing method to prompt consideration of revisions, such as changing a smoking scene corresponding to a prohibited behavior to a scene of chewing gum.
[0192] The input information generation step generates corrected input information requesting content with the prohibited behavior corrected, if the prohibited behavior is included in the work information. The content generation step generates content with the prohibited behavior corrected by inputting the generated corrected input information into a machine learning model.
[0193] As a result, the information processing method can automatically generate animated or drama-like video content by generating corrected input information, for example, by correcting a smoking scene corresponding to a prohibited behavior indicated by prohibited information to a scene of chewing gum.
[0194] The acquisition step acquires casting information (casting information 900 in the embodiment), which is setting information for the character appearing in the content, as virtual actor persona information, and the input information generation step generates input information based on the acquired casting information.
[0195] This means that, for example, when generating video content such as anime, dramas, or concert videos, the information processing method can generate input information to match the virtual actor to the characters appearing in the content. The information processing method can generate input information as if the virtual actor had appeared in multiple videos. Therefore, the information processing method can increase the recognition of the virtual actor and enhance the value of the video content in which the virtual actor appears.
[0196] Furthermore, the acquisition step acquires information about the character played by the virtual actor as casting information, and the input information generation step generates input information based on the acquired information about the character.
[0197] This allows the information processing method to clarify, for example, that actor X, identified by the virtual actor persona information, will play character A, identified by the work information, based on casting information such as "actor X plays character A." Therefore, the information processing method can generate content in which the virtual actor and the content's casting are matched.
[0198] Furthermore, the acquisition step acquires instruction information (in the embodiment, instruction information 800) related to instructions from the user, and the input information generation step generates input information based on the acquired instruction information.
[0199] As a result, the information processing method can generate input information that requests content that conforms to the instruction information indicated by the user's utterance, for example, thus enabling the generation of content that is more closely aligned with the user's instructions.
[0200] Furthermore, the acquisition step involves acquiring instruction information when access from a specific user is received by a reception unit (reception unit 130 in this embodiment), which is accessible only to that specific user.
[0201] As a result, the information processing method can acquire instruction information, such as requests from video directors, only when access is accepted by a reception unit, such as a user interface accessible only to specific users, such as a virtual actor's talent agency. Therefore, the information processing method can, for example, respond specifically to requests from video directors, etc., only in this case.
[0202] Furthermore, the acquisition step acquires a dialogue history (dialogue history 950 in the embodiment), which is a history of interactions between the virtual actor and the user, and the input information generation step generates input information based on the acquired dialogue history.
[0203] As a result, the information processing method generates input information based on the dialogue history, so it can generate input information that requests content in response to suggestions from a virtual actor based on dialogue with the user, or comments from the virtual actor about their own performance. For example, the information processing method can change the performance of a character appearing in content based on dialogue between a user, such as a video director, and a virtual character such as a virtual actor.
[0204] Furthermore, the input information generation step generates alternative input information that requests alternative content as content based on the acquired dialogue history, and the content generation step generates alternative content by inputting the generated alternative input information into a machine learning model.
[0205] As a result, the information processing method can generate alternative input information such as "Smoking scenes for actor X are prohibited. Please consider other staging options," based on dialogue history, for example, "Input (User): So, create scene 1 with character A chewing gum." Therefore, the information processing method can generate content that is more suitable than the dialogue history, such as content in which prohibited behavior has been modified, like content in which a smoking scene has been changed to a gum-chewing scene.
[0206] Furthermore, the input information generation step generates alternative input information based on the artwork information.
[0207] This allows the information processing method to generate alternative input information requesting alternative content by, for example, paying homage to past content indicated by work information linked to the dialogue history. Specifically, the information processing method can generate alternative input information that includes text such as, "If you can also propose an homage to past works, please propose it if you think it's a very good suggestion." Therefore, the information processing method can generate alternative content such as paying homage to past content.
[0208] Furthermore, the information processing method further includes a determination step that determines the consistency of user characteristics based on the dialogue history, and the input information generation step generates input information based on the determined consistency.
[0209] As a result, the information processing method can generate input information only when it is determined that the user's characteristics are consistent, such as when a user persona is maintained, for example, when a video director interacting with a virtual actor is identified, or when it is determined that the user is the same person. Therefore, the information processing method can generate input information that instructs the creation of content that more accurately reflects the interaction with the user, while maintaining the consistency of the user persona.
[0210] Furthermore, the acquisition step acquires information about multiple virtual actors as virtual actor persona information, and the input information generation step generates information requesting content suitable for multiple virtual actors as input information, based on the acquired information about multiple virtual actors.
[0211] Thus, the information processing method can be applied even when there are multiple virtual actors, by generating information requesting content suitable for multiple virtual actors based on acquired information about multiple virtual actors.
[0212] Furthermore, the input information generation step generates information requesting content related to dialogue between multiple virtual actors as input information.
[0213] This allows the information processing method to generate input information that, for example, requests a fan event as content in which a virtual actor interacts with other virtual actors who are co-stars in the anime, as well as with fans, about acting, etc. Therefore, the information processing method can generate input information that allows the virtual actor to gain a deeper understanding of the director's intentions, communicate their own acting ideas, reflect those intentions and ideas in their acting (generated videos, etc.), and initiate relationships with other virtual actors.
[0214] Furthermore, the content generation step generates content in which a virtual actor appears in real space as a robot.
[0215] This allows the information processing method to generate content, for example, for a virtual actor, given a robotic body, to interact with a user in the real world. Therefore, the information processing method can provide the user with an experience similar to interacting with a real actor.
[0216] Furthermore, the content generation step generates input information by inputting the acquired work information and virtual actor persona information into a separate machine learning model from the machine learning model.
[0217] As a result, the information processing method generates input information by using machine learning models, making it easier to generate input information even when various types of information, such as work information and virtual actor persona information, are in an indeterminate form such as images or audio rather than text.
[0218] Furthermore, the input information generation step generates at least one of the following as input information to the machine learning model: text, images, and audio, which will be input to the large-scale language model.
[0219] This allows the information processing method to generate content not only by inputting text, but also by inputting images, audio, and other information into a large-scale language model. Therefore, the information processing method can, for example, respond to the user with content output from a large-scale language model by inputting not only text received from the user, but also input generated based on images and audio. As a result, the information processing method can generate moving images and other content through interaction with the user.
[0220] For example, information processing methods such as ChatGPT (Generative Pre-trained Transformer) (GPT-4o), which are large-scale language models, can generate images as content that follow the flow of the conversation.
[0221] As an example, let's consider a case where an information processing device (information processing device 100 in this embodiment) generates an image based on the instruction information, "Generate an image of grilled mochi on a plate," and then receives the instruction information, "Make the number of mochi two." In this case, the information processing method can generate an image of "two grilled mochi on a plate" based on the instruction information, "Make the number of mochi two," through the flow of the dialogue.
[0222] (7. Hardware Configuration) The information processing device 100 and the like according to the embodiments of this disclosure described above are realized by a computer 1000 having the configuration shown in Figure 26, for example. The information processing device 100 will be explained as an example. Figure 26 is a hardware configuration diagram showing an example of a computer that realizes the functions of the information processing device. The computer 1000 has a processing circuitry 1100, RAM 1200, ROM 1300, secondary storage device 1400, communication interface 1500, input / output interface 1600, display unit 1700, camera unit 1800, microphone 1900, and speaker 2000. The parts of the computer 1000 are connected by a bus 1050.
[0223] The processing circuit 1100 operates based on a program stored in the ROM 1300 or secondary storage device 1400, and controls each part. For example, the processing circuit 1100 loads the program stored in the ROM 1300 or secondary storage device 1400 into the RAM 1200 and executes processing corresponding to various programs.
[0224] ROM 1300 stores boot programs such as the BIOS (Basic Input Output System) executed by the processing circuit 1100 when the computer 1000 starts up, as well as programs that depend on the computer 1000's hardware.
[0225] The secondary storage device 1400 is a computer-readable recording medium that non-temporarily records programs executed by the processing circuit 1100 and data used by such programs. Specifically, the secondary storage device 1400 is a recording medium that records programs for each process of the information processing device 100 according to the embodiment of this disclosure, which is an example of program data 1450.
[0226] The communication interface 1500 is an interface for the computer 1000 to connect to an external network 1550. The communication interface 1500 corresponds to the communication unit 110 of the information processing device 100. For example, the processing circuit 1100 receives data from other devices or transmits data generated by the processing circuit 1100 to other devices via the communication interface 1500.
[0227] The input / output interface 1600 is an interface for connecting the input / output device 1650 and the computer 1000. For example, the processing circuit 1100 receives data from input devices such as a microphone 1900 or a touch panel via the input / output interface 1600. The processing circuit 1100 also transmits data to output devices such as a display unit 1700 or a speaker 2000 via the input / output interface 1600. The input / output interface 1600 may also function as a media interface for reading programs recorded on a predetermined recording medium (media). Examples of media include optical recording media such as DVDs (Digital Versatile Discs) and PDs (Phase Change Rewritable Disks), magneto-optical recording media such as MOs (Magneto-Optical Disks), tape media, magnetic recording media, or semiconductor memory.
[0228] The display unit 1700 is an interface for displaying information processed by the computer 1000. The display unit 1700 is, for example, a liquid crystal display or an organic electroluminescent display (Organic Electro Luminescence Display). Alternatively, the display unit 1700 may be a touch panel display device or an image projection device.
[0229] The camera unit 1800 is an interface for the computer 1000 to capture images. The microphone 1900 is an interface for the computer 1000 to capture sound. The speaker 2000 is an interface for the computer 1000 to output processed sound. The various parts of the computer 1000 are connected by the bus 1050. Each interface does not necessarily have to be located inside the computer 1000, but may be located outside the computer 1000 via a network or the like. Furthermore, each part of the computer 1000 may be controlled by a circuit different from the processing circuit 1100. For example, the display unit 1700 may be controlled not by the processing circuit 1100, but by a circuit dedicated to display processing provided within the display unit 1700.
[0230] For example, when computer 1000 functions as an information processing device 100 according to an embodiment of this disclosure, the processing circuit 1100 of computer 1000 functions as a control unit 140 by executing a program loaded onto RAM 1200. The secondary storage device 1400 stores the information processing program according to this disclosure and various data stored by the storage unit 120. The processing circuit 1100 reads and executes program data 1450 from the secondary storage device 1400, but as another example, these programs may be obtained from other devices via an external network 1550. In other words, the secondary storage device 1400 is not limited to being inside computer 1000, but may be located outside computer 1000. The processing circuit 1100 is an example of an integrated circuit, and CPU, MPU, GPU, APU, ASIC, and FPGA can all be considered integrated circuits.
[0231] (8. Addendum) The technology can also be configured as follows: (1) An information processing method comprising: an acquisition step of acquiring work information relating to content and virtual actor persona information which is information of a virtual actor who plays a character appearing in the content; and a content generation step of generating the content including the character played by the virtual actor based on the acquired work information and virtual actor persona information. (2) The information processing method according to (1), further comprising an input information generation step of generating input information for a machine learning model that requests content suitable for the virtual actor, based on the acquired work information and virtual actor persona information, wherein the content generation step generates the content by inputting the generated input information into the machine learning model. (3) The information processing method according to (2), wherein the acquisition step acquires prohibited information relating to prohibited speech and actions of the virtual actor as the virtual actor persona information, and the input information generation step generates the input information based on the acquired prohibited information. (4) The information processing method according to (3), wherein the input information generation step generates response input information requesting a response regarding the prohibited behavior if the prohibited behavior is included in the acquired work information, and the content generation step generates the response by inputting the generated response input information into the machine learning model. (5) The information processing method according to (4), wherein the input information generation step generates modification input information requesting a response regarding a modification of the prohibited behavior as response input information, and the content generation step generates a response regarding the modification by inputting the generated modification input information into the machine learning model.(6) The information processing method according to (5), wherein the input information generation step generates corrected input information requesting content in which the prohibited behavior has been corrected, if the prohibited behavior is included in the work information, and the content generation step generates content in which the prohibited behavior has been corrected by inputting the generated corrected input information into the machine learning model. (7) The information processing method according to any one of (2) to (5), wherein the acquisition step acquires casting information which is setting information of a character appearing in the content, and the input information generation step generates the input information based on the acquired casting information. (8) The information processing method according to (7), wherein the acquisition step further acquires information relating to a character played by a virtual actor as casting information, and the input information generation step generates the input information based on the acquired information relating to the character. (9) The information processing method according to any one of (2) to (8), wherein the acquisition step further acquires instruction information relating to instructions from the user, and the input information generation step generates the input information based on the acquired instruction information. (10) The information processing method according to (9), wherein the acquisition step acquires the instruction information when access from a specific user is accepted by a reception unit accessible only to a specific user. (11) The information processing method according to any one of (2) to (9), wherein the acquisition step further acquires a dialogue history which is a history of dialogue between the virtual actor and the user, and the input information generation step generates the input information based on the acquired dialogue history. (12) The information processing method according to (11), wherein the input information generation step generates alternative input information requesting an alternative to the content as the content based on the acquired dialogue history, and the content generation step generates an alternative to the content by inputting the generated alternative input information into the machine learning model.(13) The information processing method according to (12), wherein the input information generation step further generates the alternative input information based on the work information. (14) The information processing method according to any one of (11) to (13), further comprising a determination step of determining the consistency of the user's characteristics based on the dialogue history, wherein the input information generation step further generates the input information based on the determined consistency. (15) The information processing method according to any one of (2) to (13), wherein the acquisition step acquires information about a plurality of virtual actors as virtual actor persona information, and the input information generation step generates information requesting content suitable for the plurality of virtual actors as input information based on the acquired information about the plurality of virtual actors. (16) The information processing method according to (15), wherein the input information generation step generates information requesting content relating to dialogue between the plurality of virtual actors as input information. (17) The information processing method according to any one of (1) to (16), wherein the content generation step generates content in which the virtual actor appears in real space as a robot. (18) The information processing method according to any one of (2) to (16), wherein the content generation step generates input information by inputting the acquired work information and virtual actor persona information into a machine learning model other than the machine learning model. (19) The information processing method according to any one of (2) to (16), wherein the input information generation step generates at least one of text, images, and sounds to be input into a large-scale language model as input information to the machine learning model. (20) An information processing system comprising: an acquisition unit that acquires work information relating to content and virtual actor persona information which is information of a virtual actor who plays a character appearing in the content; and a content generation unit that generates content including the character played by the virtual actor based on the acquired work information and virtual actor persona information.(21) An information processing program for causing a computer to function as an information processing system, comprising: an acquisition unit that acquires work information relating to content and virtual actor persona information which is information of a virtual actor who plays a character appearing in the content; and a content generation unit that generates the content including the character played by the virtual actor based on the acquired work information and virtual actor persona information.
[0232] 10 Information Processing System 100 Information Processing Device 110 Communication Unit 120 Storage Unit 130 Reception Unit 140 Control Unit 141 Acquisition Unit 142 Judgment Unit 143 Input Information Generation Unit 144 Content Generation Unit 200 Terminal Device 210 Reception Unit 220 Communication Unit 230 Display Unit 240 Control Unit 500 Work Information 600 Virtual Actor Persona Information 800 Instruction Information 900 Casting Information 950 Dialogue History N Network
Claims
1. An information processing method comprising: an acquisition step of acquiring work information relating to content and virtual actor persona information, which is information of a virtual actor who plays a character appearing in the content; and a content generation step of generating the content, which includes the character played by the virtual actor, based on the acquired work information and virtual actor persona information.
2. The information processing method according to claim 1, further comprising an input information generation step of generating input information for a machine learning model that requests content suitable for the virtual actor, based on the acquired work information and virtual actor persona information, wherein the content generation step generates the content by inputting the generated input information into the machine learning model.
3. The information processing method according to claim 2, wherein the acquisition step acquires prohibited information regarding prohibited speech and actions of the virtual actor as virtual actor persona information, and the input information generation step generates the input information based on the acquired prohibited information.
4. The information processing method according to claim 3, wherein the input information generation step generates response input information requesting a response regarding the prohibited behavior if the prohibited behavior is included in the acquired work information, and the content generation step generates the response by inputting the generated response input information into the machine learning model.
5. The information processing method according to claim 4, wherein the input information generation step generates modification input information requesting a response regarding the modification of the prohibited behavior as the response input information, and the content generation step generates a response regarding the modification by inputting the generated modification input information into the machine learning model.
6. The information processing method according to claim 5, wherein the input information generation step generates corrected input information requesting content in which the prohibited behavior has been corrected, if the prohibited behavior is included in the work information, and the content generation step generates content in which the prohibited behavior has been corrected by inputting the generated corrected input information into the machine learning model.
7. The information processing method according to claim 2, wherein the acquisition step acquires casting information, which is setting information for characters appearing in the content, and the input information generation step generates the input information based on the acquired casting information.
8. The information processing method according to claim 7, wherein the acquisition step further acquires information relating to a character played by the virtual actor as the casting information, and the input information generation step generates the input information based on the acquired information relating to the character.
9. The information processing method according to claim 2, wherein the acquisition step further acquires instruction information relating to instructions from the user, and the input information generation step generates the input information based on the acquired instruction information.
10. The information processing method according to claim 9, wherein the acquisition step is to acquire the instruction information when access from a specific user is accepted by a reception unit accessible only to that specific user.
11. The information processing method according to claim 2, wherein the acquisition step further acquires a dialogue history which is a history of the dialogue between the virtual actor and the user, and the input information generation step further generates the input information based on the acquired dialogue history.
12. The information processing method according to claim 11, wherein the input information generation step generates alternative input information requesting alternative content as content based on the acquired dialogue history, and the content generation step generates alternative content by inputting the generated alternative input information into the machine learning model.
13. The information processing method according to claim 12, wherein the input information generation step further generates the alternative input information based on the work information.
14. The information processing method according to claim 11, further comprising a determination step of determining the consistency of the user's characteristics based on the dialogue history, wherein the input information generation step generates the input information based on the determined consistency.
15. The information processing method according to claim 2, wherein the acquisition step acquires information about a plurality of virtual actors as the virtual actor persona information, and the input information generation step generates information requesting content suitable for the plurality of virtual actors as the input information, based on the acquired information about the plurality of virtual actors.
16. The information processing method according to claim 15, wherein the input information generation step generates information requesting content relating to a dialogue between the plurality of virtual actors as the input information.
17. The information processing method according to claim 1, wherein the content generation step generates content in which the virtual actor appears in real space as a robot.
18. The information processing method according to claim 2, wherein the content generation step generates the input information by inputting the acquired work information and virtual actor persona information into a machine learning model separate from the machine learning model.
19. The information processing method according to claim 2, wherein the input information generation step generates at least one of text, images, and audio to be input to a large-scale language model as input information to the machine learning model.
20. An information processing system comprising: an acquisition unit that acquires work information relating to content and virtual actor persona information, which is information of a virtual actor who plays a character appearing in the content; and a content generation unit that generates the content, including the character played by the virtual actor, based on the acquired work information and virtual actor persona information.
21. An information processing program for causing a computer to function as an information processing system, comprising: an acquisition unit that acquires work information relating to content and virtual actor persona information, which is information of a virtual actor who plays a character appearing in the content; and a content generation unit that generates the content, including the character played by the virtual actor, based on the acquired work information and virtual actor persona information.
Citation Information
Patent Citations
Method and apparatus for producing animation
JP2010020781A
Virtual character creating system as preliminary stage of project by virtual character
JP2021028792A
Movie generation device and movie generation system
JP2024066971A