Computer systems and programs
The computer system addresses inefficient NPC response times by pre-preparing hypothetical conversations using AI, ensuring natural and engaging interactions in video games.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2026-04-08
AI Technical Summary
Existing video game technologies struggle with inefficient and unnatural NPC response times due to the use of generation units, leading to slowed conversational pacing and user dissatisfaction.
A computer system utilizing a generation unit to pre-prepare hypothetical conversation information, enabling quick response control by selecting hypothetical responses based on user inputs, and incorporating server-side and terminal-side AI to manage and generate NPC responses efficiently.
The system reduces response generation time, allowing for natural and efficient NPC conversations, enhancing user engagement by maintaining a conversational pace similar to human interactions.
Smart Images

Figure 2026060040000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a computer system and the like.
Background Art
[0002] Communication with NPCs (Non Player Characters) that appear in video games is an important factor affecting the charm of the game world. For example, various technologies have been devised regarding speech, which is a representative example of communication.
[0003] Patent Document 1 describes a technique in which a plurality of lines and a plurality of display conditions are associated and stored in advance, and even if the user does not input a line, a line that conforms to the display conditions is automatically selected according to the game situation at that time and displayed within the game screen.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] In recent years, the use of a generation unit represented by generative AI (Artificial Intelligence) in video games has been explored.
[0006] The problem to be solved by the present invention is to realize the control when a user and an NPC communicate by using a generation unit.
Means for Solving the Problems
[0007] The first invention for solving the above problems is a computer system that performs conversation control to realize a conversation with a user by repeatedly performing response control in response to user input, and comprises: an assumed conversation information preparation means (for example, the assumed conversation information preparation unit 210 in Figure 25, step S24 in Figure 27) that prepares assumed conversation information by predicting future assumed user input and generating a hypothetical response in response to the predicted user input based on conversation situation information including information on past user input; and a response control execution means (for example, the response control execution unit 220 in Figure 25, step S208 in Figure 30) that performs response control of a new response in response to new user input using the assumed conversation information.
[0008] The time required to generate an NPC's response to user input may be longer than the response time for a conversation between natural people. According to the first invention, a computer system can pre-prepare hypothetical conversation information, including user input in a hypothetical conversation and hypothetical responses to it. The computer system can quickly perform response control of an NPC simply by selecting a hypothetical response corresponding to a new user input from the hypothetical conversation information. In other words, control of communication between NPCs appearing in a video game can be realized using the generation unit.
[0009] The second invention is a computer system in which, in the above-described computer system, the assumed conversation information preparation means includes a hypothetical response generation / failure determination means (for example, the hypothetical response generation / failure determination unit 212 in Figure 25, the past user input information 740 of the conversation situation information 730 in Figure 8, and the loop C in Figure 29) that determines whether or not to generate the hypothetical response based on the conversation situation information.
[0010] According to the second invention, the computer system will be able to determine whether or not to generate a hypothetical response based on conversational context information.
[0011] The third invention is a computer system in which, in the above-described computer system, the assumed conversation information preparation means includes a hypothetical response generation / failure determination means (for example, the hypothetical response generation / failure determination unit 212 in Figure 25, step S116 in Figure 29) that determines whether or not to generate the hypothetical response based on the inference result of the inferred user input.
[0012] According to the third invention, the computer system can determine whether or not to generate a hypothetical response based on the inferred result of the inferred user input.
[0013] The fourth invention is a computer system in which, in the above-described computer system, the assumed conversation information preparation means infers a plurality of assumed user inputs, and the assumed response generation / determination means determines whether or not to generate the assumed response for each of the assumed user inputs (for example, steps S105 and S106 in Figure 29).
[0014] According to the fourth invention, the computer system can reduce the time required to generate hypothetical responses by determining whether or not to generate a hypothetical response for each of the multiple inferred user inputs.
[0015] The fifth invention is a computer system in which, in the above-described computer system, the assumed conversation information preparation means prepares assumed conversation information by performing an inference of the assumed user input for the assumed response and repeating the inference and generation of the assumed response, thereby advancing the exchange of assumed conversation.
[0016] According to the fifth invention, the computer system can prepare hypothetical conversation information by repeatedly inferring user input and generating hypothetical responses, thereby advancing the exchange of a hypothetical conversation.
[0017] The sixth invention is a computer system in which, in the above-described computer system, the assumed conversation information preparation means performs the estimation of a plurality of assumed user inputs and the generation of a hypothetical response for each of the assumed user inputs, and prepares assumed conversation information for a tree-structured assumed conversation by repeating the estimation and generation of the hypothetical response by performing the estimation of a plurality of assumed user inputs for the hypothetical response.
[0018] According to the sixth invention, the computer system becomes capable of preparing hypothetical conversation information for a tree-structured hypothetical conversation.
[0019] The seventh invention is a computer system in which, in the above-described computer system, the assumed conversation information preparation means controls the number of layers for each branch destination in the tree-structured assumed conversation by controlling the number of repetitions (for example, steps S114 to S118 in Figure 29).
[0020] According to the seventh invention, the computer system can control how far it predicts the exchange of a hypothetical conversation at each branch in a tree-structured hypothetical conversation.
[0021] The eighth invention is a computer system in which, in the above-described computer system, the assumed conversation information preparation means determines the importance of each branch destination and controls the number of repetitions based on that importance (for example, steps S70 to S78 in Figure 28).
[0022] According to the eighth invention, the computer system can determine the importance of each branching point and control how far it predicts the expected conversational exchange based on that importance.
[0023] The ninth invention is a computer system in which, in the above-described computer system, the assumed conversation information preparation means controls the number of repetitions based on the user information of the user (for example, steps S74 to S78 in Figure 28).
[0024] According to the ninth invention, the computer system can control the number of repetitions of user input and NPC responses based on the user information of the user.
[0025] The tenth invention is a computer system that further includes response time length determination means (for example, the response time determination unit 208 in FIG. 25, the response control definition data 562 in FIG. 10, step S20 in FIG. 27) for determining the response time length which is the length of the response in the above computer system, and the assumed conversation information preparation means generates the assumed response based on the response time length (for example, the description of the response character count limit in the generation condition description 752 of the comprehensive generation instruction information 750 in FIG. 13).
[0026] According to the tenth invention, the computer system can generate an assumed response based on the response time length.
[0027] The eleventh invention is a computer system in the above computer system, where the response control item candidates include the lines of the NPC (non-player character), the speaking voice of the NPC, the motion of the NPC, and the effects during the response, and the response control execution means executes the response by controlling at least one of the control item candidates (for example, step S40 in FIG. 28, step S208 in FIG. 30).
[0028] According to the eleventh invention, the computer system can express the response by at least one of the lines of the NPC, the speaking voice of the NPC, the motion of the NPC, and the effects during the response.
[0029] The twelfth invention is a computer system in the above computer system, where the response includes the control contents of a plurality of the control item candidates, and the assumed conversation information preparation means generates the assumed response by variably determining the generation order of the control item candidates (for example, step S40 in FIG. 28, step S106 in FIG. 29).
[0030] According to the twelfth invention, the computer system can generate a hypothetical response by variably determining the generation order of candidate control items.
[0031] The 13th invention is a computer system in which, in the above-described computer system, the assumed conversation information preparation means generates the assumed response using a predetermined generation unit (for example, the server-side generation AI 20 in Figure 1), the generation unit has a server system-side generation unit and a user terminal-side generation unit (for example, the server-side generation unit 230 and the user terminal-side generation unit 230t in Figure 1), and the assumed conversation information preparation means variably determines which generation unit, the server system-side generation unit or the user terminal-side generation unit, to generate the assumed response (for example, step S40 in Figure 28).
[0032] According to the 13th invention, the computer system can variably determine whether to have the generation unit on the server system side or the generation unit on the user terminal side generate the hypothetical response.
[0033] The fourteenth invention is a computer system in which, in the above-described computer system, the assumed conversation information preparation means generates the assumed response using a predetermined generation unit, the generation unit has a detailed generation unit (for example, the detailed generation AI23 (23a, 23b, ...) in Figure 11) that generates the control content of the response, and a general generation unit (for example, the general generation AI21 in Figure 11) that causes the detailed generation unit to generate the control content of the response based on given generation instruction information, and the assumed conversation information preparation means generates the assumed response by providing the generation instruction information to the general generation unit (for example, step S60 in Figure 28).
[0034] According to the 14th invention, the computer system can generate a hypothetical response by using a detailed generation unit and a general generation unit, and by providing generation instruction information to the general generation unit.
[0035] The 15th invention is a computer system in which the assumed conversation information preparation means generates the assumed response using a predetermined generation unit (for example, steps S42 to S44 and S60 in Figure 28), determines the importance of each branch destination in the tree-structured assumed conversation (for example, steps S76 to S78 in Figure 28), and regenerates the assumed response related to the branch destination using the generation unit based on the importance (for example, assumed response content change definition data 566 in Figure 10 and step S80 in Figure 28).
[0036] According to the 15th invention, the computer system becomes capable of regenerating hypothetical responses related to each branch in a tree-structured hypothetical conversation based on the importance of each branch.
[0037] The sixteenth invention is a computer system in which, if the new user input and the presumed user input do not satisfy predetermined compatibility conditions, the response control execution means generates a new response (for example, steps S216 to S220 in Figure 31), and if the new user input and the presumed user input satisfy the compatibility conditions, the response based on the assumed response is used as the new response and response control is executed (for example, steps S202 to S208 in Figure 30).
[0038] According to the 16th invention, the computer system can generate a new response and control the response if the new user input and the inferred user input do not satisfy predetermined compatibility conditions, and can control the response with a hypothetical response prepared in advance as assumed conversation information if the compatibility conditions are met.
[0039] The 17th invention is a computer system in which, in the above-described computer system, the response control execution means executes response control with a predetermined specific response as the new response if the new user input and the presumed user input do not satisfy predetermined compatibility conditions (for example, step S210 in Figure 31), and executes response control with a response based on the assumed response as the new response if the new user input and the presumed user input satisfy the compatibility conditions.
[0040] According to the 17th invention, the computer system can perform response control as a new response if the new user input and the inferred user input do not satisfy the matching conditions.
[0041] The 18th invention is a program for causing a computer system to perform conversation control to realize a conversation with a user by repeatedly performing response control in response to user input, and is a program for causing the computer system to function as follows: a pre-planned conversation information preparation means that prepares pre-planned conversation information by predicting future pre-planned user input and generating a hypothetical response in response to the pre-planned user input based on conversation situation information including information on past user input; and a response control execution means that performs response control of a new response in response to new user input using the pre-planned conversation information.
[0042] According to the 18th invention, it is possible to realize a program that can enable a computer system to perform the same functions as the first invention. [Brief explanation of the drawing]
[0043] [Figure 1] A system configuration diagram showing an example of a content delivery system. [Figure 2] A diagram illustrating an example of content. [Figure 3] A diagram illustrating the principle of response control related to conversational events. [Figure 4] A diagram to explain the information in the anticipated conversation. [Figure 5] A diagram showing an example of the data structure for first user input information, first hypothetical response information, inferred user input information, and hypothetical response information. [Figure 6] A diagram to explain the continuation of the forecast. [Figure 7] A diagram to explain importance. [Figure 8] A diagram showing an example of the data structure for conversation context information. [Figure 9] A diagram showing an example of the data structure of character initial settings data. [Figure 10] A diagram illustrating the machine learning of a comprehensive generative AI. [Figure 11] This diagram illustrates an example configuration for server-side generated AI. [Figure 12] A diagram illustrating the machine learning of a comprehensive generative AI. [Figure 13] A diagram showing an example of a comprehensive generation instruction information written in natural language. [Figure 14] A diagram illustrating machine learning in text generation AI. [Figure 15] A diagram showing an example of detailed generation instruction information for text generation written in natural language. [Figure 16] A diagram illustrating machine learning in speech generation AI. [Figure 17] A diagram showing an example of detailed generation instruction information for speech generation written in natural language. [Figure 18] A diagram illustrating machine learning in motion generation AI. [Figure 19] A diagram showing an example of detailed motion generation instruction information written in natural language. [Figure 20] A diagram illustrating machine learning in VFX generation AI. [Figure 21] A diagram showing an example of detailed generation instruction information for VFX generation written in natural language. [Figure 22] A diagram illustrating the machine learning process used in AI for generating sound effects (SE). [Figure 23] A diagram showing an example of detailed generation instruction information for SE generation written in natural language. [Figure 24] A diagram showing examples of programs and data stored by a server system. [Figure 25] A diagram showing an example of the functional configuration of the server processing unit. [Figure 26] A diagram showing an example of the data structure of conversation management data. [Figure 27] A flowchart illustrating the flow of processing performed by the server system from the occurrence of a conversation event during content delivery. [Figure 28] A flowchart illustrating the process of preparing anticipated conversation information. [Figure 29] Flowchart continuing from Figure 28. [Figure 30] A flowchart illustrating the flow of the response continuation process. [Figure 31] Flowchart continuing from Figure 30. [Figure 32] A diagram showing an example of comprehensive generation instruction information for additional preparation, written in natural language. [Figure 33] A flowchart illustrating the flow of additional preparation processes. [Figure 34] A diagram illustrating a modified example. [Figure 35] A diagram illustrating a modified example. [Figure 36] A diagram illustrating a modified example. [Figure 37] A diagram illustrating a modified example. [Modes for carrying out the invention]
[0044] Examples of embodiments of the present invention will be described below, but it goes without saying that the embodiments to which the present invention can be applied are not limited to the following embodiments.
[0045] Figure 1 is a system configuration diagram showing an example of the configuration of a content provision system according to this embodiment. The content provision system 1000 is a computer system that provides content offering virtual experiences in a virtual space, such as video games, virtual activities, and shopping, to multiple registered users. In other words, the content provision system 1000 provides virtual experiences to users.
[0046] The content provision system 1000 is a computer system that includes a server system 1100 and user terminals 1500 for each user, all connected via a network 9 for data communication.
[0047] Network 9 refers to a communication path capable of data transmission. In other words, Network 9 includes not only LANs (Local Area Networks) using dedicated lines (dedicated cables) or Ethernet (registered trademark) for direct connections, but also telephone networks, cable networks, and the Internet.
[0048] The server system 1100 is a computer system that performs various processes such as managing and controlling registered user information and various controls related to content provision (for example, controlling the progress of a game), and has a database 1140.
[0049] The server system 1100 has a control board 1150 mounted on the main unit 1101. The control board 1150 is equipped with various microprocessors such as a CPU (Central Processing Unit) 1151, a GPU (Graphics Processing Unit), and a DSP (Digital Signal Processor), various IC memories 1152 such as VRAM, RAM, and ROM, and a communication device 1153. Some or all of the functions mounted on the control board 1150 may be implemented using an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or a SoC (System on a Chip).
[0050] Although Figure 1 depicts the server system 1100 as a single server device, it may also be implemented using multiple devices. For example, the server system 1100 may be configured with multiple servers, each responsible for a specific function, connected to each other via an internal bus or network 9 for data communication.
[0051] The server system 1100 has a server-side generated AI 20. The server-side generating AI 20 is implemented using a machine learning-based AI model on hardware employing a multi-core architecture (e.g., a group of GPUs and memory, a group of AI chips, etc.). The hardware for the server-side generating AI 20 is not limited to being exclusive to the server system 1100. The server-side generating AI 20 may be implemented as a single multimodal generating AI model, or it may be configured as an AI group having one or more generating AIs for each type of data to be generated (e.g., text, speech, motion, etc.). In other words, the server-side generating AI 20 is the generation unit 230 on the server system side.
[0052] User terminal 1500 serves as a Man-Machine Interface (MMIF) for user 2 to play content as a player. For example, if the content to be played is a game, user terminal 1500 functions as a gameplay terminal for user 2, who is the player. Although only one user terminal 1500 is depicted in Figure 1, in actual operation, it is common for multiple user terminals 1500 to communicate and connect to the server system 1100 simultaneously.
[0053] The user terminal 1500 is a computer system that can connect to the network 9, such as a personal computer, smartphone, wearable computer, portable game console, home game console, or tablet computer.
[0054] The user terminal 1500 is a computer comprising an operation input device, an image display device, a communication device, and a control board 1550 for performing calculations. Examples of operation input devices include a touch panel 1506, a keyboard, a game controller, and a mouse. Examples of image display devices include a touch panel 1506, a head-mounted display, and a glasses-type display.
[0055] The control board 1550 is equipped with a CPU 1551, various microprocessors such as a GPU and DSP, various IC memories 1552 such as VRAM, RAM, and ROM, and a communication module 1553 that connects to the network 9. These elements mounted on the control board 1550 are electrically connected via bus circuits and the like, enabling data reading and writing, and signal transmission and reception. Part or all of the control board 1550 may be an ASIC, FPGA, or SoC.
[0056] The control board 1550 stores programs and various data necessary to realize the functions of the user terminal 1500 in the IC memory 1552. The user terminal 1500 realizes the functions of a play terminal for playing content by executing a predetermined application program.
[0057] The user terminal 1500 has a terminal-side generated AI 30. The terminal-side generation AI 30 is implemented using a machine learning-based AI model on hardware employing a multi-core architecture (e.g., a group of GPUs and memory, a group of AI chips, etc.). The hardware of the terminal-side generation AI 30 is not limited to being exclusive to the user terminal 1500. The terminal-side generation AI 30 may be implemented as a single multimodal generation AI model, or it may be configured as an AI group having one or more generation AIs for each type of data to be generated (e.g., text, voice, motion, etc.). In other words, the terminal-side generation AI 30 is the generation unit 230t on the user terminal side. In the following explanation, the terminal-side generation AI 30 will be assumed to be a standalone generation AI that generates character dialogue text.
[0058] Figure 2 is a diagram illustrating an example of content. The content provision system 1000 provides a video game as content, in which a player character 4 and NPCs 5 (5a, 5b, ...) appear in a game space constructed in a virtual three-dimensional space. The game genre can be set as appropriate, but for example, as shown in the example of game screen W2, it may be an action RPG (Role Playing Game) in which the story unfolds as the player character 4 and NPCs 5 fight together as allies against enemy NPCs 6.
[0059] Figure 3 is a diagram illustrating the principle of response control related to conversational events. During gameplay, when a conversation event occurs, it becomes possible for player character 4 and NPC 5 to converse within that event.
[0060] For example, when User 2 inputs a statement from Player Character 4 as the first line of conversation (First User Input), the response-controlled NPC 5 responds to Player Character 4's statement (First Response). User 2 responds to this response from NPC 5 with the next input (Second User Input). The response-controlled NPC 5 then responds to that (Second Response). In this way, the response control in response to user input is repeated to realize a conversation between the user and NPC 5.
[0061] The player character 4's speech may be, for example, text input using a software keyboard, or voice input using the microphone installed in the user terminal 1500.
[0062] The NPC 5 who interacts with player character 4 is called the "opponent character." The opposing character changes depending on the situation and the type of conversation event that occurs. For example, in a conversation event where player character 4 buys fruit at a fruit shop, the opposing character is the shopkeeper NPC. Also, for example, in a conversation event that occurs during combat, the opposing character will be the ally NPC 5.
[0063] The first line of dialogue is not limited to the first user input; it may also be a predetermined line of dialogue that prompts NPC5 to speak. Hereafter, we will assume that the first line of dialogue is the first user input.
[0064] Conversation events have a conversation objective, and they end when that objective is achieved. The "conversation objective" is determined for each conversation event. For example, suppose a conversation event involves player character 4 purchasing fruit at a fruit shop. In this example, the conversation objective is considered achieved if the user input during the conversation includes three elements: the name of the product to be purchased, the quantity of each product to be purchased, and approval of the final purchase price. Alternatively, in the case of a conversation event during combat, the conversation objective may be considered achieved if any hints or advice regarding defeating enemy NPC 6 appear in the user input or in the response of NPC 5.
[0065] Traditionally, NPC5's responses were pre-written by the game developer for each conversation event. However, this meant that NPC5 would only give the same response for the same conversation event. Therefore, the content delivery system 1000 generates NPC5's responses using server-side AI 20 and terminal-side AI 30, so that even for the same conversation event, NPC5's responses are not necessarily the same.
[0066] However, when generating responses for NPC5 using a generation AI, it is necessary to create generation instruction information to give to the generation AI, provide this information to the generation AI, and then wait for the output of the generated data. In cases where the generation AI is implemented using an external device, the communication time with that external device will also be included in the waiting time.
[0067] Furthermore, the waiting time is often longer than the time it takes from speaking to responding in a conversation between two natural people, which can slow down the pace of the conversation between player character 4 and NPC 5, and in some cases, make user 2 feel that it is unnatural.
[0068] Therefore, in the content provision system 1000, as a background process, when the first user input is received, a generating AI is used to predict further anticipated conversations based on the information from the first user input, and future responses are prepared in advance.
[0069] The assumed conversation information 770 generated by background processing includes a first user input, a first hypothetical response which is a prediction of the first response of NPC5 in response to that input, a predicted user input which is the result of predicting the next user input in response to the first response, and a hypothetical response which is a prediction of the response in response to the predicted user input.
[0070] The first hypothetical response is generated by a generation AI that includes a closed question. The inferred user input includes the answer to this closed question. The hypothetical responses to inferred user input are generated by the AI based on past user input and responses from the start of the conversation up to that point, and include the next closed question to satisfy the conversational objective.
[0071] While User 2 is making their first user input and listening to NPC 5's first response, or while User 2 is making their second user input, background processing is completed.
[0072] When user 2 provides a second user input, the server system 1100 searches the estimated user inputs in the assumed conversation information 770 for an estimated user input that matches the content of the second user input (identical or highly similar). Then, it adopts the hypothetical response corresponding to the retrieved estimated user input as the second response of NPC 5 and controls the response.
[0073] At first glance, one might suspect that the time required from the first user input to the start of response control for the second response will be longer. However, the total time required is much shorter than a continuous generation method where the first response is generated in response to the first user input, and then the second response is generated in response to the second user input. Therefore, by using the generation AI, it becomes possible to realize a conversation between user 2 and NPC 5 that is lighter and closer to a natural conversational pace than before.
[0074] Figure 4 is a diagram illustrating the hypothetical conversational information 770 in actual operation based on the principle explained in Figure 3. The assumed conversation information 770 has multiple layers, such as the first user input information 772 relating to the first user input being the first layer, the first hypothetical response information 774 which forms the basis of the first response being the second layer, the inferred user input information 776 relating to the inferred user input for the first response being the third layer, the hypothetical response information 778 relating to the hypothetical response corresponding to the inferred user input being the fourth layer, and so on.
[0075] The assumed conversation information 770 is the result of prediction by the generating AI, and therefore has a so-called branching tree structure where it branches off from the previous layer to predictable candidates and connects to the next layer, and the amount of information in each layer tends to increase as the number of layers increases.
[0076] Figure 5 shows an example of the data structure for the first user input information 772, the first hypothetical response information 774, the inferred user input information 776, and the hypothetical response information 778.
[0077] Each of these pieces of information has an input branch ID and an output branch ID as a common data structure to indicate the branching relationship (position in the tree structure) in the tree structure of the assumed conversation information 770. However, the input branch ID of the first user input information 772 is a predetermined value indicating an end (0 in the example in Figure 5), and the output branch ID of the assumed response information 778 is initially set to an initial value indicating that it is an end in the tree structure (NULL in the example in Figure 5).
[0078] The first user input information 772 includes the input branch ID and output branch ID, and the input text entered by user 2.
[0079] The first hypothetical response information 774 includes an input branch ID and an output branch ID, a predicted probability of occurrence, importance, and a response representation information 775 generated by the generating AI.
[0080] The server-side generated AI20 and the terminal-side generated AI30 are implemented, at least in part, as GPT (Generative Pretrained Transformer) type LLM (Large Language Models). The generated data consists of content that is judged to have a high probability of occurrence based on the learning results, assuming the input data. The "predicted occurrence probability" is the probability of occurrence at the time of generation.
[0081] "Importance" is an indicator of the importance of the content of the given hypothetical response in the current conversation event, and can also be called a weighted value. Importance is set to an initial value that has a positive relationship with the predicted probability of occurrence, and is changed based on the conditions at the time the conversation event occurs.
[0082] The response control applied to the NPC5, the other character, uses multiple control items to achieve various expressions in conversation-based communication. These control items include dialogue text (including narration), the character's voice when reading the dialogue text, character motion, visual effects (VFX), and auditory effects (SE, BGM).
[0083] The response representation information 775 is data for executing control, prepared for each control item used in the response control in the given hypothetical response. For example, if the control item is spoken voice, it corresponds to voice data, and if the control item is motion, it corresponds to motion data.
[0084] The inferred user input information 776 includes the input branch ID and output branch ID, the predicted probability of appearance, the importance level, and the inferred input text predicted / inferred by the generating AI.
[0085] The assumed response information 778 includes the input branch ID and output branch ID, the predicted probability of occurrence, the importance level, and the response expression information 775.
[0086] Alternatively, the server-side generating AI 20 may output first user input information 772, first hypothetical response information 774, estimated user input information 776, and hypothetical response information 778. Or, the server system 1100 may create this information by separating the data generated by the server-side generating AI 20 and associating it with the input / output branch ID.
[0087] Figure 6 is a diagram illustrating the continuation of the forecast. When the server system 1100 generates assumed conversation information 770 as background processing, it decides whether to continue predicting subsequent assumed conversations based on the "importance" of each terminal assumed response information 778. Terminal assumed response information 778 is identified by the fact that the branching ID is a predetermined value indicating a terminal (for example, NULL) (see Figure 5).
[0088] Figure 7 is a diagram illustrating importance. When generating hypothetical response information 778, the server system 1100 sets the initial value of importance to have a positive relationship with its predicted occurrence probability. Then, the server system 1100 changes the importance from its initial value by referring to the importance change definition data 564, which is predefined in response to the conversation event that occurred.
[0089] The importance change definition data 564 stores the "change requirements regarding conversation situation information" and "the changes in importance when those requirements are met" that are provided to the server-side generation AI 20 along with the generation instruction information.
[0090] Figure 8 shows an example of the data structure of conversation situation information 730. Conversation situation information 730 is a collective term for information that collects various pieces of information describing the content at the time of generation instruction, user 2, and the conversation history exchanged in previous conversation events.
[0091] For example, conversation situation information 730 includes the opponent character ID 731, which indicates NPC5 as the opponent character, a copy of the opponent character's personality description text 732, and a copy of the opponent character's speech pattern description text 733.
[0092] The character description text 732 and the speech pattern description text 733 for the opponent character are copied from the character initial settings data 520, for example, as shown in Figure 9.
[0093] Character initial setup data 520 is prepared for each of the 5 NPCs. Each character initial setup data 520 includes a unique character ID 522, character model data 524, and personality information 530 for each of the 5 NPCs.
[0094] Personality information 530 includes basic setting information 531, personality description text 532, speech pattern description text 533, gesture characteristic description text 534, and character relationship information 535. Of course, other data may be included as appropriate.
[0095] Basic settings information 531 stores information such as the character's name, age, and race. Personality description text 532 is a natural language description of the character's personality. Speech style description text 533 is a natural language description of the character's speech style. Gesture characteristics description text 534 is a natural language description of the character's gestures, body language, and other characteristics.
[0096] Character relationship information 535 is various information that indicates the relationship between a character and other characters from the perspective of that character. For example, one character relationship information 535 stores the ID of the character it relates to and a relationship description text in association with each other. The relationship description text is a natural language explanation that describes the relationship between the character it relates to and the character it relates to (what the character thinks). For example, it might say, "I respect them very much and feel like using polite language with them," or "I neither like nor dislike them."
[0097] Returning to Figure 8, the conversation situation information 730 includes a copy of the prerequisite explanation text 734 for the event that occurred, a copy of the conversation purpose explanation text 735 for the event that occurred, a copy of the content situation log 736, and a copy of the character situation log 737.
[0098] The prerequisite explanation text 734 for the occurring event and the conversation purpose explanation text 735 for the occurring event are copied from conversation event definition data 550, for example, as shown in Figure 10.
[0099] A conversation event definition data 550 is prepared for each conversation event. Each conversation event definition data 550 includes an event ID 552, event occurrence requirements 554 that define what must be met for the event to occur, and partner character settings 556. Note that if the conversation event is for casual conversation without a particular purpose, the event occurrence requirement 554 may be set to "random occurrence".
[0100] The opponent character setting 556 is data that specifies NPC5, the character that will be the conversation partner of player character 4. For example, it may be the ID of a specific NPC, such as an NPC who works at a certain fruit shop, or a predetermined value that specifies that NPC5 located within a predetermined range from player character 4 when an event occurs will be the opponent character.
[0101] Furthermore, the conversation event definition data 550 includes a background explanation text 558, which is an explanatory text describing the outline of the conversation event, and a conversation purpose explanation text 560, which is an explanatory text describing the purpose of the conversation in the conversation event.
[0102] Furthermore, there may be conversation events for casual chats without a specific purpose. In that case, the conversation purpose explanation text 560 may include a predetermined value indicating "none," a time limit indicating the duration of the conversation event, and a maximum number of times the conversation between player character 4 and NPC 5 can be repeated.
[0103] Furthermore, the conversation event definition data 550 includes response restriction definition data 562, importance change definition data 564 (see Figure 7), and assumed response content change definition data 566. Of course, other information and data may also be included as appropriate.
[0104] The response restriction definition data 562 specifies the restrictions on responses to the conversation event. For example, it stores specified values such as the character limit for the response dialogue per response (single response character limit) and the number of responses. The single response character limit is set so that the number of characters is not exceeded when the response dialogue is read aloud.
[0105] The character limit per response may also be defined using a function that takes the values of various data included in the conversation situation information 730 (for example, the number of opposing characters) as variables. Furthermore, the response limit definition data 562 may also include a definition to randomly set the limit each time using random number generation.
[0106] The hypothetical response content change definition data 566 is data that treats the response in the conversation event as an opportunity and incorporates advice to give to user 2 in the conversation event, as well as the dramatic effects related to the response, into the content of the hypothetical response.
[0107] The hypothetical response content change definition data 566 includes conversation situation requirements, assumed conversation information requirements, and change instruction information. The conversation situation requirements indicate the conditions that conversation situation information 730 (see Figure 8) must satisfy in order to perform the change to the hypothetical response. The assumed conversation information requirements indicate the conditions that assumed conversation information 770 must satisfy in order to perform the change to the hypothetical response. The change instruction information is a prompt that describes in natural language how to change the content of the corresponding hypothetical response to the server-side generated AI 20. Alternatively, it may be data that indicates which wording of the content of the corresponding hypothetical response should be changed and how.
[0108] Returning to Figure 8, the conversation status information 730 includes a copy of the content status log 736, a copy of the character status log 737, past user input information 740, past response information 742, a copy of the user information 744, and a conversation objective fulfillment checklist 746.
[0109] The server system 1100 automatically stores content status logs and character status logs at a given timing or predetermined interval while providing content. The content status log copy 736 is a copy of the content status log for the most recent (or past predetermined time including the most recent) period, and the character status log copy 737 is a copy of the character status log for the most recent (or past predetermined time including the most recent) period.
[0110] Content status log copy 736 and the original content status log store various data describing the game's progress at the time of recording. For example, Content Status Log Copy 736 includes the play time at the time of recording (e.g., elapsed time since the start of play), the game world date and time, game progress, the game stage ID and event ID being played, background object data, etc. Of course, other data may also be included as appropriate.
[0111] Character Status Log 708 and the original Character Status are logs related to the content status, but they are character-specific logs prepared for each character appearing (4 player characters, 5 NPCs, and 6 enemy NPCs). For example, character status log copy 736 includes character ID, play time at the time of recording, location coordinates, character level, and ability parameter values (e.g., HP, MP, skill points, etc.). Other information such as status ailments, equipment information, and group affiliation name may also be included as appropriate.
[0112] Past user input information 740 includes a historical record of user input text since the start of the conversation event.
[0113] Past response information 742 contains a historical record of the dialogue text that NPC5 has responded to since the start of the conversation event.
[0114] User information copy 744 is a predetermined type of data associated with the account of user 2 managed by the server system 1100.
[0115] The Conversation Objective Fulfillment Checklist 746 is a checklist of keywords from the conversation objective description text 560 (see Figure 10) of the conversation event that occurred. When this list is created, the initial value for all keywords is initialized to a predetermined value indicating "not fulfilled".
[0116] Returning to Figure 7, the server system 1100 searches for any "change requirements regarding conversation situation information" in the importance change definition data 564 of the conversation event that occurred for each of the hypothetical response information 778 at the end of the assumed conversation information 770 (see Figures 5 and 6) that the conversation situation information 730 (see Figure 8) satisfies. Then, it applies the "importance change content" of the "change requirements regarding conversation situation information" that it satisfies to the importance of the hypothetical response information 778 and changes it.
[0117] Returning to Figure 6, the server system 1100 decides, based on the changed importance, whether to predict and infer further beyond the terminal assumed response information 778 in the tree structure.
[0118] The importance level will take values ranging from a minimum of "0" to a maximum of "1.0". In the example in Figure 6, the importance level was changed for each terminal hypothetical response information 778. As a result, for example, the importance level of hypothetical response information 778a is 0.75, hypothetical response information 778b is 0.6, hypothetical response information 778c is 0.2, and hypothetical response information 778d is 0.1.
[0119] The server system 1100 decides not to predict further assumed conversations for terminal hypothetical response information 778 (778c, 778d) whose importance does not reach a threshold value (e.g., 0.5).
[0120] On the other hand, the server system 1100 sets terminal hypothetical response information 778 (778a, 778b) whose importance has reached a threshold value (e.g., 0.5) as new branching points and decides to predict the subsequent hypothetical conversation, inheriting the flow of the hypothetical conversation up to that point. Predicting the hypothetical response to the predicted user input for further branching points is called "continuing the prediction".
[0121] For terminal hypothetical response information 778 (778a, 778b) whose importance has reached a threshold value (e.g., 0.5), the branch ID is set to the newly configured branch point ID. In the example in Figure 6, the predetermined value NULL is replaced with "TP52", which is the newly configured branch point ID.
[0122] The extent to which the prediction should continue is specified by the "number of repetitions," which is determined by considering user input and response as one unit of repetition and associating them with the hypothetical response information 778 where the importance has reached a threshold and a branching point has been set.
[0123] Figure 11 is a diagram illustrating an example configuration of the server-side generated AI20. The server-side generation AI 20 includes a general generation AI 21 and detailed generation AIs 23 (23a, 23b, ...) prepared for each control item.
[0124] Control items are control items in response control and correspond to the type of expression of the response. Examples of control items include dialogue text (including narration), spoken voice of a character reading the dialogue text, character motion, visual effects (VFX), and auditory effects (SE, BGM).
[0125] Therefore, the server-side generation AI 20 includes a text generation AI 23a that generates text for the first hypothetical response, inferred user input, and hypothetical response; a speech voice generation AI 23b that generates spoken voice; and a motion generation AI 23c that generates motion. It also includes a VFX generation AI 23d that generates VFX, and a SE generation AI 23e that generates SE and BGM.
[0126] The overall generation AI 21 determines the order in which to execute the control items for each response control of NPC 5, and which detailed generation AI 23 will be assigned to handle the generation (assigned generation AI). It then generates detailed generation instruction information for the assigned generation AI and provides it to each AI, causing them to generate response expression information 775 (775a, 775b, ...) based on their respective generated data. The overall generation AI 21 can be said to be in charge of supervision or production, directing the generation of response expression information 775 by the various detailed generation AI 23.
[0127] Furthermore, the overall generation AI 21 may set its assigned generation AI from among the detailed generation AI of the server-side generation AI 20 and the detailed generation AI of the terminal-side generation AI 30.
[0128] For example, the terminal-side generation AI 30 in this embodiment has a text generation AI 33a. Therefore, in this embodiment, the overall generation AI 21 can set the AI responsible for generating the response dialogue text to either the server-side text generation AI 23a or the terminal-side text generation AI 33a.
[0129] If the terminal-side generation AI 30 is configured to have multiple generation AIs for different control items, similar to the server-side generation AI 20, then similarly, a separate generation AI may be assigned to each control item, including control items other than dialogue text, such as spoken voice and motion.
[0130] Figure 12 is a diagram illustrating the machine learning of the integrated generative AI21. The integrated generation AI21 is pre-trained to take comprehensive generation instruction information prepared for training, including the text of the first user input, and annotations, and output a list of generation order control items for the first assumed response, as well as detailed generation instruction information for the corresponding generation AI.
[0131] Furthermore, the integrated generation AI21 is pre-trained to take the comprehensive generation instruction information and annotations prepared for training as input and output detailed generation instruction information to generate inferred input text for inferred user input.
[0132] Furthermore, the integrated generation AI21 is pre-trained to take the comprehensive generation instruction information and annotations prepared for training as input and output a list of generation order control items for each assumed response, as well as detailed generation instruction information for the corresponding generation AI.
[0133] Furthermore, the integrated generation AI supports generation instructions in natural language.
[0134] Figure 13 shows an example of a comprehensive generation instruction information 750, which is given by the server-side generation AI 20 to the overall generation AI 21, written in natural language. The comprehensive generation instruction information 750 includes a generation purpose description 751, a generation condition description 752, and reference information 753. Of course, other descriptions may also be included as appropriate.
[0135] The generation purpose description 751 and the generation condition description 752 are templates. The specification in generation purpose description 751 regarding the number of first hypothetical responses to create, the number of estimated user inputs, and the number of estimated hypothetical responses (for example, the part "up to the top 3 with predicted occurrence probabilities") is just an example and can be set as appropriate. For example, the specification may be set to create and estimate up to a predetermined upper limit number of items whose predicted occurrence probability is above a predetermined threshold value.
[0136] Reference information 753 includes a copy of conversation situation information 730 (753a), a copy of the conversation purpose description text 560 for the conversation event that occurred (753b), and a copy of the personality description text 532 for the other character (753c).
[0137] Furthermore, the reference information 753 includes a copy of the speech pattern description text 533 of the opponent character (753d), empty data that serves as a sample for the first hypothetical response information 774 (753e), and a copy of the input text for the first user input information 772 (753f).
[0138] Furthermore, the reference information 753 includes empty data (753g) that serves as a sample for the inferred user input information 776, empty data (753h) that serves as a sample for the assumed response information 778, and a copy (753j) of the response restriction definition data 562 for the conversation event that occurred. Of course, other data may be included in reference information 753 as appropriate.
[0139] Returning to Figure 12, the list of generation control items in the order of generation of the first hypothetical response and the list of generation control items in the order of generation of the hypothetical response show the control items applied to NPC5 in each single response control in the order of generation. If a control item includes dialogue text, the generation order of the dialogue text is set to 1st.
[0140] Candidate control items in the generation order control item list include dialogue text (including narration), spoken voice of a character reading the dialogue text, character motion, visual effects (VFX), and auditory effects (SE, BGM). Which of these control items is applied is determined by the overall generation AI 21 based on various information contained in the overall generation instruction information 750, such as the conversation situation, conversation purpose, the personality of the other character, and response restrictions.
[0141] For example, if NPC5 mutters without moving, the dialogue text and spoken audio are selected as control items. If the control items include dialogue text, the overall generation AI21 prioritizes the generation of the dialogue text as the first item. The second and subsequent items may be assigned in order of the control items that are expected to take the longest to generate.
[0142] The overall generation AI 21 may be composed of multiple AIs. For example, it may be configured as a first overall generation AI responsible for generating the list and detailed generation instruction information related to the first hypothetical response, and a second overall generation AI responsible for generating the list and detailed generation instruction information related to the hypothetical response.
[0143] Figure 14 is a diagram illustrating the machine learning of the text generation AI23a. The text generation AI23a is implemented, for example, as a Large Language Model (LLM) of the Generative Pretrained Transformer (GPT) type.
[0144] The text generation AI23a is pre-trained to take detailed generation instruction information prepared for training, the first user input text, and annotations as input, and output the dialogue text for the first hypothetical response.
[0145] Furthermore, the text generation AI23a is pre-trained to take detailed generation instruction information prepared for training, the dialogue text of the first assumed response, and annotations as input, and output the inferred input text of the inferred user input.
[0146] Furthermore, the text generation AI23a is pre-trained to take detailed generation instruction information prepared for training, as well as inferred user input text and annotations, and output hypothetical response dialogue text for the inferred user input.
[0147] Furthermore, the text generation AI23a supports generation instructions in natural language.
[0148] Figure 15 shows an example of detailed generation instruction information 762(762a) for text generation provided to the text generation AI 23a, which is written in natural language. The detailed generation instruction information 762(762a) includes a generation purpose description 751, a generation condition description 752, and reference information 753, similar to the general generation instruction information 750.
[0149] The reference information 753 of the detailed generation instruction information 762 (762a) includes a copy of the conversation situation information 730 (753a), a copy of the conversation purpose explanation text 560 of the conversation event that occurred (753b), and a copy of the personality explanation text 532 of the other character (753c).
[0150] Furthermore, the reference information 753 includes a copy of the speech pattern description text 533 of the opponent character (753d), and empty data (753e) that serves as a sample for the first hypothetical response information 774.
[0151] Furthermore, the reference information 753 includes a copy of the response restriction definition data 562 of the conversation event that occurred (753j), a copy of the dialogue text of the response expression information 755 of the first assumed response information 774 (753k), a copy of the inferred input text of the inferred user input information 776 (753m), and so on.
[0152] Figure 16 is a diagram illustrating the machine learning process of the speech generation AI 23b. The speech generation AI23b is pre-trained to take detailed speech generation instruction information and annotations prepared for training as input and output speech response expression information 775b, and it supports generation instructions in natural language.
[0153] Figure 17 shows an example of detailed generation instruction information 762b for speech generation, which is provided to the speech generation AI 23b, which is written in natural language. Detailed generation instruction information 762b for speech generation includes the generation purpose, generation conditions, and reference information. The reference information includes the dialogue text to be converted into speech and information about the personality and tone of voice of the other character.
[0154] The generation purpose and generation conditions are fixed phrases. The dialogue text to be converted into speech is the dialogue text of response expression information 775 of the first hypothetical response information 774 or hypothetical response information 778 (see Figure 5). The personality and tone of voice of the opposing character are copied from the reference information 753 of the comprehensive generation instruction information 750 (see Figure 12).
[0155] Figure 18 is a diagram illustrating the machine learning of motion generation AI23c. Motion generation AI23c is pre-trained to take detailed motion generation instruction information and annotations prepared for training as input and output motion data response representation information 775c, and it supports generation instructions in natural language.
[0156] Figure 19 shows an example of detailed motion generation instruction information 762c provided to the motion generation AI 23c, which is written in natural language. The detailed motion generation instruction information 762c includes the generation purpose, generation conditions, and reference information. The reference information includes the model data of the opponent character, the dialogue text to match the motion, the opponent character's personality, and the opponent character's gesture characteristics.
[0157] The generation purpose and generation conditions are fixed phrases. The dialogue text is a copy of the dialogue text from the response expression information 775 of the first hypothetical response information 774 or hypothetical response information 778 (see Figure 5). The model data of the opponent character is copied from the character model data 524 of the character initial setting data 520 (see Figure 9). For the opponent character's personality and gesture characteristics, the relevant information is copied from the reference information 753 of the comprehensive generation instruction information 750.
[0158] Figure 20 is a diagram illustrating the machine learning process of the VFX generation AI23d. The VFX generation AI 23d is pre-trained to take detailed generation instructions and annotations prepared for training as input and output VFX response representation information 775d, and it supports generation instructions in natural language.
[0159] Figure 21 shows an example of detailed generation instruction information 762d for VFX generation, which is provided to the VFX generation AI 23d written in natural language. Detailed generation instruction information 762d for VFX generation includes the generation purpose, generation conditions, and reference information. The reference information includes the model data of the opponent character, the dialogue text to match the motion, and the personality of the opponent character.
[0160] The generation purpose and generation conditions are standardized. For dialogue text, the dialogue text from response expression information 775 of the first hypothetical response information 774 or hypothetical response information 778 is copied. For the opponent character's model data, the character model data 524 from character initial setting data 520 is copied. For the opponent character's personality, the relevant information is copied from reference information 753 of comprehensive generation instruction information 750.
[0161] Figure 22 is a diagram illustrating the machine learning process of SE generation AI23e. The SE generation AI23e is pre-trained to take detailed generation instruction information and annotations prepared for training as input and output SE response representation information 775e, and it supports generation instructions in natural language.
[0162] Figure 23 shows an example of detailed generation instruction information 762e for SE generation, which is provided to the SE generation AI23e written in natural language. Detailed generation instruction information 762e for SE generation includes the generation purpose, generation conditions, and reference information. The reference information includes the model data of the opponent character, the dialogue text, and the personality of the opponent character.
[0163] The generation purpose and generation conditions are standardized. For dialogue text, the dialogue text from response expression information 775 of the first hypothetical response information 774 or hypothetical response information 778 is copied. For the opponent character's model data, the character model data 524 from character initial setting data 520 is copied. For the opponent character's personality, the relevant information is copied from reference information 753 of comprehensive generation instruction information 750.
[0164] Figure 24 shows an example of programs and data stored by the server system 1100. The server system 1100 stores the server program 501, the distribution client program 503, the trained AI model 508, user information 600 for each user 2, and play data 700 in the IC memory 1152. It also stores content initial setup data 510 in the database 1140. Of course, other data may be stored as appropriate.
[0165] The content initial setup data 510 may be stored in the IC memory 1152. The server program 501 may include a generation AI program 502 for realizing the functions of server-side generation AI 20 and detailed generation AI 23. Alternatively, the generation AI program 502 may be stored separately.
[0166] The server system 1100 executes and processes the server program 501 on the CPU 1151, thereby realizing the function of a server processing unit 200s, as shown in Figure 25.
[0167] The server processing unit 200s performs various controls related to content provision. Specifically, the server processing unit 200s includes a user information management unit 202, a billing control unit 204, a content provision control unit 206, a response time determination unit 208, an assumed conversation information preparation unit 210, a response control execution unit 220, and a generation unit 230.
[0168] The User Information Management Unit 202 executes the prescribed user registration procedure and controls the registration and management of user information 600 for each user 2.
[0169] The billing control unit 204 processes payments related to content. For example, it processes payments for playing content and for purchasing paid items usable in the content. Once the billing control unit 204 has processed a payment, it instructs the user information management unit 202 to update the billing information contained in the user information 600.
[0170] The content provision control unit 206 controls the progress of the content and provides it to the user terminal 1500. In this embodiment, the content is a video game, so the control unit will be responsible for controlling the progress of gameplay.
[0171] The response time determination unit 208 determines the response time length, which is the length of the response, as a limitation on the response. In this embodiment, this is determined by referring to the response limitation definition data 562 (see Figure 10) for each conversation event that occurs.
[0172] The anticipated conversation information preparation unit 210 prepares anticipated conversation information 770 by predicting future user inputs and generating hypothetical responses corresponding to those predicted user inputs, based on conversation situation information 730 which includes at least past user input information. The anticipated conversation information preparation unit 210 also has a hypothetical response generation / failure determination unit 212.
[0173] The hypothetical response generation determination unit 212 determines whether or not to generate a hypothetical response corresponding to the inferred user input based on the conversation situation information 730.
[0174] The response control execution unit 220 performs response control for a new response corresponding to a new user input using the assumed conversation information 770. A new response corresponding to a new user input means that if the first user input is a new user input, one of the first assumed responses selected becomes the new response. If the second user input to one of the response-controlled first assumed responses is a new user input, one of the assumed responses corresponding to an inferred user input that is the same as or similar in content to the second user input becomes the new response.
[0175] The generation unit 230 predicts the input to the closed question from user 2 and generates response representation information 775 (see Figure 11), which includes execution data representing the response of NPC 5. This is implemented by a so-called generative AI, which is the server-side generative AI 20.
[0176] The generation unit 230 includes a detailed generation unit 232 and a general generation unit 234. The detailed generation unit 232 is prepared for each control item of the NPC5's response expression and generates response expression information 775 for each control item of the response's control content. The detailed generation AI 23 corresponds to this. The overall generation unit 234 causes the detailed generation unit 232 to generate the response's control content based on the given generation instruction information. The overall generation AI 21 corresponds to this.
[0177] Returning to Figure 24, the distribution client program 503 is executed by the user terminal 1500, causing the user terminal 1500 to function as a man-machine interface for various control and content provision systems 1000 as a game client.
[0178] A pre-trained AI model 508 is prepared for each AI that makes up the server-side generation AI 20 (overall generation AI 21, detailed generation AI 23).
[0179] The content initial setup data 510 stores various initial setup data necessary for the execution of the content. For example, the content initial setup data 510 includes character initial setup data 520 (see Figure 9) and conversation event definition data 550 (see Figure 10).
[0180] User information 600 is prepared for each user 2 who has completed the prescribed registration procedure, and stores various information associated with that user 2. One user information 600 includes, for example, a user account, content save data, and billing information (for example, billing history associating billing date and time with billing amount, billing history statistics such as cumulative billing amount, number of billings, and billing frequency).
[0181] Play data 700 stores various information related to content provision and is updated by the content provision control unit 206. Play data 700 includes character control data 704 for each character such as player character 4 and NPC 5, original content status log 706, original character status log 708, and conversation management data 720. Of course, other data may also be included as appropriate.
[0182] The original content status log 706 and the original character status log 708 are the originals of content status log copy 736 and character status log copy 737, respectively.
[0183] The conversation management data 720 stores various data related to conversation events. The conversation management data 720 includes, for example, an event ID 722 indicating the conversation event that occurred, a partner character ID 724, conversation status information 730 (see Figure 8), and overall generation instruction information 750 (see Figure 13), as shown in Figure 26. The conversation management data 720 also includes detailed generation instruction information 762, assumed conversation information 770 (see Figures 4 and 5), and applicable input branch ID 780.
[0184] The applicable branch ID 780 indicates which branch point the conversation event has reached among the branch points set in relation to the assumed conversation information 770.
[0185] Figure 27 is a flowchart illustrating the flow of processing performed by the server system 1100 from the occurrence of a conversation event, and shows the control related to NPC 5 conversing with player character 4. The explanation of the control of player character 4 in response to user input has been omitted.
[0186] When a conversation event occurs during content delivery (YES in step S10), the server system 1100 sets NPC5 as the character to interact with (step S12). Then, it receives the first user input that initiates the conversation and creates and saves the first user input information 772 (step S14). The first user input can now be referenced at any time as past user input.
[0187] Next, the server system 1100 sets the outbound branch ID of the first user input information 772 to the applied inbound branch ID 780 (step S16).
[0188] Next, the server system 1100 initializes the conversation status information 730 (see Figure 8) (step S18). In the initial state, since no conversation has started, all information except past user input information 740 and past response information 742 is set. The conversation objective fulfillment checklist 746 shows that all keywords of the conversation objective of the conversation event that occurred are in the "unfulfilled" state.
[0189] Next, the server system 1100 sets conversation limits such as the conversation time limit (step S20), creates comprehensive generation instruction information 750 (step S22), and executes the assumed conversation information preparation process (step S24).
[0190] Figures 28 and 29 are flowcharts illustrating the flow of the process for preparing anticipated conversation information. As shown in Figure 28, in the assumed conversation information preparation process, the server system 1100 provides the comprehensive generation instruction information 750 to the server-side generation AI 20 (step S40), causing it to generate response expression information for the first assumed response (step S42).
[0191] In step S40, the overall generation AI 21 generates a list of generation order control items for the first hypothetical response, settings for the responsible generation AIs for each control item related to the first hypothetical response, and detailed generation instruction information 762 for each responsible generation AI for each control item (see Figure 12). Alternatively, the overall generation AI 21 may output the generation control item list and the responsible generation AIs, and the server system 1100 may generate the detailed generation instruction information 762 for the responsible generation AIs.
[0192] Furthermore, the central generation AI 21 provides each responsible generation AI with detailed generation instruction information 762 in the order indicated by the generation order control item list for the first hypothetical response, causing them to generate various response expression information 775 (see Figure 11). At this point, the response expression information 775a of the dialogue text for the first hypothetical response is also generated.
[0193] Then, the server system 1100 creates and stores a first hypothetical response information 774 containing the various response representation information 775 that have been generated (step S44).
[0194] Next, the server system 1100 generates inferred user input for each generated first hypothetical response (step S50). Specifically, the overall generation AI 21 provides the text generation AI 23a with detailed generation instruction information 762a for text and the dialogue text of the first hypothetical response to be processed, causing it to generate inferred input text for the inferred user input.
[0195] Next, the server system 1100 creates and saves the inferred user input information 776, which includes the generated inferred user input text (step S52).
[0196] Next, the server system 1100 generates response representation information 775 of the assumed response for each generated inferred user input (step S60). Specifically, the central generation AI 21 generates a list of control items for the generation order of hypothetical responses, settings for the generation AI responsible for each control item, and detailed generation instruction information 762 for each responsible generation AI, for each inferred user input. Furthermore, the central generation AI 21 provides the detailed generation instruction information 762 to each responsible generation AI in the generation order indicated by the generation order control item list, causing them to generate various response expression information 775 for the hypothetical responses to the inferred user input being processed. At this point, response expression information 775a of the dialogue text for the hypothetical responses to the inferred user input being processed is also generated.
[0197] Next, the server system 1100 creates and saves hypothetical response information 778 containing the various response representation information 775 that have been generated (step S62).
[0198] After executing steps S40 to S62, the first tier of hypothetical response information 774, the second tier of inferred user input information 776, and the third tier of hypothetical response information 778, which constitute the basic part of the assumed conversation information 770, will be prepared.
[0199] Next, the server system 1100 performs the predictive continuation setting process (steps S70 to S78). In the prediction continuation setting process, the server system 1100 determines and sets an initial value of importance for each terminal hypothetical response information 778 based on its predicted occurrence probability (step S70).
[0200] Next, the server system 1100 performs preparatory processing for changing importance based on the conversation situation information 730 (step S72). For example, it performs tasks such as determining user 2's play tendencies and determining whether user 2 is a high-spending user. The type of preparatory processing depends on the setting of the "change requirements regarding conversation situation information" in the importance change definition data 564 (see Figure 7).
[0201] Next, the server system 1100 changes the importance level of each of the terminal hypothetical response information 778 based on the importance change definition data 564 (step S74). Then, terminal hypothetical response information 778 whose changed importance level has reached a predetermined high importance threshold is considered to be worth predicting the subsequent assumed conversation, and an outbound branch ID is set as a new branching point (step S76).
[0202] Next, the server system 1100 sets the number of repetitions for each of the terminal hypothetical response information 778 that has been designated as a new branching point, according to its importance (step S78).
[0203] Next, the server system 1100 refers to the hypothetical response content change definition data 566 (see Figure 10) for the conversation event that occurred and changes the dialogue text of the generated hypothetical response information 778 (step S80).
[0204] Specifically, the server system 1100 searches for hypothetical response information 778 that satisfies two requirements: "change requirements regarding conversation situation information" and "requirements regarding assumed conversation information." The integrated generation AI 21 then provides the text generation AI 23a with change instruction information for the definition data that satisfies these two requirements, causing it to regenerate the response expression information 775a of the dialogue text of the retrieved hypothetical response information 778. Furthermore, the integrated generation AI 21 causes the detailed generation AI 23, which generates response expression information 775 other than the dialogue text, to regenerate each response expression information 775 based on the regenerated dialogue text. The response expression information 775 of the retrieved hypothetical response information 778 is then overwritten with the information thus regenerated.
[0205] Moving on to Figure 29, the server system 1100 then executes loop C for each assumed response of assumed response information 778 whose number of repetitions set in step S78 is not "0" (steps S100 to S120).
[0206] In loop C, the server system 1100 generates inferred user input for the assumed response to be processed (step S102), and creates and saves inferred user input information 776 for each generation (step S104).
[0207] Next, the server system 1100 compares the past user input information 740 of the conversation status information 730 with the conversation objective satisfaction checklist 746 of the conversation event that occurred to determine whether all conversation objectives have been satisfied for each generated inferred user input. Then, it selects the inferred user inputs that have not all been satisfied (step S105).
[0208] Next, the server system 1100 generates hypothetical response expression information 775 for each selected inferred user input in the generation order set by the overall generation AI 21 (step S106), and creates and saves hypothetical response information 778 for each generation (step S108). This creates a new terminal hypothetical response information 778, extending the branch of the tree structure.
[0209] Next, the server system 1100 initializes and modifies the importance level for each generated hypothetical response information 778 (step S110), sets the new terminal hypothetical response information 778 whose importance level is above a predetermined standard as a new branching point, and sets the branching ID (step S112).
[0210] Next, the server system 1100 determines whether all conversational objectives have been satisfied based on the past user input information 740 and past response information 742 of the conversation situation information 730, the inferred user input and hypothetical response generated in loop C (step S114).
[0211] If the condition is not met (NO in step S114), the server system 1100 subtracts "1" from the number of repetitions of the assumed response information 778 that is the target of processing in loop C (step S116), and terminates loop C (step S120).
[0212] If the condition is met (YES in step S114), the server system 1100 changes the number of repetitions of the assumed response information 778 that is the target of processing in loop C to "0" (step S116) and terminates loop C (step S120).
[0213] Next, if there are any hypothetical response information 778 remaining from the hypothetical response information 778 set in step S78 whose repetition count is not "0" (TRUE in step S122), the process returns to step S100. If there is nothing remaining (FALSE in step S122), the assumed conversation information 770 is considered complete, and the server system 1100 terminates the assumed conversation information preparation process.
[0214] Returning to Figure 27, the server system 1100 initiates response control to cause the other character NPC5 to provide a first response corresponding to the first user input (step S142). Specifically, the server system 1100 selects the first hypothetical response information 774, which has the highest importance among the assumed conversation information 770, and initiates response control to the other character NPC5 by applying its response expression information 775 (see Figure 5). As a result, the other character NPC5 provides a new response.
[0215] When response control is performed, the server system 1100 adds and updates the conversation status information 730 with new past response information 742 (step S144).
[0216] Next, the server system 1100 sets the outbound branch ID of the first hypothetical response information 774 selected in step S140 as the applicable inbound branch ID (step S146), and accepts new user input (step S152).
[0217] When new user input is received, the server system 1100 updates the conversation status information 730 by adding the new past user input information 740 (step S154).
[0218] The server system 1100 compares the past user input information 740 of the conversation status information 730 with the conversation objective satisfaction checklist 746 for the conversation event that occurred to confirm that the latest conversation objective has been satisfied. If, as a result, not all of the conversation objectives have been "satisfied" (NO in step S156), the server system 1100 executes the response continuation process (step S158).
[0219] Figures 30 and 31 are flowcharts illustrating the flow of the response continuation process. In the response continuation process, the server system 1100 first searches the assumed conversation information 770 for presumed user input information 776 that matches the new user input (step S200). For example, it searches for presumed user input information 776 that matches or is nearly the same as the new user input, or that has some words but indicates the same content. This can also be rephrased as searching for presumed user input information 776 that satisfies the matching conditions with the new user input.
[0220] If there is a matching guessed user input information 776 (YES in step S202), the server system 1100 searches for hypothetical response information 778 corresponding to the retrieved guessed user input information 776 (step S204). Specifically, it searches for hypothetical response information 778 in which the outbound branch ID of the guessed user input information 776 retrieved in step S200 matches the inbound branch ID.
[0221] Next, the server system 1100 selects the most important of the retrieved hypothetical response information 778 (step S206). Then, it starts executing response control on the NPC 5, the other character, by applying the response expression information 775 of the selected hypothetical response information 778 (step S208). As a result, the NPC 5 makes a new response to the new user input.
[0222] Since NPC5 has responded, the server system 1100 updates the conversation status information 730 by adding the new past response information 742 (step S230), and sets the outbound branch ID of the selected hypothetical response information 778 to the applied inbound branch ID 780 (step S232).
[0223] On the other hand, if the search in step S200 does not yield any matching presumed user input information 776 (NO in step S202), the process moves to Figure 31, and the server system 1100 starts response control to cause the other character NPC5 to give a predetermined specific response (step S210). As a specific response, the server system causes the other character to speak a line of dialogue requesting further user input (for example, "Excuse me. Could you please repeat that?").
[0224] Then, the server system 1100 accepts user input again (step S212).
[0225] While User 2 enters the same content as the most recent new user input, the server system 1100 discards the incorrect assumed conversation information 770 (step S214). Then, the server system 1100 prepares the assumed conversation information 770 again (step S216). Specifically, it treats the new user input as the first user input and performs the assumed conversation information preparation process (see Figures 28 to 29).
[0226] Next, the server system 1100 selects the first hypothetical response information 774 with the highest predicted occurrence probability from the re-prepared hypothetical conversation information 770 (step S218). Then, it starts response control applying the response expression information 775 of the first hypothetical response information 774 (step S220), and proceeds to step S230 in Figure 27.
[0227] Since NPC5 has responded, the server system 1100 updates the conversation status information 730 by adding the new past response information 742 (step S230), and sets the outbound branch ID of the selected hypothetical response information 778 to the applied inbound branch ID 780 (step S232).
[0228] Next, the server system 1100 determines whether the inferred user input information 776 corresponding to the inferred user input information 778 selected in step S206, and the inferred user input information 778 corresponding to said inferred user input information 776, are present in the assumed conversation information 770.
[0229] If the assumed conversation information 770 is re-prepared, it is determined whether the assumed conversation information 770 contains the inferred user input information 776 corresponding to the inferred user input information 778 selected in step S218, and the inferred user input information 778 corresponding to said inferred user input information 776.
[0230] If the determination is affirmative (YES in step S234), the server system 1100 assumes that the assumed conversation information 770 is still available and exits the response continuation process.
[0231] If the determination is negative (NO in step S234), the server system 1100 considers that it has run out of assumed conversation information 770. Then, it creates a comprehensive generation instruction information 750AD for preparation for additional assumed conversation information 770, starting from the assumed response controlled in step S208 (step S236; see Figure 32), and executes the additional preparation process (step S238). This can also be described as regenerating the assumed conversation information 770.
[0232] Figure 33 is a flowchart illustrating the flow of the additional preparation process. In the additional preparation process, the server system 1100 first provides the server-side generation AI 20 with the comprehensive generation instruction information 750AD for additional preparation (step S250), causing it to generate an estimated user input for the assumed response controlled in step S208 (step S252).
[0233] Specifically, the overall generation AI 21 sets the assigned generation AI and generates detailed generation instruction information 762a (see Figure 15) for text, treating the hypothetical response dialogue text controlled in step S208 as the "first hypothetical response dialogue text". This is then provided to the assigned generation AI to generate the inferred user input.
[0234] Next, the server system 1100 creates new inferred user input information 776 for each generated inferred user input and adds it to the assumed conversation information 770 (step S254). At this time, the input branch ID of the new inferred user input information 776 is initialized to the output branch ID of the assumed response information 778 of the assumed response controlled in step S208 (or step S220). The output branch ID of the new inferred user input information 776 is newly set.
[0235] Next, the server system 1100 generates response expression information 775 of a hypothetical response corresponding to the suggested user input for each generated inferred user input (step S256). Then, it creates hypothetical response information 778 for each generation and adds it to the assumed conversation information 770 (step S258).
[0236] Next, the server system 1100 performs a prediction continuation setting process for each of the created hypothetical response information 778 (step S260), and completes the additional preparation process. At this stage, the initially created assumed conversation information 770 has grown into a tree structure with branches that extend through the inferred user input information 776 generated in step S252, and then to the assumed response information 778 generated in step S256.
[0237] Once the server system 1100 has finished the additional preparation process, it returns to Figure 30 and terminates the response continuation process.
[0238] When the response continuation process ends, the system returns to Figure 27 and then to step S152, where it receives new user input regarding the response that was executed by the NPC5 character during the response continuation process.
[0239] If it is determined that all conversational objectives have been satisfied in step S156, that is, that the conversational objectives have been achieved (YES in step S156), the server system 1100 terminates the conversation (step S290).
[0240] As described above, according to this embodiment, control of communication between the user and the NPC can be realized using the generation unit.
[0241] In other words, the time required to generate an NPC response to user input may be longer than the response time for a conversation between natural people. However, according to this embodiment, predicted conversation information, including user input in a hypothetical conversation and a hypothetical response to it, can be prepared in advance. Therefore, by simply selecting a hypothetical response corresponding to new user input from the anticipated conversation information, it becomes possible to quickly control the response of NPC5.
[0242] [Variation] Although examples of embodiments to which the present invention is applied have been described above, the forms to which the present invention can be applied are not limited to the above forms, and components can be added, omitted, or modified as appropriate.
[0243] (Variation 1) For example, the content provision system 1000 may be implemented using a P2P (Peer to Peer) architecture with multiple user terminals 1500. In this case, programs and data corresponding to the functional division are stored in the user terminals 1500, and the functions corresponding to the server processing unit 200s in the above embodiment are distributed and implemented by the user terminals 1500, which act as P2P nodes. The same effects as in the above embodiment can be obtained with this configuration as well.
[0244] (Variation 2) Furthermore, in the above embodiment, the content provision system 1000 may be implemented not as a client-server type, but as a single computer system that was designated as the user terminal 1500 in the above embodiment.
[0245] Specifically, as shown in Figure 34, for example, the user terminal 1500 stores all the data stored by the server system 1100 in the above embodiment (see Figure 24). However, instead of the server program 501 and the distribution client program 503, a content provision program 505 is provided as an application program for the user terminal 1500.
[0246] In the above embodiment, the content provision program 505 implements all of the functional units of the server system 1100 (see Figure 25) as a terminal processing unit 200t on the user terminal 1500. In this modified example, the content provision program is executed on the user terminal 1500. The processing flow in the above embodiment (see Figures 27 to 31 and 33) can be interpreted by replacing the execution entity from the server system 1100 to the user terminal 1500.
[0247] (Variation 3) In the above embodiment, the assumed conversation information 770 was prepared in the assumed conversation information preparation process (see Figures 28 and 29) before the response control was performed in the first assumed response. However, the configuration may also be such that the response control is performed when the first assumed response information 774 is created in the assumed conversation information preparation process.
[0248] Specifically, as shown in Figure 35, the assumed conversation information preparation process in the above embodiment is divided into a first assumed conversation preparation process that generates a first assumed response, and a second assumed conversation preparation process that generates response expression information for the inferred user input and assumed response.
[0249] Then, instead of step S24, the server system 1100 performs the first hypothetical conversation preparation process (step S25). In the first hypothetical conversation process, the server system 1100 creates the first preparation comprehensive generation instruction information 750A as shown in Figure 36 and provides it to the server-side generation AI 20 to generate the response expression information 775 for the first hypothetical response, and creates and saves the first hypothetical response information 774.
[0250] Next, the server system 1100 selects the first assumed response information 774 with the highest appearance prediction probability among the created and saved first assumed response information 774, and starts response control for the first user input to the NPC 5 of the opponent character (step S142).
[0251] Then, the server system 1100 updates the conversation situation information 730 (step S144). When setting the application input branch ID 780 (step S146), it starts the second assumed conversation information preparation process (step S 147).
[0252] In the second assumed conversation information process, the server system creates the second comprehensive generation instruction information 750B as shown in FIG. 37 and gives it to the server-side generation AI 20 to generate the response expression information of the inferred user input and the assumed response.
[0253] According to this configuration, it is possible to execute the second assumed conversation information preparation process as background processing from the start of receiving a new user input until user 2 finishes the input.
Explanation of Signs
[0254] 20... Server-side generation AI 21... Overall generation AI 23... Detailed generation AI 30... Terminal-side generation AI 200s... Server processing unit 202... User information management unit 204... Billing control unit 206... Content provision control unit 208... Response time determination unit 210... Assumed conversation information preparation unit 212... Assumed response generation determination unit 220... Response control execution unit 230... Generation unit 232... Detailed generation unit 234... Overall generation unit 501... Server program 510... Content initial setting data 520...Character initial settings data 550...Conversation event definition data 556... Opponent character settings 560... Conversation purpose explanation text 562...Response limit definition data 564…Importance Change Definition Data 566... Defined data for changing the assumed response content 600... User Information 700... Play data 730...Conversation status information 731... Opponent's Character ID 740... Past user input information 742…Past response information 750... Comprehensive generation instruction information 755... Response representation information 762…Detailed generation instruction information 770...Expected conversation information 772...First user input information 774...First Assumption Response Information 775...Response representation information 776... Estimated user input information 778...Hypothetical response information 780...Applicable Branch ID
Claims
1. A computer system that performs conversational control to enable conversation with a user by repeatedly performing response control in response to user input, A means for preparing anticipated conversation information that prepares anticipated conversation information by predicting future user inputs and generating hypothetical responses corresponding to those predicted user inputs, based on conversation situation information including past user inputs. A response control execution means that performs response control for a new response in response to new user input using the assumed conversation information, A computer system equipped with the following features.
2. The aforementioned means for preparing anticipated conversation information is: A means for determining whether or not to generate the hypothetical response based on the conversation situation information, Having, The computer system according to claim 1.
3. The aforementioned means for preparing anticipated conversation information is: A means for determining whether or not to generate the hypothetical response based on the inference result of the inferred user input, Having, The computer system according to claim 1.
4. The aforementioned assumed conversation information preparation means predicts multiple predicted user inputs, The means for determining whether or not to generate a hypothetical response determines whether or not to generate a hypothetical response for each of the inferred user inputs. The computer system according to claim 2 or 3.
5. The assumed conversation information preparation means prepares assumed conversation information by performing an inference of the assumed user input in response to the assumed response and repeating the inference and generation of the assumed response, thereby advancing the exchange of the assumed conversation. The computer system according to claim 1.
6. The assumed conversation information preparation means prepares assumed conversation information for a tree-structured assumed conversation by performing multiple assumed user inputs estimation and generating hypothetical responses for each of the assumed user inputs estimation, and by repeating the estimation and generation of hypothetical responses by performing multiple assumed user inputs estimation for the hypothetical responses. The computer system according to claim 1.
7. The aforementioned assumed conversation information preparation means controls the number of layers for each branch destination in the tree-structured assumed conversation by controlling the number of repetitions. The computer system according to claim 6.
8. The aforementioned anticipated conversation information preparation means determines the importance of each branch destination and controls the number of repetitions based on that importance. The computer system according to claim 7.
9. The aforementioned anticipated conversation information preparation means controls the number of repetitions based on the user information of the user. The computer system according to claim 7.
10. A response time length determination means for determining the response time length, which is the length of the response. Furthermore, The aforementioned assumed conversation information preparation means generates the assumed response based on the response time length. The computer system according to claim 1.
11. Candidate control items for responses include the NPC's (Non-Player Character) dialogue, the NPC's speech, the NPC's motion, and the visual effects during the response. The response control execution means executes a response by controlling at least one of the candidate control items. The computer system according to claim 1.
12. The response includes control content for multiple candidate control items. The aforementioned assumed conversation information preparation means generates the assumed response by variably determining the generation order of the control item candidates. The computer system according to claim 11.
13. The aforementioned assumed conversation information preparation means generates the assumed response using a predetermined generation unit, The generation unit comprises a generation unit on the server system side and a generation unit on the user terminal side. The aforementioned assumed conversation information preparation means variably determines whether to have the server system's generation unit or the user terminal's generation unit generate the assumed response. The computer system according to claim 1.
14. The aforementioned assumed conversation information preparation means generates the assumed response using a predetermined generation unit, The generation unit comprises a detailed generation unit that generates response control content and a general generation unit that causes the detailed generation unit to generate response control content based on given generation instruction information. The aforementioned assumed conversation information preparation means generates the assumed response by providing the generation instruction information to the overall generation unit. The computer system according to claim 1.
15. The aforementioned means for preparing anticipated conversation information is: The process involves generating the hypothetical response using a predetermined generation unit, Determining the importance of each branch in the aforementioned tree-structured hypothetical conversation, Based on the aforementioned importance, the hypothetical response relating to the branch destination is regenerated using the generation unit, To do The computer system according to claim 6.
16. The response control execution means generates a new response if the new user input and the estimated user input do not satisfy the predetermined compatibility conditions, and executes response control using the response based on the assumed response as the new response if the new user input and the estimated user input satisfy the compatibility conditions. The computer system according to claim 1.
17. The response control execution means executes response control using a predetermined specific response as the new response if the new user input and the presumed user input do not satisfy the predetermined compatibility conditions, and executes response control using a response based on the assumed response as the new response if the new user input and the presumed user input satisfy the compatibility conditions. The computer system according to claim 1.
18. A program for causing a computer system to perform conversation control, which involves repeatedly performing response control in response to user input to enable conversation with the user, A hypothetical conversation information preparation means prepares hypothetical conversation information by predicting future user inputs and generating hypothetical responses corresponding to those predicted user inputs, based on conversation situation information including past user inputs. A response control execution means that performs response control for a new response in response to new user input using the assumed conversation information, A program for causing the aforementioned computer system to function.
Citation Information
Patent Citations
Game system and computer program used therefor
JP2010131082A