Virtual space management system and virtual space management method
The virtual space management system addresses the coordination issue between non-player character utterances and actions by using a learning model to generate synchronized speech and motion, enhancing user experience through natural and consistent behavior.
Patent Information
- Application Number
- PCT/JP2024/000585
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-12
- Publication Date
- 2025-07-17
AI Technical Summary
Existing virtual space management systems face issues with the coordination between the utterances and actions of non-player characters, leading to a sense of discomfort for users.
A virtual space management system that utilizes a first learning model to generate synchronized speech and motion text information for non-player characters, along with an instruction information generation unit to coordinate their actions, ensuring consistent and natural behavior.
The system achieves coordinated speech and motion of non-player characters, providing a more natural and harmonious user experience without a sense of incongruity.
Smart Images

Figure JP2024000585_17072025_PF_FP_ABST
Abstract
Description
Virtual space management system and virtual space management method
[0001] The present invention relates to a virtual space management system and a virtual space management method.
[0002] In a conventional virtual space provision service that provides users with access to a virtual space, attempts have been made to encourage user behavior in the virtual space by having a non-player character managed by the operator of the virtual space provision service speak to a user avatar operated by the user. For example, in Patent Document 1 below, a metaverse virtual three-dimensional space is generated by a metaverse management device. Within the virtual three-dimensional space, a player character exists that is controlled by the metaverse management device, which receives operation data generated by the user using a user terminal. A pseudo-player character control device is combined with the metaverse management device. The pseudo-player character control device sends a pseudo-player character into the virtual three-dimensional space created by the metaverse management device. The pseudo-player character control device generates pseudo-operation data in the same data format as the operation data and sends it to the metaverse management device to operate the pseudo-player character.
[0003] Patent No. 7235376
[0004] In many cases, the speech of a non-player character in a conversation between a user avatar and the non-player character is generated by an external dialogue AI (Artificial Intelligence) that is different from the platform system that manages the virtual space. Meanwhile, the actions of the non-player character are determined by a client device of the platform system. This poses a problem of insufficient coordination between the speech and actions of the non-player character, which can cause a sense of discomfort to the user.
[0005] The present invention aims to achieve natural behavior by coordinating the speech and actions of non-player characters.
[0006] A virtual space management system according to one embodiment of the present invention includes a text generation unit that generates speech text information indicating the speech of a first avatar and action text information indicating the action of the first avatar using a first learning model, an instruction information generation unit that generates instruction information that instructs the action of the first avatar based on the action text information, and a control information generation unit that generates an image of the first avatar moving based on the instruction information and generates first control information for outputting the image together with audio based on the speech text information.
[0007] A virtual space management method according to one aspect of the present invention generates speech text information indicating the speech of a first avatar and action text information indicating the actions of the first avatar using a first learning model, generates instruction information instructing the actions of the first avatar based on the action text information, generates an image in which the first avatar moves based on the instruction information, and generates control information for outputting the image together with audio based on the speech text information.
[0008] According to one aspect of the present invention, the speech and actions of non-player characters can be coordinated to realize natural behavior.
[0009] FIG. 1 is a block diagram showing the configuration of a virtual space management system 1 according to an embodiment. FIG. 2 is a schematic diagram showing an example of a display image VP of a virtual space VS. FIG. 3 is a block diagram showing the configuration of a base server 10. FIG. 4 is a block diagram showing the configuration of a text generation server 20. FIG. 5 is a block diagram showing the configuration of a gateway 30. FIG. 6 is a block diagram showing the configuration of a user terminal 40. FIG. 7 is a schematic diagram showing the flow of data in the virtual space management system 1. FIG. 8 is a schematic diagram showing the flow of data in a system 2 according to a comparative example.
[0010] A. Embodiment A-1. Overall Configuration FIG. 1 is a block diagram showing the configuration of a virtual space management system 1 according to an embodiment. The virtual space management system 1 includes an infrastructure server 10, a text generation server 20, a gateway 30, and a user terminal 40. The infrastructure server 10 is an example of a virtual space control device. The text generation server 20 is an example of a text generation device. The gateway 30 is an example of a connection device. The infrastructure server 10, the text generation server 20, and the gateway 30 are interconnected by a network N. Although one user terminal 40 is shown in FIG. 1, multiple user terminals 40 may be included in the virtual space management system 1.
[0011] The infrastructure server 10 manages the virtual space VS (see FIG. 2 ). The infrastructure server 10 stores, for example, three-dimensional space data that reproduces buildings, objects, or natural environments that constitute the virtual space VS, or three-dimensional avatar data that reproduces the appearance of an avatar. The infrastructure server 10 also controls the movement of an avatar in the virtual space VS based on instruction information. The instruction information is information that specifies a command that specifies the movement of the avatar. The infrastructure server 10 generates an image of the avatar that has been moved based on the instruction information, and transmits it to the user terminal 40 together with an image showing the background portion of the virtual space VS. Note that the functions of the infrastructure server 10 may be realized by multiple computers. In other words, the infrastructure server 10 may be referred to as a platform system for the virtual space VS.
[0012] In this embodiment, control information for outputting an image showing an avatar (particularly a character avatar AC) and the sound of the avatar is referred to as first control information, and control information for outputting an image showing the virtual space VS that serves as the background of the avatar is referred to as second control information. The first control information and the second control information may be transmitted to the user terminal 40 as a set of data. The second control information may also include sound data showing sounds (such as background sounds) in the virtual space VS.
[0013] The text generation server 20 generates text information for controlling the movement of a character avatar AC (see FIG. 2) using a first learning model LM1 (see FIG. 4). The character avatar AC is an example of a first avatar. The text generation server 20 generates character utterance text information indicating the speech of the character avatar AC and action text information indicating the actions of the character avatar AC. The first learning model LM1 is a learning model for realizing dialogue AI. More specifically, the first learning model LM1 outputs character utterance text information indicating a response to a dialogue partner (user avatar AU in this embodiment) based on the speech of the dialogue partner. The first learning model LM1 also outputs action text information indicating, in text, the content of an action that matches the content indicated by the character utterance text information.
[0014] The gateway 30 is disposed between the text generation server 20 and the infrastructure server 10, and connects the text generation server 20 and the infrastructure server 10. In this embodiment, the gateway 30 generates instruction information that instructs the action of the character avatar AC based on the action text information generated by the text generation server 20.
[0015] The user terminal 40 is a terminal used by the user U to access the virtual space VS. The user terminal 40 is, for example, an information processing terminal such as a smartphone, a tablet terminal, a personal computer, or a head-mounted display. In this embodiment, the user terminal 40 is assumed to be a smartphone.
[0016] The virtual space management system 1 is a system for providing a service using a virtual space VS to a user U. In this embodiment, a communication service that allows a user U to have various experiences by acting as an avatar in the virtual space VS will be described as an example of a service using the virtual space VS. The communication service (hereinafter simply referred to as the "communication service") provided by the virtual space management system 1 allows an avatar corresponding to the user U (hereinafter referred to as a "user avatar AU") to interact with avatars corresponding to other users U (hereinafter referred to as "other avatars"), and allows users to tour famous places that exist in a real space recreated within the virtual space VS. The user avatar AU is an example of a second avatar.
[0017] Furthermore, in the communication service, a character avatar AC accompanies the user avatar AU and supports his / her actions in the virtual space VS. The character avatar AC, for example, introduces events in the virtual space VS that match the preferences of the user U, and guides the user U to locations in the virtual space VS specified by the user U. The character avatar AC is an example of a non-player character.
[0018] FIG. 2 is a schematic diagram showing an example of a display image VP of the virtual space VS. The display image VP shown in FIG. 2 is displayed, for example, on the display 401 (see FIG. 6 ) of the user terminal 40. The display image VP depicts the virtual space VS including buildings and roads, a user avatar AU, and a character avatar AC. The user avatar AU behaves based on the operation of the user U on the user terminal 40. Specifically, the user avatar AU moves in a direction specified by the user U and utters utterances directly from the user U. Meanwhile, the character avatar AC behaves based on text information generated by the text generation server 20. Specifically, the character avatar AC moves in a direction specified by the action text information and utters utterances corresponding to the character utterance text information.
[0019] 2 shows the virtual space VS as seen from a field of view that includes the user avatar AU. However, the display image VP may be an image of the virtual space VS as seen from the line of sight of the user avatar AU. In this case, the display image VP does not include the user avatar AU, but includes the character avatar AC and the portion of the virtual space VS that is within the field of view of the user avatar AU.
[0020] A-2. Infrastructure Server 10 Fig. 3 is a block diagram showing the configuration of the infrastructure server 10. The infrastructure server 10 includes a communication device 101, a storage device 102, a processing device 103, and a bus 120 that interconnects these devices.
[0021] The communication device 101 communicates with other devices (the gateway 30 and the user terminal 40) using wireless communication or wired communication. In this embodiment, the communication device 101 has an interface connectable to a network N and communicates with other devices via the network N.
[0022] The storage device 102 is a recording medium readable by the processing device 103. The storage device 102 includes, for example, a nonvolatile memory and a volatile memory. The nonvolatile memory is, for example, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), and an electrically erasable programmable read-only memory (EEPROM). The volatile memory is, for example, a random access memory (RAM). The storage device 102 stores a program PG1. The program PG1 is a program for operating the infrastructure server 10.
[0023] The processing device 103 includes one or more central processing units (CPUs). The one or more CPUs are an example of one or more processors. Each of the processors and CPUs is an example of a computer. The processing device 103 reads a program PG1 from the storage device 102. The processing device 103 executes the program PG1 to function as a transmission / reception control unit 110 and a control information generation unit 111.
[0024] At least one of the transmission / reception control unit 110 and the control information generation unit 111 may be configured by a circuit such as a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), an FPGA (Field Programmable Gate Array), etc. Details of the transmission / reception control unit 110 and the control information generation unit 111 will be described later.
[0025] A-3. Text Generation Server 20 Fig. 4 is a block diagram showing the configuration of the text generation server 20. The text generation server 20 includes a communication device 201, a storage device 202, a processing device 203, and a bus 220 that interconnects these devices. The hardware description of the communication device 201, storage device 202, and processing device 203 is the same as that of the communication device 101, storage device 102, and processing device 103 of the infrastructure server 10.
[0026] The storage device 202 stores the program PG2. The program PG2 is a program for operating the text generation server 20. The processing device 203 reads the program PG2 from the storage device 202. The processing device 203 executes the program PG2 to function as a text generation unit 211. The storage device 202 also stores a first learning model LM1. The first learning model LM1 is used to generate text by the text generation unit 211. Details of the text generation unit 211 will be described later.
[0027] A-4. Gateway 30 Fig. 5 is a block diagram showing the configuration of the gateway 30. The gateway 30 includes a communication device 301, a storage device 302, a processing device 303, and a bus 320 that interconnects these devices. The hardware description of the communication device 301, storage device 302, and processing device 303 is the same as that of the communication device 101, storage device 102, and processing device 103 of the infrastructure server 10.
[0028] The storage device 302 stores a program PG3. The program PG3 is a program for operating the gateway 30. The processing device 303 reads the program PG3 from the storage device 302. The processing device 303 executes the program PG3 to function as an instruction information generation unit 311. Details of the instruction information generation unit 311 will be described later.
[0029] A-5. User Terminal 40 Fig. 6 is a block diagram showing the configuration of the user terminal 40. The user terminal 40 includes a display 401, a microphone 402, a speaker 403, an input device 404, a communication device 405, a storage device 406, a processing device 407, and a bus 420 that interconnects these devices.
[0030] The display 401 is a display device (for example, various display panels such as a liquid crystal display panel or an organic EL display panel) that displays information to the outside. Note that, if the user terminal 40 is a head-mounted display, a projection device that projects an image onto lenses placed in front of both eyes of the user U is provided instead of the display 401.
[0031] The microphone 402 collects sounds generated around the user terminal 40 and outputs them as audio data. In this embodiment, the sounds collected by the microphone 402 are mainly the speech of the user U. The speaker 403 outputs sounds that can be heard around the user terminal 40 based on the input audio data. In this embodiment, the sounds output by the speaker 403 are mainly the speech of the character avatar AC or another user U (other avatar).
[0032] The input device 404 is an input device (for example, a keyboard, a mouse, a switch, a button, or a sensor) that accepts an input from the outside. For example, the display 401 and the input device 404 may be integrated, like a touch panel.
[0033] The hardware description of the communication device 405, storage device 406, and processing device 407 is the same as that of the communication device 101, storage device 102, and processing device 103 of the infrastructure server 10. The storage device 406 stores program PG4. Program PG4 is a program for operating the user terminal 40. The processing device 407 reads program PG4 from the storage device 406. By executing program PG4, the processing device 407 functions as an input control unit 411 and an output control unit 412.
[0034] The input control unit 411 transmits information input from the input device 404 and the microphone 402 (hereinafter referred to as "input information") to the infrastructure server 10. The input information is, for example, information indicating the movement direction of the user avatar AU in the virtual space VS, the speech of the user U, or settings required for using the communication service.
[0035] The output control unit 412 displays an image on the display 401 and outputs sound from the speaker 403 based on the first control information and the second control information transmitted from the infrastructure server 10. The first control information and the second control information are, for example, images showing the character avatar AC, the other person's avatar, and the virtual space VS, or the speech of the character avatar AC and the other person's avatar.
[0036] A-4. Details of Each Functional Unit Next, the details of the transmission / reception control unit 110 and control information generation unit 111 of the infrastructure server 10, the text generation unit 211 of the text generation server 20, and the instruction information generation unit 311 of the gateway 30 will be described.
[0037] The transmission / reception control unit 110 of the infrastructure server 10 controls the transmission and reception of information between devices connected to the infrastructure server 10. For example, when a user avatar AU and a character avatar AC are conversing, the transmission / reception control unit 110 transcribes the user U's speech based on audio data indicating the user U's speech transmitted from the user terminal 40 to generate user utterance text information, and transmits the user utterance text information to the text generation server 20. Alternatively, audio data indicating the user U's speech (data before transcription) may be transmitted to the text generation server 20. In this embodiment, the transmission / reception control unit 110 transmits the user utterance text information to the text generation server 20. Furthermore, the transmission / reception control unit 110 transmits, for example, first control information and second control information generated by a control information generation unit 111 (described later) to the user terminal 40.
[0038] The text generation unit 211 of the text generation server 20 generates character utterance text information indicating the speech of the character avatar AC and action text information indicating the actions of the character avatar AC using the first learning model LM1. For example, the text generation unit 211 inputs user utterance text information transmitted from the infrastructure server 10 to the first learning model LM1. The first learning model LM1 determines the content of a response that matches the content of the user U's utterance based on the user utterance text information, and outputs it as character utterance text information. The first learning model LM1 also outputs the action of the character avatar AC that matches the content of the character utterance text information as action text information. The text generation unit 211 transmits a set of the character utterance text information and the action text information to the gateway 30.
[0039] Specific examples of character utterance text information and action text information are shown below. For example, if the user avatar AU waves to the character avatar AC and says "hello," the character utterance text information may be output as "hello," and the action text information may be output as "wave." For example, if the user avatar AU says to the character avatar AC, "There's an event in the plaza today," the character utterance text information may be output as "Great! Can I come with you?" and the action text information may be output as "Follow the user avatar AU."
[0040] That is, the text generation unit 211 generates character utterance text information corresponding to the content of the utterance of a user avatar AU different from the character avatar AC, and generates action text corresponding to the content of the character utterance text information. The user avatar AU is an example of an avatar different from the character avatar AC. The character utterance text information corresponding to the content of the utterance of the user avatar AU is character utterance text information that indicates a natural response to the utterance of the user avatar AU. Furthermore, the action information corresponding to the content of the character utterance text is text information that indicates a natural action to be performed while uttering the content of the character utterance text. The character utterance text information and the action text information may be content that reflects, for example, the nature (personality) of the character avatar AC and the relationship with the user avatar AU (for example, whether this is the first meeting, or whether they have a master-servant relationship like pets or an equal relationship like partners, etc.).
[0041] The instruction information generation unit 311 of the gateway 30 generates instruction information instructing the character avatar AC to perform an action based on the action text information. In this embodiment, the instruction information generation unit 311 selects at least one command corresponding to the action indicated by the action text information from among multiple commands that define the action of the character avatar AC. The instruction information includes information that identifies the at least one command selected by the instruction information generation unit 311.
[0042] The multiple commands that define the behavior of the character avatar AC are, for example, the same commands as the commands (hereinafter referred to as "designated commands") that the user U can specify when instructing the user avatar AU to perform certain actions. The designated commands may be interpreted by the control information generation unit 111 of the infrastructure server 10. For example, the designated commands that can be specified include "Command 1: move forward," "Command 2: move backward," "Command 3: move right," "Command 4: move left," "Command 5: jump," and "Command 6: wave."
[0043] If the specified commands include a command for executing the action indicated by the action text information, the instruction information generating unit 311 generates information identifying the command as instruction information. For example, if the action text information is "wave your hand," the instruction information generating unit 311 generates instruction information identifying "Command 6."
[0044] Furthermore, if the specified commands do not include a command for executing the action indicated by the action text information, the instruction information generation unit 311 generates instruction information by combining multiple commands to achieve the action. For example, if the action text information is "Follow character avatar AC," the instruction information generation unit 311 sequentially selects commands 1 to 4 in accordance with the movement of character avatar AC, and generates instruction information that specifies the selected command.
[0045] The control information generation unit 111 of the infrastructure server 10 generates an image of the character avatar AC moving based on instruction information and generates first control information for outputting the image together with sound based on character utterance text information. The control information generation unit 111 generates the first control information by performing movement control to control the movement of the character avatar AC based on the instruction information and voice synthesis processing to generate voice data reading the character utterance text information in the voice of the character avatar AC. The control information generation unit 111 also controls the movement of the user avatar AU based on instruction information from the user terminal 40 (instruction information based on operations performed by the user U on the user terminal 40). The control information generation unit 111 also performs acoustic processing on the user U's voice data transmitted from the user terminal 40, converting the user U's voice into the voice of the user avatar AU. The control information generation unit 111 may perform these processes using a game engine such as UNREAL ENGINE (registered trademark).
[0046] The control information generator 111 also generates an image (hereinafter referred to as a "background image") showing the virtual space VS around the user avatar AU and the character avatar AC, and generates second control information for outputting the background image. The background image reflects the actions of other avatars as necessary.
[0047] The first control information and second control information generated by the control information generator 111 are transmitted to the user terminal 40 by the transmission / reception controller 110 of the infrastructure server 10. Based on the first control information and the second control information, the user terminal 40 outputs an image of the character avatar AC moving in the background virtual space VS, together with audio indicating the speech of the character avatar AC.
[0048] 7 is a schematic diagram showing the flow of data in the virtual space management system 1. First, the user U operates the user terminal 40 to instruct the user avatar AU to speak and move (step S101). The operation of the user terminal 40 includes, for example, inputting spoken voice into the microphone 402 and instructing the user avatar AU to move using the input device 404 (for example, pressing a button to instruct a specific motion or pressing a button indicating a movement direction). The spoken voice input into the microphone 402 is converted into audio data. The instruction to move the user avatar AU using the input device 404 is converted into a command, and instruction information (information specifying the type of command) indicating the movement of the user avatar AU is generated.
[0049] The user terminal 40 transmits to the infrastructure server 10 audio data representing the speech input by the user U and instruction information representing the movement of the user avatar AU (step S102). The infrastructure server 10 generates an image of the user avatar AU moving based on the designation information representing the movement of the user avatar AU. The infrastructure server 10 also performs acoustic processing on the audio data as necessary to convert the voice of the user U into the voice of the user avatar AU. The infrastructure server 10 also generates a background image representing the virtual space VS around the user avatar AU and the character avatar AC. The generation of the background image is performed sequentially by the infrastructure server 10.
[0050] The infrastructure server 10 generates user utterance text information based on the voice data representing the uttered voice, and transmits the user utterance text information to the gateway 30 (step S103). The gateway 30 transmits the user utterance text information received from the infrastructure server 10 to the text generation server 20 (step S104).
[0051] Although not shown in the figures, the image in which the user avatar AU moves and the background image are also transmitted from the infrastructure server 10 to the user terminal 40 and displayed on the user terminal 40. In addition, as necessary, audio data indicating the voice of the user avatar AU is also transmitted to the user terminal 40 and output on the user terminal 40.
[0052] The text generation server 20, which has received the user utterance text information from the gateway 30, generates character utterance text information indicating the utterance of the character avatar AC based on the content of the utterance of the user avatar AU. The text generation server 20 also generates action text information indicating the action of the character avatar AC based on the content of the character utterance text information.
[0053] The text generation server 20 transmits the character utterance text information and the action text information to the gateway 30 (step S105). The gateway 30 selects a command for realizing the action of the character avatar AC indicated by the action text information, and generates instruction information that identifies the command. The gateway 30 transmits the instruction information and the character utterance text information to the infrastructure server 10 (step S106).
[0054] The infrastructure server 10 generates an image of the character avatar AC moving, based on the designation information indicating the movement of the character avatar AC. The infrastructure server 10 also generates audio data indicating the voice of the character avatar AC by reading aloud the text indicated by the character utterance text information using a voice synthesis program. The infrastructure server 10 transmits image information indicating the image of the character avatar AC moving and audio data indicating the voice of the character avatar AC to the user terminal 40 as first control information. The infrastructure server 10 also transmits image information indicating a background image around the user avatar AU and the character avatar AC to the user terminal 40 as second control information (step S107).
[0055] Based on the second control information, user terminal 40 displays a background image on display 401. Based on the first control information, user terminal 40 also displays an image operated by character avatar AC on display 401, and outputs the speech of character avatar AC from speaker 403 (step S108).
[0056] FIG. 8 is a schematic diagram showing the flow of data in a system 2 according to a comparative example. The system 2 illustrated in FIG. 8 includes a platform server 10A, a text generation server 20A, and a user terminal 40A. Steps S201 and S202 are similar to steps S101 and S102 in FIG. 7. That is, the user U operates the user terminal 40A to instruct the user avatar AU to speak and move (step S201). The user terminal 40A transmits audio data representing the speech input by the user U and instruction information representing the movement of the user avatar AU to the platform server 10 (step S202). The platform server 10A generates an image in which the user avatar AU moves based on the designation information representing the movement of the user avatar AU. The platform server 10A also performs acoustic processing on the audio data as necessary to convert the user U's voice into the voice of the user avatar AU. The platform server 10A also generates a background image.
[0057] The infrastructure server 10A converts the user U's speech into user utterance text information indicating the speech of the user avatar AU and transmits it to the text generation server 20A (step S203). The text generation server 20A generates character utterance text information indicating the speech of the character avatar AC based on the user utterance text information. The text generation server 20A transmits the character utterance text information to the infrastructure server 10A (step S204).
[0058] Upon receiving the character utterance text information, the infrastructure server 10A uses a voice synthesis program to read out the text indicated by the character utterance text information, thereby generating audio data representing the voice of the character avatar AC. The infrastructure server 10A also determines the movements of the character avatar AC. The movements of the character avatar AC are determined, for example, using a learning model for determining movements. The infrastructure server 10A selects a combination of commands to realize the movements of the character avatar AC, and generates instruction information. The infrastructure server 10A also generates an image of the character avatar AC moving based on the instruction information.
[0059] The infrastructure server 10A transmits image information showing an image of the character avatar AC in action and audio data showing the speech of the character avatar AC to the user terminal 40A as first control information. The infrastructure server 10A also transmits image information showing a background image around the user avatar AU and the character avatar AC to the user terminal 40A as second control information (step S205).
[0060] Based on the second control information, user terminal 40A displays a background image on display 401. Based on the first control information, user terminal 40A also displays an image operated by character avatar AC on display 401 and outputs the speech of character avatar AC from speaker 403 (step S206).
[0061] A-5. Summary of the embodiment As described above, in the virtual space management system 1 according to the embodiment, the speech and actions of the character avatar AC are determined by the same learning model (first learning model LM1). Therefore, compared to the system 2 shown in FIG. 8, in which the speech and actions of the character avatar AC are determined by different devices (different learning models), the consistency between the speech and actions is improved, enabling more natural behavior.
[0062] Furthermore, in the virtual space management system 1, the content of the movements of the character avatar AC is determined by text, which allows for greater freedom of movement and improves the naturalness of the movements, compared to when the content of the movements of the character avatar AC is determined by a combination of commands, for example.
[0063] Furthermore, in the virtual space management system 1, unlike the system 2, the actions of the character avatar AC are performed by an external device (text generation server 20) separate from the infrastructure server 10. In other words, in the virtual space management system 1, the dialogue AI is an external element separate from the infrastructure system of the virtual space VS. Therefore, the components that make up the virtual space management system 1 (specifically, the dialogue AI and the infrastructure system) are loosely coupled, allowing for the construction of systems that are highly independent of each other. For example, update processes for the dialogue AI and the infrastructure system can be performed at independent times. In this case, the update process can be performed by anyone with knowledge of the system to be updated, thereby increasing the flexibility of system operation.
[0064] B: Modifications Modifications of the above-described embodiment are shown below. Two or more of the following modifications may be arbitrarily selected and combined as long as they are not mutually contradictory.
[0065] B1: First Modification In the embodiment described above, the virtual space management system 1 included a base server 10, a text generation server 20, a gateway 30, and a user terminal 40. This is not limiting, and for example, the functions of the gateway 30 may be realized by the text generation server 20. While it is possible to include the functions of the gateway 30 in the base server 10, it is more preferable to provide the functions of the gateway 30 in the text generation server 20 from the viewpoint of separating the function of motion control and the function of action determination.
[0066] B2: Second Modification In the above-described embodiment, the infrastructure server 10 transmits user utterance text information to the text generation server 20. However, this is not limiting, and for example, voice data indicating the user U's utterance may be transmitted to the text generation server 20. In this case, the text generation server 20 converts the user U's utterance into user utterance text information. Furthermore, the text generation server 20 may generate character utterance text information and action text information taking into account the intonation or tone of the user U's utterance.
[0067] The infrastructure server 10 may also transmit to the text generation server 20 at least one of information specifying the movement of the user avatar AU (e.g., the direction and speed of movement when the user avatar AU is moving, or the type of movement when the user avatar AU is making a motion) or a background image. In this case, the text generation server 20 may determine the content of the speech or movement of the character avatar AC based on at least one of the movement of the user avatar AU or the background image. For example, the text generation server 20 may cause the character avatar AC to perform an action or speak in accordance with the movement of the user avatar AU. The text generation server 20 may also detect the state of the virtual space VS around the character avatar AC and the user avatar AU based on the background image, and cause the character avatar AC to perform an action or speak in accordance with the state.
[0068] C: Others (1) In the above-described embodiments, ROM, RAM, etc. are given as examples of storage devices 102, 202, 302, and 406. However, storage devices 102, 202, 302, and 406 may also be flexible disks, magneto-optical disks (e.g., compact disks, digital versatile disks, Blu-ray (registered trademark) disks), smart cards, flash memory devices (e.g., cards, sticks, key drives), CD-ROMs (Compact Disc-ROMs), registers, removable disks, hard disks, floppy (registered trademark) disks, magnetic strips, databases, servers, or other suitable storage media.
[0069] (2) In the above-described embodiments, the described information, signals, etc. may be represented using any of a variety of different technologies. For example, data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be referred to throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.
[0070] (3) In the above-described embodiment, input and output information may be stored in a specific location (for example, a memory) or may be managed using a management table. Input and output information may be overwritten, updated, or added to. Output information may be deleted. Input information may be transmitted to another device.
[0071] (4) In the above-described embodiment, the determination may be made based on a value represented by one bit (0 or 1), a Boolean value (true or false), or a comparison of numerical values (e.g., a comparison with a predetermined value).
[0072] (5) The order of the process procedures, sequences, flowcharts, etc. illustrated in the above-described embodiments may be rearranged unless it is consistent. For example, the methods described in this disclosure present elements of various steps using an example order, and are not limited to the particular order presented.
[0073] (6) Each function illustrated in Figures 3 to 6 is realized by any combination of hardware and / or software. Furthermore, the method of realizing each functional block is not particularly limited. That is, each functional block may be realized using a single device that is physically or logically coupled, or may be realized using two or more physically or logically separated devices that are directly or indirectly connected (for example, using wires, wirelessly, etc.) and these multiple devices. A functional block may be realized by combining software with the single device or the multiple devices.
[0074] (7) The programs exemplified in the above-described embodiments should be broadly construed to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, etc., regardless of whether they are called software, firmware, middleware, microcode, hardware description language, or by other names.
[0075] Software, instructions, information, etc. may also be transmitted or received over a transmission medium. For example, if software is transmitted from a website, server, or other remote source using wired technologies (such as coaxial cable, fiber optic cable, twisted pair, Digital Subscriber Line (DSL)), and / or wireless technologies (such as infrared, microwave), then these wired and / or wireless technologies are included within the definition of transmission media.
[0076] (8) In each of the foregoing embodiments, the terms "system" and "network" are used interchangeably.
[0077] (9) The information, parameters, etc. described in this disclosure may be expressed using absolute values, relative values from a predetermined value, or corresponding other information.
[0078] (10) In the above-described embodiments, the portable device may be a mobile station (MS). Those skilled in the art may also refer to a mobile station as a subscriber station, mobile unit, subscriber unit, wireless unit, remote unit, mobile device, wireless device, wireless communication device, remote device, mobile subscriber station, access terminal, mobile terminal, wireless terminal, remote terminal, handset, user agent, mobile client, client, or some other appropriate terminology. In this disclosure, the terms "mobile station," "user terminal," "user equipment (UE)," "terminal," etc. may be used interchangeably.
[0079] (11) In the above-described embodiments, the terms "connected," "coupled," or any variation thereof refers to any direct or indirect connection or coupling between two or more elements, and may include the presence of one or more intermediate elements between two elements that are "connected" or "coupled" to each other. The coupling or connection between elements may be physical, logical, or a combination thereof. For example, "connected" may be read as "access." As used in this disclosure, two elements may be considered to be "connected" or "coupled" to each other using at least one of one or more wires, cables, and printed electrical connections, as well as electromagnetic energy having wavelengths in the radio frequency range, microwave range, and optical (both visible and invisible) range, as some non-limiting and non-exhaustive examples.
[0080] (12) In the above embodiments, the phrase "based on" does not mean "based only on," unless otherwise specified. In other words, the phrase "based on" means both "based only on" and "based at least on."
[0081] (13) As used in this disclosure, the terms "determining" and "determining" may encompass a wide variety of actions. "Determining" and "determining" may include, for example, judging, calculating, computing, processing, deriving, investigating, looking up, searching, inquiring (e.g., searching in a table, database, or other data structure), ascertaining something as "determining" or "determining," and the like. Furthermore, "judgment" and "decision" may include regarding receiving (e.g., receiving information), transmitting (e.g., sending information), input, output, and accessing (e.g., accessing data in memory) as having been "judgment" or "decision." Furthermore, "judgment" and "decision" may include regarding resolving, selecting, choosing, establishing, comparing, etc. as having been "judgment" or "decision." In other words, "judgment" and "decision" may include regarding some action as having been "judgment" or "decision." Furthermore, "judgment" may be interpreted as "assuming," "expecting," "considering," etc.
[0082] (14) In the above embodiments, when "include," "including," and variations thereof are used, these terms are intended to be inclusive, similar to the term "comprising." Furthermore, the term "or," as used in this disclosure, is not intended to be an exclusive or.
[0083] (15) In this disclosure, when articles are added by translation, such as a, an, and the in English, the disclosure may include the noun following these articles being plural.
[0084] (16) In this disclosure, the term "A and B are different" may mean "A and B are different from each other." The term may also mean "A and B are each different from C." Terms such as "separate" and "combined" may also be interpreted in the same way as "different."
[0085] (17) The aspects / embodiments described in this disclosure may be used alone, in combination, or switched depending on the implementation. Notification of predetermined information (e.g., notification that "X is true") is not limited to being explicit, but may be implicit (e.g., not notifying the predetermined information).
[0086] 1...Virtual space management system, 10...Infrastructure server, 20...Text generation server, 30...Gateway, 40...User terminal, 103...Processing device, 110...Transmission / reception control unit, 111...Control information generation unit, 203...Processing device, 211...Text generation unit, 303...Processing device, 311...Instruction information generation unit, 407...Processing device, 411...Input control unit, 412...Output control unit, AC...Character avatar, AU...User avatar, LM1...First learning model, N...Network, U...User, VS...Virtual space.
Claims
1. A virtual space management system comprising: a text generation unit that generates speech text information indicating the speech of a first avatar and motion text information indicating the motion of the first avatar using a first learning model; an instruction information generation unit that generates instruction information for instructing the motion of the first avatar based on the motion text information; and a control information generation unit that generates first control information for generating an image in which the first avatar moves based on the instruction information and outputting the image together with speech based on the speech text information.
2. The virtual space management system according to claim 1, wherein the instruction information generation unit selects at least one command corresponding to the motion indicated by the motion text information from a plurality of commands that define the motion of the first avatar, and the instruction information includes information for specifying the at least one selected command.
3. The virtual space management system according to claim 1, wherein the text generation unit generates the speech text information corresponding to at least one of the content of the speech or the motion of a second avatar different from the first avatar, and generates the motion text information corresponding to the content of the speech text information.
4. The virtual space management system according to claim 1, wherein the text generation unit is provided in a text generation device that generates text information using the first learning model, the control information generation unit is provided in a virtual space control device that generates second control information for outputting an image indicating a virtual space, and the instruction information generation unit is provided in a connection device that connects the text generation device and the virtual space control device.
5. A virtual space management method comprising: generating speech text information indicating the speech of a first avatar and motion text information indicating the motion of the first avatar using a first learning model; generating instruction information for instructing the motion of the first avatar based on the motion text information; and generating control information for generating an image in which the first avatar moves based on the instruction information and outputting the image together with speech based on the speech text information.
Citation Information
Patent Citations
Method and program for creating and / or editing work
JP2023023978A
Animating body language for avatars
US20210158589A1