Information processing system, information processing method, and program

The information processing system enhances avatar guidance in virtual spaces by using an AI model to generate contextually appropriate actions based on user and space data, addressing limitations of stereotypical support in conventional systems.

JP7861975B1Active Publication Date: 2026-05-19CLUSTER INC
View PDF 10 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
CLUSTER INC
Filing Date
2025-07-18
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Conventional technologies for assisting avatar movement in virtual spaces provide limited stereotypical guidance and lack flexibility in adapting to user information and virtual space dynamics.

Method used

An information processing system that includes a first acquisition unit for dialogue and virtual space information, a generation unit to create prompts for an AI model, and an output unit to control agent avatars based on user and space data, enabling flexible support through an AI model like a large language model.

Benefits of technology

Enables flexible and contextually appropriate guidance and interaction by agent avatars in virtual spaces, adapting to user inputs and space dynamics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007861975000001_ABST
    Figure 0007861975000001_ABST
Patent Text Reader

Abstract

To enable flexible support from agent avatars based on user information and virtual space information. [Solution] An information processing system comprising: a first acquisition unit that acquires dialogue information relating to an interaction with an avatar in a virtual space identified from a plurality of virtual spaces, and virtual space information relating to the virtual space; a generation unit that generates a prompt for input to an AI model from at least the dialogue information and the virtual space information; a second acquisition unit that acquires control information relating to the actions of an agent interacting with the avatar in the virtual space by inputting the prompt to the AI ​​model; and an output unit that outputs the control information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing system, an information processing method, and a program.

Background Art

[0002] Conventionally, a technology for assisting the operation of an avatar operated by a user in a virtual space is known. For example, Patent Document 1 describes a technology for generating a specific object that enables an avatar to move to a specific position or a specific area in a virtual space.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] The technology described in Patent Document 1 further utilizes an agent avatar that performs various guidance processes in a virtual space to assist the movement of the avatar to a specific position or a specific area. However, in the conventional technology, the user support by the agent avatar is limited to stereotypical movement support and guidance support, and it has been difficult to provide flexible guidance and dialogue corresponding to user information and virtual space information.

[0005] The present invention has been made in consideration of such circumstances, and one of the objectives is to provide an information processing system, an information processing method, and a program that enable flexible support by an agent avatar according to user information and virtual space information.

Means for Solving the Problems

[0006] Embodiments of the present invention are information processing systems comprising: a first acquisition unit that acquires dialogue information relating to an interaction with an avatar in a virtual space identified from a plurality of virtual spaces, and virtual space information relating to the virtual space; a generation unit that generates a prompt for input to an AI model from at least the dialogue information and the virtual space information; a second acquisition unit that acquires control information relating to the actions of an agent interacting with the avatar in the virtual space by inputting the prompt to the AI ​​model; and an output unit that outputs the control information. [Effects of the Invention]

[0007] According to this embodiment, flexible support by agent avatars can be enabled according to user information and virtual space information. [Brief explanation of the drawing]

[0008] [Figure 1] This is a diagram illustrating the overview of the information processing system 1 according to the embodiment. [Figure 2] This figure shows an example of the configuration of the information processing system 1 according to the embodiment. [Figure 3] This figure shows an example of the structure of user data 58A. [Figure 4] This figure shows an example of the configuration of virtual space mesh data 58B. [Figure 5] This figure shows an example of the configuration of virtual space static data 58C. [Figure 6] This figure shows an example of the configuration of virtual space dynamic data 58D. [Figure 7] This figure shows an example of the structure of user history data 58E. [Figure 8] This figure shows an example of interaction with an agent avatar (AA) in a virtual space. [Figure 9] This figure shows an example of a prompt PR generated by the generation unit 130. [Figure 10]This figure shows an example of the behavior of an agent avatar (AA) controlled based on control information. [Figure 11] This figure shows another example of the behavior of an agent avatar (AA) controlled based on control information. [Figure 12] This is a sequence diagram illustrating the processing flow performed through the cooperation of the terminal device 10, the virtual reality service server 50, and the information processing device 100. [Figure 13] This flowchart shows an example of controlling the transmission of dialogue information and virtual space information by the virtual reality service server 50. [Figure 14] This flowchart shows another example of the control of the transmission of dialogue information and virtual space information by the virtual reality service server 50. [Modes for carrying out the invention]

[0009] [overview] Hereinafter, an information processing system 1 according to an embodiment of the present invention will be described with reference to the drawings. Figure 1 is a diagram illustrating the overview of the information processing system 1 according to an embodiment. The information processing system 1 includes, for example, one or more terminal devices 10, a virtual reality (VR) service server 50, and an information processing device 100.

[0010] One or more terminal devices 10 are devices used by a user, such as a mobile phone, smartphone, PC (Personal Computer), tablet device, HMD (Head Mounted Display), or game device. One or more terminal devices 10 are equipped with a virtual reality application 20A and operate in cooperation with a virtual reality service server 50 to provide virtual reality services to the user. In another embodiment, the terminal devices 10 may access the virtual reality service server 50 via a web browser without being equipped with a virtual reality application 20A and receive virtual reality services.

[0011] In this embodiment, the virtual reality service is a metaverse platform service in which, for example, each user can gather as an avatar in a virtual space provided by the virtual reality service server 50, or in a virtual space they themselves have created (using tools and parts provided by the virtual reality service), from various device environments, and enjoy interaction and events. The user may be an individual or a corporation operating a virtual store on the virtual reality service. Hereinafter, the avatar representing the user may be referred to as the user avatar UA. The user avatar UA is an avatar that appears in the virtual space from the user's own perspective.

[0012] In this service, users can enter (log in) to a virtual space created by themselves or other users as a user avatar (UA) and explore the entered virtual space. To ensure that users can fully enjoy the virtual space, the virtual reality service server 50 provides an agent avatar (AA) that guides the user avatar (UA) to predetermined locations (e.g., coordinates in the virtual space) or predetermined objects (e.g., polygon objects placed in the virtual space) and provides explanations about those locations and objects. The agent avatar (AA) is a supportive character that autonomously interacts with, guides, and explains to the user. Users can interact with the agent avatar (AA) or have it provide guidance and explanations by inputting text or voice into the terminal device 10.

[0013] As shown in FIG. 1, the user starts the virtual reality application 20A on the terminal device 10, enters a certain virtual space, and interacts with the agent avatar AA by inputting information such as text and voice into the terminal device 10 (S1). At this time, the user may start a web browser on the terminal device 10 to access the virtual reality service and enter a certain virtual space (hereinafter, the virtual reality application 20A may be read as a web browser). The terminal device 10 transmits the input information as interaction information to the virtual reality service server 50 (S2). When the virtual reality service server 50 receives the input information from the user, it identifies the virtual space in which the user has entered from among the plurality of virtual spaces managed by the virtual reality service server 50 (S3). As described above, in this virtual reality service, each user can generate a virtual space, so the virtual reality service server 50 needs to identify the virtual space in which the user has entered.

[0014] When the virtual reality service server 50 identifies the virtual space, it transmits the interaction information input by the user and the virtual space information related to the identified virtual space to the information processing device 100 (S4). When the information processing device 100 receives the interaction information and the virtual space information, it generates a prompt from the interaction information and the virtual space information, and inputs it into the learned AI model 162 (for example, a learned large language model (LLM: Large Language Model)) to obtain control information regarding the operation of the agent avatar AA in the virtual space (S5). Here, the prompt requests the AI model 162 to output control information regarding the operation to be executed by the agent avatar AA on the premise of the interaction information input by the user and the virtual space information regarding the virtual space entered by the user. In S4, the virtual reality service server 50 transmits only the identification information (virtual space ID) of the virtual space to the information processing device 100, and the AI model 162 may obtain the virtual space information from the virtual reality service server 50 via an MCP (Model Context Protocol) protocol server or the like based on the received identification information.

[0015] When the information processing device 100 acquires control information from the AI model 162, it outputs the acquired control information to the virtual reality service server 50. When the virtual reality service server 50 receives the control information from the information processing device 100, based on the control information, it controls the operation of the agent avatar AA in the virtual space (S7) and causes it to be displayed on the display device of the terminal device 10 (S8). As a result, the user can confirm the operation of the agent avatar AA and move the user avatar UA to a predetermined location or object, or receive an explanation about a predetermined location or object.

[0016] Furthermore, in the present embodiment, since a plurality of users can enter different virtual spaces, by specifying the target virtual space from among the plurality of virtual spaces and inputting information about the specified virtual space, appropriate control information for the agent avatar AA can be acquired according to each user and virtual space. That is, flexible support by the agent avatar can be enabled according to user information and virtual space information. Hereinafter, the configuration and operation of the information processing system 1 will be described in detail.

[0017] [Functional Configuration of Information Processing System 1] FIG. 2 is a diagram showing an example of the functional configuration of the information processing system 1 according to the embodiment. In FIG. 2, the terminal device 10, the virtual reality service server 50, and the information processing device 100 communicate with each other via the network NW. The network NW is a communication network such as the Internet or an intranet including wired or wireless communication lines, and may include various networks such as a public communication network, a mobile phone network, a local area network (LAN), and a wide area communication network (WAN). Also, the network NW may have a configuration in which different network types are combined. For example, the terminal device 10 may be connected via wireless communication (such as Wi-Fi or LTE / 5G), and the virtual reality service server 50 and the information processing device 100 may be connected via a wired LAN or the like. As another aspect, the information processing device 100 may be integrated with the virtual reality service server 50.

[0018] The terminal device 10 is a computer device that includes, for example, a communication unit 12, an input unit 14, a display unit 16, a control unit 18, and a storage unit 20. The storage unit 20 is a storage device that stores the virtual reality application 20A. The virtual reality application 20A is an application program that works in cooperation with the virtual reality service server 50 to realize a virtual reality service, and the division of functions between the virtual reality application 20A and the virtual reality service server 50, as described below, may be changed as appropriate. For example, the functions of the virtual reality application 20A may be enriched to control the operation of the agent avatar AA in S7 above. Conversely, the virtual reality application 20A may be a thin client application. In that case, the virtual reality application 20A only needs to have at least the function of outputting data received from the virtual reality service server 50 (screen output and audio output).

[0019] The communication unit 12 has the function of sending and receiving data between the virtual reality service server 50 and the information processing device 100 via the network NW. The communication unit 12 includes a wireless communication module, an antenna, and a network interface circuit. The communication unit 12 may support wireless communication methods such as Wi-Fi (wireless LAN), LTE, 5G, or wired communication methods such as wired LAN.

[0020] The input unit 14 has the function of accepting text input, voice input, gesture input, touch input, etc. from the user. The input unit 14 may be composed of, for example, a keyboard, microphone, touch panel, camera, motion sensor, etc. More specifically, if the terminal device 10 is a PC (personal computer) or smartphone, the input unit 14 includes input means such as a physical keyboard, a software keyboard displayed on a touch panel, a mouse, and a trackpad. On the other hand, if the terminal device 10 is an HMD (Head Mounted Display), the input unit 14 includes input means such as a gesture recognition sensor, eye-tracking device, motion controller, and microphone for voice input.

[0021] The display unit 16 has the function of visually presenting various information to the user, such as images of the virtual space generated by the virtual reality application 20A, and guidance / explanations by the agent avatar AA. The display unit 16 may be composed of, for example, a liquid crystal display, an organic EL display, an HMD, etc.

[0022] The control unit 18 executes the virtual reality application 20A stored in the memory unit 20, controls the display of the virtual space on the display unit 16 in response to user input, and has the function of comprehensively controlling the operation of the entire terminal device 10, including controlling data transmission and reception via the communication unit 12. The control unit 18 is composed of a computing device including, for example, a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit).

[0023] In addition to the virtual reality application 20A, the memory unit 20 can temporarily or permanently store user operation history, cache data related to the virtual space, dialogue logs, configuration information, and the like. The memory unit 20 is composed of a storage medium such as flash memory, SSD, or HDD.

[0024] The virtual reality service server 50 is a computer device that includes, for example, a communication unit 52, a control unit 54, and a configuration unit 56. The storage unit 58 stores information such as, for example, user data 58A, virtual space mesh data 58B, virtual space static data 58C, virtual space dynamic data 58D, and user history data 58E.

[0025] The communication unit 52 includes, for example, communication equipment such as a network interface card (NIC), a router connection port, a LAN controller, and an optical communication module. As a result, the communication unit 52 has the function of receiving user input data and status information from the terminal device 10, control information and processing results from the information processing device 100, and transmitting virtual space update information, avatar behavior information, audio / video data, etc. to the outside.

[0026] The control unit 54 is a software function unit for comprehensively controlling the operation of the entire virtual reality service server 50, and is implemented by a computing device including information processing circuits such as a CPU, memory, chipset, and bus controller. The control unit 54 performs processing related to the creation, maintenance, updating, and management of the virtual space while referring to various data stored in the storage unit 58. More specifically, the control unit 54 performs authentication processing and determines access rights for each user based on user data 58A, and generates / updates rendering information and object behavior information of the virtual reality space using virtual space mesh data 58B, virtual space static data 58C, and virtual space dynamic data 58D. The control unit 54 also stores history information regarding actions performed by the user as a user avatar UA as user history data 58E.

[0027] Furthermore, the control unit 54 analyzes user input information (text information, gesture information, voice instructions, etc.) received from the terminal device 10 via the communication unit 52 and updates the state in the virtual space in real time. By distributing the updated virtual space information to the terminal devices of other participating users, it becomes possible to provide a multi-user virtual reality service synchronously.

[0028] The settings unit 56 is a software function unit that manages various setting information in the virtual reality service and supports the control processing of the virtual space by the control unit 54. Like the control unit 54, it is implemented by a computing device that includes information processing circuits such as a CPU, memory, chipset, and bus controller. The settings unit 56 holds, for example, avatar setting information for each user, parameters related to operation rights, operation restrictions, environmental conditions within the virtual space, and information related to service usage conditions. As a result, the appearance and behavior of the avatar used by each user in the virtual space, the type of operation interface (voice, gaze, gestures, etc.), speaking rights, and movement range are controlled according to individual settings.

[0029] Furthermore, the settings unit 56 can also manage information related to the structure and environment of the virtual space, such as the map layout, initial placement of objects, and display settings such as lighting conditions and sound effects. In addition, it can receive changes input via the communication unit 52 based on user requests for setting changes and reflect them in the operation of the control unit 54. This allows for the flexible provision of virtual reality services according to the user's preferences and usage environment.

[0030] Furthermore, the setting unit 56 has a function to set an agent avatar AA in response to a user entering a virtual space. More specifically, for example, when a user enters one of several virtual spaces, the setting unit 56 assigns at least one agent avatar AA from among the several agent avatar AAs to the user or virtual space and displays it on the terminal device 10. The agent avatar AA may be freely set by each user entering the virtual space according to their own preferences and tastes, and in that case, the setting data such as the appearance and behavior of the agent avatar AA may be set by each user entering the space. In another embodiment, the agent avatar AA may be set by the user who is the creator of the virtual space as an avatar suitable for that virtual space, and in that case, a common agent avatar AA will be set and displayed to users entering that virtual space.

[0031] The information processing device 100 is a computer device that includes, for example, a communication unit 110, a first acquisition unit 120, a generation unit 130, a second acquisition unit 140, and an output unit 150. The storage unit 160 stores information such as an AI model 162.

[0032] The communication unit 110 includes, for example, communication equipment such as a network interface card (NIC), a router connection port, a LAN controller, and an optical communication module. As a result, the communication unit 110 has the function of receiving input data from the virtual reality service server 50 and transmitting control information output by the AI ​​model 162 to the outside. The acquisition function of the first acquisition unit 120 and the output function of the output unit 150, described below, can be realized using the communication unit 110.

[0033] The first acquisition unit 120, the generation unit 130, the second acquisition unit 140, and the output unit 150 are software function units for realizing the functions of the information processing device 100, and are implemented by an arithmetic unit including information processing circuits such as a CPU, memory, chipset, and bus controller.

[0034] The first acquisition unit 120 acquires from the virtual reality service server 50 dialogue information relating to the interaction between the user avatar UA and the agent avatar AA in one of the multiple virtual spaces, and virtual space information relating to that virtual space, in response to the user avatar UA interacting with the agent avatar AA in one of the multiple virtual spaces. In this embodiment, the virtual reality service server 50 and the information processing device 100 communicate via an API (Application Programming Interface) provided by the virtual reality service server 50, and the first acquisition unit 120 acquires the dialogue information and virtual space information via the API. Here, the virtual space information includes some or all of the user data 58A, virtual space mesh data 58B, virtual space static data 58C, virtual space dynamic data 58D, and user history data 58E, which will be described later.

[0035] The generation unit 130 generates prompts from the dialogue information and virtual space information acquired by the first acquisition unit 120 to cause the AI ​​model 162 to output control information regarding the actions that the agent avatar AA should perform. As will be described later, the generation unit 130 generates prompts that instruct the AI ​​model 162 to select an appropriate action command from the candidate action commands for the avatar that have been pre-implemented and prepared by the virtual reality service.

[0036] The second acquisition unit 140 acquires control information regarding the actions of the agent avatar AA in the virtual space by inputting the prompt generated by the generation unit 130 to the AI ​​model 162. Here, the AI ​​model 162 can be configured, for example, by a large language model (LLM) based on the Transformer architecture. Such an AI model 162 is trained specifically for natural language processing tasks and can infer appropriate response actions (guidance, explanation, direction, etc.) that the agent avatar AA should perform based on the dialogue information input by the user and the prompt containing information about the virtual space in question. By learning large amounts of dialogue data and knowledge data related to the virtual space in advance, the AI ​​model 162 has the ability to contextually understand the user's input intent and output natural dialogue responses and action instructions appropriate to the situation.

[0037] In Figure 2, for the sake of explanation, the AI ​​model 162 is shown as a single model. However, the present invention is not limited to such a configuration, and the AI ​​model 162 may be composed of multiple models. More specifically, the AI ​​model 162 may be composed of, for example, a dialogue generation model that generates dialogue responses and an action generation model that controls the actions of the agent avatar AA. By separating the execution of dialogue processing and action control, specialized learning and optimization for each process become possible, resulting in more natural dialogue responses and highly accurate action control.

[0038] The output unit 150 outputs the control information acquired by the second acquisition unit 140 to the virtual reality service server 50. The control unit 54 of the virtual reality service server 50 controls the operation of the agent avatar AA in the virtual space based on the output control information and outputs it to the terminal device 10. The display unit 16 of the terminal device 10 displays the output operation of the agent avatar AA in the virtual space.

[0039] [Data structure of Information Processing System 1] Next, the data structure of the information processing system 1 will be described with reference to Figures 3 to 7. Figure 3 shows an example of the structure of user data 58A. User data 58A is, for example, a combination of user ID, user avatar setting data, agent avatar setting data, and user-generated virtual space ID.

[0040] The User ID is identification information used to identify each user of the virtual reality service. User Avatar Settings Data is the settings information for the user avatar UA set by the user. More specifically, User Avatar Settings Data includes data such as the appearance of the user avatar UA (clothing, hairstyle, face shape, etc.), voice, display name, and user attributes (age, gender, address, nationality, etc.). Agent Avatar Settings Data includes data such as the appearance of the agent avatar AA, voice, behavior, personality, tone of voice, speaking speed, choice of polite / informal language, response policy to specific words, and response policy according to user attributes. Users can set User Avatar Settings Data and Agent Avatar Settings Data on the virtual reality application 20A, for example, by using the avatar creation tool provided by the virtual reality service or by selecting items (selection of clothing parts, selection of hairstyle parts, etc.). The User-Generated Virtual Space ID is identification information used to identify the virtual space created by the user. In the virtual reality service of this embodiment, a user can create one or more virtual spaces; therefore, in the example shown in Figure 3, multiple virtual space IDs are set for a single user.

[0041] Figure 4 shows an example of the configuration of virtual space mesh data 58B. As an example, Figure 4 shows virtual space mesh data 58B identified by the virtual space ID "V100" of user data 58A in Figure 3. More specifically, virtual space mesh data 58B includes polygon mesh information for constructing three-dimensional structures, backgrounds, terrain, objects, etc., in the virtual space. Here, polygon mesh information consists of multiple vertex information, edge information, and face information, and each vertex may be assigned normal vectors, texture coordinates (U, V), material attributes, shading information, etc., in addition to position coordinates (X, Y, Z).

[0042] The virtual reality service provides users with parts (lines, shapes, etc.) and objects (roofs, walls, chairs, plants, etc.) as polygon meshes that can be used to create virtual spaces. Users can create virtual spaces by freely combining these parts and objects on the virtual reality application 20A. In the virtual reality service, users can create virtual spaces through intuitive operation, but in the background, the placement information, shape information, and identification information of the parts and objects that the user places in the virtual space are recorded as polygon meshes.

[0043] Furthermore, the virtual space mesh data 58B may also include configuration information regarding interaction areas for each avatar or virtual object, namely hitboxes and operation trigger areas for reaction detection in response to user input or gaze input, thereby enabling high-precision control of user interaction within the virtual space.

[0044] As shown in Figure 4, the virtual space mesh data 58B includes, for example, navigation points NP and entry points EP. These navigation points NP and entry points EP are set by the user who created the virtual space (hereinafter sometimes simply referred to as the creator) on the virtual reality application 20A.

[0045] Navigation points NP are designated by the creator as distinctive locations within the virtual space. As described later, by setting navigation points NP, the agent avatar AA can prioritize guiding other user avatars UA (hereinafter sometimes simply referred to as visitors) that enter the virtual space to the navigation points NP. Navigation points NP may be specified as coordinate data indicating a predetermined location, or as an object placed by the creator. The predetermined location could be, for example, a point near the entrance of a cafe in the field, or the center of a shopping street in the case of a shopping street. On the other hand, the object could be, for example, a seating object placed by the user in the case of a cafe, or a sign object in the case of a shopping street.

[0046] The entry point (EP) is set by the creator as the initial position of a visitor upon entering the virtual space. Similar to the navigation point (NP), the entry point (EP) may be specified as coordinate data in the virtual space or as an object. By setting the entry point (EP), the creator can present the virtual space to the visitor in accordance with the virtual reality experience they wish to provide.

[0047] Figure 5 shows an example of the configuration of virtual space static data 58C. Virtual space static data 58C is created and managed for each virtual space generated by the creator, for example, by the creator on the virtual reality application 20A. Virtual space static data 58C is associated with, for example, a virtual space ID, virtual space name, virtual space description, entry point, entry point coordinates, entry point description, navigation point, navigation point coordinates, navigation point description, navigation order, agent avatar initial position, and creator instructions.

[0048] The virtual space name is information indicating the name assigned to the virtual space. The virtual space description is descriptive information about the virtual space. Creators can freely set the virtual space name and description according to the theme and concept of the virtual space they have created. The entry point is information indicating the initial position of a visitor entering the virtual space, and the entry point coordinates are information indicating the coordinates of the entry point within the virtual space. The entry point description is descriptive information about the entry point. Creators can, for example, set the entry point to a location they deem optimal, taking into account the visitor's virtual reality experience.

[0049] Navigation points are information indicating distinctive locations within the virtual space, and navigation point coordinates are information indicating the coordinates of the navigation points within the virtual space. Navigation point descriptions are descriptive information about the navigation points. Creators can set navigation points at desired locations, such as appealing points in the virtual space they have created or areas they want visitors to pay attention to. In Figure 5, as an example, four locations, from navigation point NP1 to navigation point NP4, are set as navigation points, but the present invention is not limited to such a configuration, and creators can set any number of navigation points.

[0050] The navigation order is information indicating the priority given to the agent avatar AA when guiding visitors to navigation points. By setting the navigation order, creators can provide visitors with a virtual reality experience in the order they deem optimal. The agent avatar initial position is information indicating the initial position of the agent avatar AA displayed when a visitor enters the virtual space. This information, such as the virtual space name, virtual space description, entry point, entry point coordinates, entry point description, navigation points, navigation point coordinates, navigation point description, navigation order, and agent avatar initial position, can be freely set by the creator, or some or all of the items can be left blank and automatically generated by the AI ​​model 162. For example, the prompt may only contain the contents of the virtual space mesh data 58B, and some or all of the items of the virtual space static data 58C may be extracted by the AI ​​model 162 based on the virtual space mesh data 58B.

[0051] Creator instructions are directives given by the creator to the AI ​​model 162 to determine the behavior and various settings of the agent avatar AA within the virtual space. More specifically, creator instructions include various instructions regarding the agent avatar AA's settings, such as its personality, tone of voice, speaking speed, choice of polite / informal language, response policies to specific words, and behavior in accordance with background information. Creators specify creator instructions based on the worldview of the virtual space they have created and the virtual reality experience they want to give visitors. Unlike numerical information such as the navigation order mentioned above, creators can flexibly specify directives using natural language, such as policies regarding conversations with the agent avatar AA (for example, ignoring or rejecting conversations related to intellectual property other than virtual reality services or settings outside the virtual space (the fourth wall), guiding in a gentlemanly tone, etc.) and policies regarding guidance to navigation points (for example, prioritizing guidance to outdoors over indoors, first guiding to a map of the shopping district and letting the user choose a destination, etc.). Furthermore, if there is a conflict between the creator's instructions and the agent avatar settings data (for example, if the agent avatars have completely opposite personalities), the creator's instructions may take precedence over the agent avatar settings data (they may overwrite it).

[0052] As explained with reference to Figure 3, in this embodiment, one or more agent avatars are set up for each user, so the agent avatar setting data is stored in user data 58A. In another embodiment, if one or more agent avatars are set up for each virtual space, the agent avatar setting data may be set by the creator as a data item in the virtual space static data 58C. In that case, the setting unit 56 sets agent avatar AA for each user based on the agent avatar setting data in the virtual space static data 58C when the user enters the virtual space.

[0053] Figure 6 shows an example of the configuration of virtual space dynamic data 58D. Virtual space dynamic data 58D is created and managed for each virtual space generated by the creator, and is updated in real time by the settings unit 56 while a visitor is in the virtual space. For example, virtual space dynamic data 58D associates the virtual space ID, the user who is in the virtual space, the location of the user who is in the virtual space, and the line of sight of the user who is in the virtual space.

[0054] "Currently Entering Users" is information that identifies a user who has entered the virtual space. "Currently Entering Users' Location" is information indicating the coordinates of the currently entering user's (user avatar UA) location. "Currently Entering Users' Line of Sight" is information indicating the direction of the currently entering user's line of sight (for example, a combination of azimuth angle φ and elevation angle θ). By combining the currently entering user's location and the currently entering user's line of sight, it is possible to track in real time where the currently entering user is looking within the virtual space. The setting unit 56 may, for example, set a user as a currently entering user when that user enters the virtual space, and remove that user from the list of currently entering users when that user leaves the virtual space (logs out). Alternatively, the setting unit 56 may remove a user from the list of currently entering users when a certain period of time has elapsed since that user last performed an operation within the virtual space.

[0055] Figure 7 shows an example of the configuration of user history data 58E. User history data 58E is created and managed for each user of the virtual reality service and is updated in real time by the configuration unit 56 while each user is in the virtual space. User history data 58E is, for example, associated with a timestamp, virtual space ID, navigation point, interaction, and action.

[0056] A timestamp is information indicating the point in time when a user performed some kind of interaction or action in the virtual space. The interaction is information indicating the content of the interaction the user had at the time of the timestamp. The action is information indicating the action the user performed at the time of the timestamp. In Figure 7, as an example, movement between navigation points is recorded as an action, but the action may include any action provided by the virtual reality service. For example, an action may include taking photos of avatars in the virtual space, sending a friend request to another user's avatar, or purchasing avatar items (clothes, accessories, etc.). Here, taking a photo means taking a screenshot in the virtual space, and sending a friend request means sending a request to enable friend-specific functions (e.g., messaging function) by following each other. Also, in Figure 7, as an example, the interaction and action at a navigation point are recorded, but the present invention is not limited to such a configuration, and instead of a navigation point, for example, the interaction and action at a coordinate point in the virtual space may be recorded.

[0057] [Operation of Information Processing System 1] Next, the operation of the information processing system 1 will be explained with reference to Figures 8 to 11. First, the user launches the virtual reality application 20A on the terminal device 10 and enters a virtual space. For example, the user enters a virtual space by displaying a list of virtual spaces that are currently accepting entries from the menu screen of the virtual reality application 20A and selecting the virtual space they wish to enter. Once the user enters the virtual space, the setting unit 56 refers to the user data 58A to obtain user avatar setting data and sets the user avatar UA based on the obtained user avatar setting data. At the same time, the setting unit 56 refers to the user data 58A to obtain agent avatar setting data and sets the agent avatar AA for the user based on the agent avatar setting data. The control unit 54 places and displays the set user avatar UA and agent avatar AA on the display unit 16 of the terminal device 10. Next, the user interacts with the agent avatar AA by inputting information such as text or voice into the terminal device 10.

[0058] Figure 8 shows an example of interaction with an agent avatar AA in a virtual space. As an example, Figure 8 shows the timing when a user enters a virtual space identified by the virtual space ID "V100". First, when a user enters the virtual space with virtual space ID "V100" on the virtual reality application 20A, the virtual reality application 20A sends the virtual space ID "V100" to the virtual reality service server 50. When the virtual reality service server 50 receives the virtual space ID "V100", the control unit 54 uses the virtual space ID "V100" as a key, refers to the virtual space static data 58C in Figure 5, places the user avatar UA at the entry point, and places the agent avatar AA at the agent avatar's initial position, and displays them on the virtual reality application 20A.

[0059] Simultaneously, the virtual reality service server 50 transmits the virtual space ID "V100" to the information processing device 100. When the information processing device 100 receives the virtual space ID "V100", the first acquisition unit 120 uses the virtual space ID "V100" as a key to acquire virtual space information from the virtual reality service server 50. At this time, if the user spoke at the same time as entering, the virtual reality application 20A transmits the speech information to the virtual reality service server 50, and the first acquisition unit 120 may acquire the speech information from the virtual reality service server 50 along with the virtual space information. Here, the virtual space information includes at least a portion of the user data 5BA, virtual space mesh data 58B, virtual space static data 58C, virtual space dynamic data 58D, and user history data 58E. Here, it is assumed that the virtual reality service server 50 transmits all of this data to the information processing device 100 as virtual space information.

[0060] When the first acquisition unit 120 acquires virtual space information, the generation unit 130 generates an initial prompt for input to the AI ​​model 162 based on this information. Here, the initial prompt is used to cause the AI ​​model 162 to output control information regarding the initial actions to be performed by the agent avatar AA, based on the virtual space information (and, if necessary, speech information) into which the user has entered. For example, the initial prompt includes instruction information that, along with the virtual space information, instructs the user to "select and output the initial action to be performed by the agent avatar AA from the command candidates." Details of the command candidates will be described later. The initial prompt may also include user data 58A.

[0061] The second acquisition unit 140 acquires control information regarding the initial operation of the agent avatar AA in the virtual space by inputting the generated initial prompt into the AI ​​model 162. The output unit 150 outputs the control information acquired by the second acquisition unit 140 to the virtual reality service server 50 via an API for execution of operations. When the virtual reality service server 50 receives the control information, the control unit 54 controls the initial operation of the agent avatar AA in the virtual space based on the output control information and outputs it to the terminal device 10. The display unit 16 of the terminal device 10 displays the outputted initial operation of the agent avatar AA in the virtual space. Figure 8 shows an example where, as an initial operation, the agent avatar AA displays an initial message (for example, "Welcome to [Virtual Space Name]!") in the dialog box DB (i.e., it shows an example where "Dialogue" is selected from the operation commands described later). After being placed in the initial position, if a new user enters, the agent avatar AA similarly displays an initial message (for example, "Welcome to [Virtual Space Name]!") in the dialog box DB.

[0062] Next, a user who has entered the virtual space may, for example, wish to visit a recommended location in the virtual space and input text or voice (for example, "Take me to a recommended location") into the terminal device 10 to inquire about the recommended location. When the user inputs text or voice, the terminal device 10 sends the input information (and may include all information in the dialog box DB) along with the virtual space ID of the virtual space the user is currently in as dialogue information to the virtual reality service server 50. Upon receiving the dialogue information and the virtual space ID, the virtual reality service server 50 identifies the virtual space associated with the virtual space ID from among multiple virtual spaces.

[0063] When the virtual reality service server 50 identifies a virtual space, it transmits the virtual space ID of the identified virtual space to the information processing device 100. When the information processing device 100 receives the virtual space ID, the first acquisition unit 120 uses the received virtual space ID as a key to acquire virtual space information and dialogue information from the virtual reality service server 50. The virtual space information acquired at this time is the same as the virtual space information at the time of initial operation, however, if, for example, the user's field of view in the virtual space has changed since the initial time, or if the virtual space information has changed since the initial time, the content of the acquired virtual space information will change. When the first acquisition unit 120 acquires the dialogue information and virtual space information, the generation unit 130 generates a prompt PR for input to the AI ​​model 162 based on this information.

[0064] Figure 9 shows an example of a prompt PR generated by the generation unit 130. In Figure 9, the prompt PR includes, for example, information areas from area A1 to area A7.

[0065] Area A1 is an area that contains instructions to be given to the AI ​​model 162. More specifically, for example, Area A1 instructs the agent avatar AA to select and output control information regarding the actions it should perform from a list of action command candidates, based on dialogue information, virtual space information, and other information. In the example of prompt PR in Figure 9, the instruction is to select an appropriate action command from the list of action command candidates described in Area A6, but the present invention is not limited to such a configuration, and the virtual reality service server 50 may be instructed to output an executable script according to a predetermined programming language. In that case, the instructions may be in a sequence that first instructs the AI ​​model 162 to identify an appropriate action command, and then instructs it to output a script that implements the identified action command.

[0066] Area A2 is the area where dialogue information is recorded. In the example prompt PR in Figure 9, in addition to the user's utterance data ("Take me to a recommended place"), it also includes the agent's previous utterance data ("Welcome to Sky Cafe!"). By including not only the user's utterance data but also the preceding dialogue data in the prompt PR in this way, it is possible to obtain control information regarding the operation of the agent avatar AA that is more in line with the user's intent.

[0067] Area A3 is an area for describing virtual space information. In Figure 9, the brackets [] represent the content of the data described within the brackets. In the example of prompt PR in Figure 9, the virtual space information includes all of the virtual space mesh data 58B, virtual space static data 58C, and virtual space dynamic data 58D. However, the present invention is not limited to such a configuration, and only some of the content of the virtual space mesh data 58B, virtual space static data 58C, and virtual space dynamic data 58D may be described. In this way, by including not only the virtual space mesh data 58B but also the virtual space static data 58C, which represents the static characteristics of the virtual space, and the virtual space dynamic data 58D, which represents the real-time status of the virtual space, in prompt PR, it is possible to obtain control information regarding the operation of the agent avatar AA that is more convenient for the user.

[0068] Area A4 is the area for recording user information. In the example prompt PR in Figure 9, the user information includes the contents of the user avatar settings data and the agent avatar settings data. As mentioned above, the user avatar settings data includes the appearance of the user avatar UA (clothing, hairstyle, face shape, etc.), voice, movement patterns, display name, status display, and user attributes (age, gender, address, nationality, etc.). The agent avatar settings data includes, for example, the appearance of the agent avatar AA, behavior, personality, tone of voice, speaking speed, choice of polite / informal language, response policy to specific words, and response policy according to user attributes. Therefore, by including this information in the prompt PR, the AI ​​model 162 can output speech and actions that are more appropriate for the user avatar UA and more in line with the image of the agent avatar AA.

[0069] For example, if the user avatar UA's appearance is set to that of a child, AI model 162 may output speech using simpler vocabulary. Conversely, if the user avatar UA's appearance is set to that of an adult, AI model 162 may output speech using more general vocabulary. Also, for example, if the agent avatar AA's personality is set to that of an introverted person, AI model 162 may output speech using a quiet tone. Conversely, if the agent avatar AA's personality is set to that of an extroverted person, AI model 162 may output speech using a lively tone.

[0070] Area A5 is the area for recording user history information. In the example prompt PR in Figure 9, the user history information includes the contents of the user history data. As mentioned above, user history data is information that records what kind of conversations or actions the user performed at what time, at what navigation point. By including such information in the prompt PR, the AI ​​model 162 can output utterances and actions that are more in line with the user's preferences. For example, if the user has frequently visited indoor facilities in the past, the AI ​​model 162 may output actions of the agent avatar AA that prioritize guiding the user to indoor facilities. Conversely, if the user has frequently visited outdoor facilities in the past, the AI ​​model 162 may output actions of the agent avatar AA that prioritize guiding the user to outdoor facilities.

[0071] Area A6 is an area that lists possible actions that agent avatar AA can perform as action commands. Action commands represent APIs for calling and executing avatar actions implemented by the virtual reality service, and include actions such as: dialogue (dialogue with another avatar without moving), move 1 (move by specifying the coordinates of the destination), move 2 (move by specifying the destination navigation point), warp (move immediately and simultaneously with another avatar by specifying the destination navigation point), follow (move by following another avatar), call (check if another avatar is within a specified range of your own avatar, and if not, transmit speech information to the other avatar), cancel dialogue (check if another avatar is within a specified range of your own avatar, and if not, end the dialogue), hide (check if another avatar is within a specified range of your own avatar, and if not, end the dialogue and hide yourself), and stand still (wait without performing any action if none of the actions are necessary). These actions are implemented by the virtual reality service as actions that can be performed in the virtual space by all avatars, including not only agent avatars (AA) but also user avatars (UA). Action commands are not limited to those listed in area A6, and may also include, for example, world crafting by agents (generating objects in the virtual world), taking photos, sending friend requests, emotes (actions such as laughing or getting angry), and changing clothes (changing the items worn by the avatar).

[0072] Although not shown in Figure 9, area A6 specifies the parameters that the AI ​​model 162 should output, as needed, in relation to each action candidate. For example, in the case of dialogue and calling, the parameters specified are to output the dialogue content and calling content, respectively. In the case of movement 1, the parameters specified are to output the coordinates of the destination, and the optional dialogue content to be displayed during movement. In the case of movement 2 and warp, the parameters specified are to output the navigation point of the destination, and the optional dialogue content to be displayed during movement. By including this list of action candidates and the parameters to be output in the prompt PR, the response results from the AI ​​model 162 can be stabilized compared to when instructions are given without providing options.

[0073] Area A7 is an area for inputting attached data to be attached to a prompt PR into the AI ​​model 162. In the example prompt PR in Figure 9, the attached data includes image data, audio data, video data, and document data. More specifically, area A7 contains link information for each attached data, and the AI ​​model 162 can retrieve the attached data by referring to this link information. The attached data may be, for example, data uploaded by the user from the virtual reality application 20A to the virtual reality service server 50, or data generated by the generation unit 130 based on dialogue information and / or virtual space information. For example, by attaching dialogue information to the prompt PR not only as text information in area A2 but also as audio data, the AI ​​model 162 can output control information that more accurately reflects the user's intentions.

[0074] Furthermore, while Figure 9 shows instructions to select a single action command candidate from multiple candidates for simplicity of explanation, the user may be instructed to select multiple action command candidates sequentially as needed, provided that this does not cause discomfort to the user. For example, by sequentially selecting Move 2 and Call, it becomes possible to implement an operation where Agent Avatar AA first moves to Navigation Point NP, then checks whether a User Avatar UA exists within a predetermined range, and if the User Avatar UA does not exist, transmits speech information to the User Avatar UA. Moreover, the prompt PR information shown in Figure 9 is merely an example, and may only be a part of it. The prompt PR information may be raw data, or it may be data that has been processed in some way beforehand.

[0075] When the generation unit 130 generates a prompt PR, the second acquisition unit 140 inputs the generated prompt to the AI ​​model 162 to acquire control information regarding the agent avatar AA's actions in the virtual space. The output unit 150 outputs the control information acquired by the second acquisition unit 140 to the virtual reality service server 50 via an API for execution of actions. When the virtual reality service server 50 receives the control information, the control unit 54 controls the agent avatar AA's actions in the virtual space based on the output control information and outputs it to the terminal device 10. The display unit 16 of the terminal device 10 displays the outputted actions of the agent avatar AA in the virtual space.

[0076] Figure 10 shows an example of the operation of agent avatar AA controlled based on control information. Figure 10 shows an example of a scene in which AI model 162 outputs control information in response to prompt PR shown in Figure 9, and agent avatar AA performs an action based on the outputted control information. The control information in Figure 10 includes, for example, the action command Move 2 and the specified parameters (destination "Navigation Point 3" and dialogue content "I will guide you to a terrace seat with a parasol that has a great view. Follow me!"). Therefore, agent avatar AA moves to navigation point NP3 while outputting the dialogue content "I will guide you to a terrace seat with a parasol that has a great view. Follow me!" to the dialog box DB.

[0077] Figure 11 shows another example of the operation of agent avatar AA controlled based on control information. Figure 11 shows another example of a scenario in which AI model 162 outputs control information in response to prompt PR shown in Figure 9, and agent avatar AA performs an action based on the outputted control information. The control information in Figure 11 includes, for example, a warp operation command and specified parameters (destination "Navigation Point 3" and dialogue content "I will guide you to a terrace seat with a parasol and a great view. Warping!"). Therefore, agent avatar AA warps to navigation point NP3 while outputting the dialogue content "I will guide you to a terrace seat with a parasol and a great view. Warping!" to the dialog box DB.

[0078] Thus, in this embodiment, the user inputs dialogue information to the terminal device 10, and the agent avatar AA is controlled accordingly based on the outputted control information. However, the present invention is not limited to such a configuration, and the control unit 54 of the virtual reality service server 50 may switch the operation of the agent avatar AA to an operation terminal (not shown) operated by the operator of the virtual reality service, or enable it to accept operations, in response to a predetermined event occurring within the virtual reality service. The predetermined event may be, for example, the user performing a predetermined operation on the virtual reality application 20A to request a handover, or it may be, for example, the user avatar UA performing aggressive behavior (e.g., dialogue containing certain prohibited words) towards the agent avatar AA or other user avatars. In that case, the operator will input an operation command directly from the operation terminal and send it to the agent avatar AA to perform the operation.

[0079] In this embodiment, the virtual reality service server 50 transmits dialogue information and virtual space information to the information processing device 100, and the second acquisition unit 140 inputs this information to the AI ​​model 162. However, the present invention is not limited to such a configuration. For example, the virtual reality service server 50 may transmit dialogue information and virtual space information to a first AI agent, and the first AI agent may input this dialogue information and virtual space information as input information to a second AI agent, which is the AI ​​model 162, via A2A (Agent to Agent) communication. Subsequently, the second AI agent may output control information to the first AI agent again via A2A communication, and the virtual reality service server 50 may receive the control information from the first AI agent. In this configuration, the first AI agent and the second AI agent communicate with each other, for example, according to the MCP (Model Context Protocol) protocol.

[0080] [Process Flow] Next, the processing flow of the information processing system 1 will be explained with reference to Figures 12 to 14. Figure 12 is a sequence diagram illustrating the processing flow performed through the cooperation of the terminal device 10, the virtual reality service server 50, and the information processing device 100.

[0081] First, the user launches the virtual reality application 20A on the terminal device 10 and enters a virtual space (S100). In response, the terminal device 10 sends entry information, including the user's user ID and the virtual space ID of the virtual space, to the virtual reality service server 50 (S102). Upon receiving the user ID and virtual space ID, the virtual reality service server 50 identifies the virtual space the user entered (S104) and sends the virtual space ID to the information processing device 100 (S108). Upon receiving the virtual space ID, the information processing device 100 uses the virtual space ID as a key to obtain virtual space information from the virtual reality service server 50 (S108), and generates an initial prompt based on the obtained virtual space information (S110). Next, the information processing device 100 inputs the generated initial prompt into the AI ​​model 162 to obtain control information regarding the initial operation of the agent avatar AA (S112).

[0082] When the information processing device 100 acquires control information, it outputs the acquired control information to the virtual reality service server 50 (S114). Based on the user ID received in S102, the virtual reality service server 50 refers to the user avatar setting data and agent avatar setting data of the user data 58A and sets the user avatar UA and agent avatar AA for the user (S116). Based on the control information received in S114, the virtual reality service server 50 displays the user avatar UA and agent avatar AA set in S116 on the terminal device 10 and controls the initial operation (S118).

[0083] Next, the user inputs information such as text and voice into the terminal device 10 as dialogue information with the agent avatar AA (S120). The terminal device 10 associates the input dialogue information with the user's user ID and the virtual space ID of the virtual space and sends it to the virtual reality service server 50 (S122). Based on the received virtual space ID, the virtual reality service server 50 identifies the virtual space the user is in from among multiple virtual spaces (S124). Next, the virtual reality service server 50 sends the virtual space ID to the information processing device 100 (S116).

[0084] When the information processing device 100 receives a virtual space ID, it uses the virtual space ID as a key to obtain virtual space information and dialogue information from the virtual reality service server 50 (S128), and generates a prompt PR based on the obtained virtual space information (S130). Next, the information processing device 100 inputs the generated prompt PR to the AI ​​model 162 to obtain control information regarding the initial operation of the agent avatar AA (S132). When the information processing device 100 obtains the control information, it outputs the obtained control information to the virtual reality service server 50 (S134). The virtual reality service server 50 controls the operation of the agent avatar AA based on the received control information and displays it on the terminal device 10 (S136). This completes the processing shown in this sequence diagram.

[0085] Thus, in this embodiment, when a user inputs dialogue information, the virtual reality service server 50 transmits the dialogue information and virtual space information to the information processing device 100, and the information processing device 100 outputs control information regarding the operation of the agent avatar AA. However, even when the user has not input dialogue information, the virtual reality experience for the user can be improved by updating or regenerating the control information as needed. In other words, the timing at which the virtual reality service server 50 transmits the dialogue information and virtual space information to the information processing device 100 may include various timings other than when the user inputs dialogue information.

[0086] As an example of transmission timing, the virtual reality service server 50 may determine whether there has been a change in the virtual space information the user is currently in, and if it determines that there has been a change in the virtual space information the user is currently in, it may send the virtual space information change information to the information processing device 100. The virtual reality service server 50 may send only the changed portion of the virtual space information, or it may send the entire virtual space information. The virtual space information change information may, for example, represent the difference obtained by image analysis of polygon mesh information in virtual space mesh data 58B, or represent the change of navigation point NP in virtual space static data 58C. Accordingly, the generation unit 130 may, together with the virtual space information change information, generate a prompt PR that instructs, for example, "Based on the virtual space information change information, change the operation of agent avatar AA as necessary," and the second acquisition unit 140 may input the prompt PR to the AI ​​model 162.

[0087] Figure 13 is a flowchart showing an example of the control of the transmission of dialogue information and virtual space information by the virtual reality service server 50. The processing shown in the flowchart in Figure 13 is repeatedly executed, for example, while a user is entering a virtual space.

[0088] First, the virtual reality service server 50 identifies the virtual space in which the user is currently present (S200). Next, the virtual reality service server 50 determines whether or not it has obtained dialogue information from the user (S202). If it determines that it has obtained dialogue information from the user, the virtual reality service server 50 sends the dialogue information and virtual space information to the information processing device 100 (S204). On the other hand, if it determines that it has not obtained dialogue information from the user, the virtual reality service server 50 determines whether or not there has been a change in the virtual space information in which the user is currently present (S206). If it determines that there has been a change in the virtual space information in which the user is currently present, the virtual reality service server 50 sends the change information of the virtual space to the information processing device 100 (S208). This completes the processing of this flowchart.

[0089] According to the processing in this flowchart, when there is a change in the virtual space information, the virtual reality service server 50 can reduce the processing load on the information processing device 100 by sending information corresponding only to the changed portion of the virtual space information to the information processing device 100.

[0090] As another example of transmission timing, the virtual reality service server 50 may measure the elapsed time since the user last entered dialogue information, and if a certain period of time has passed without any new dialogue information being entered, it may send notification information indicating the passage of a certain period to the information processing device 100. Accordingly, the generation unit 130 may generate a prompt PR instructing, for example, "No additional dialogue information has been received from the user avatar UA, and a certain period of time has passed. Please change the behavior of the agent avatar AA as needed," and the second acquisition unit 140 may input the prompt PR to the AI ​​model 162.

[0091] Figure 14 is a flowchart showing another example of the control of the transmission of interaction information and virtual space information by the virtual reality service server 50. The processing shown in the flowchart in Figure 14 is repeatedly executed, for example, while a user is entering a virtual space.

[0092] First, the virtual reality service server 50 identifies the virtual space in which the user is currently present (S300). Next, the virtual reality service server 50 determines whether or not it has obtained dialogue information from the user (S302). If it determines that it has obtained dialogue information from the user, the virtual reality service server 50 sends the dialogue information and virtual space information to the information processing device 100 (S304). On the other hand, if it determines that it has not obtained dialogue information from the user, the virtual reality service server 50 determines whether or not a certain period of time has elapsed since the last transmission of the dialogue information and virtual space information (S306). If it determines that a certain period of time has elapsed since the last transmission of the dialogue information and virtual space information, the virtual reality service server 50 sends notification information indicating the elapsed period to the information processing device 100 (S308). This completes the processing of this flowchart.

[0093] According to the processing in this flowchart, even if the user does not input additional dialogue information, the agent avatar AA can be made to perform an appropriate action based on the situation. For example, if the user avatar UA is acting independently and ignoring the agent avatar AA, the agent avatar AA can be made to perform actions such as calling out to the user, canceling the dialogue, or hiding.

[0094] As described above, this embodiment acquires dialogue information relating to interaction with an avatar in a single virtual space identified from multiple virtual spaces, and virtual space information relating to the virtual space. A prompt for input to the AI ​​model is generated from at least the dialogue information and the virtual space information. By inputting the prompt to the AI ​​model, control information relating to the actions of the agent interacting with the avatar in the virtual space is acquired, and the control information is output. This enables flexible support by the agent avatar according to user information and virtual space information. [Explanation of symbols]

[0095] 10 Terminal devices 50 Virtual reality service servers 100 Information Processing Devices 110 Communications Department 120 First acquisition part 130 Generation part 140 Second acquisition part 150 Output section 160 Storage section 162 AI Models

Claims

1. A first acquisition unit acquires dialogue information relating to interaction with an avatar in one virtual space identified from multiple virtual spaces, and virtual space information relating to the said virtual space. A generation unit that generates prompts for input to the AI ​​model from at least the dialogue information and the virtual space information, A second acquisition unit acquires control information regarding the actions of an agent interacting with the avatar in the virtual space by inputting the aforementioned prompt into the AI ​​model. It comprises an output unit that outputs the aforementioned control information, The prompt includes instruction information to cause the AI ​​model to output an action that the agent should perform. Information processing system.

2. The system further includes a control unit that controls the operation of the agent in the virtual space based on the control information. The information processing system according to claim 1.

3. The agent's actions include guiding the avatar to a predetermined location or object in the virtual space, and providing an explanation regarding the predetermined location or object. The information processing system according to claim 1.

4. A first acquisition unit that acquires dialogue information relating to an interaction with an avatar in one virtual space identified from a plurality of virtual spaces, and virtual space information relating to the said virtual space, A generation unit that generates prompts for input to the AI ​​model from at least the dialogue information and the virtual space information, A second acquisition unit acquires control information regarding the actions of an agent interacting with the avatar in the virtual space by inputting the aforementioned prompt into the AI ​​model. It comprises an output unit that outputs the aforementioned control information, The prompt is instruction information for causing the agent to output control information regarding whether or not the aforementioned action needs to be performed. The operation includes the movement of the agent in the virtual space, The aforementioned movement is the agent moving in accordance with the avatar. Information processing system.

5. The aforementioned movement is the agent moving the avatar to a predetermined location or predetermined object immediately and simultaneously with the agent. The information processing system according to claim 4.

6. The aforementioned operation includes, after the agent has moved in the virtual space, checking whether the avatar exists within a predetermined range of the agent, and if the avatar does not exist within the predetermined range, either interacting with the avatar, ceasing operations that include interacting with the avatar, or hiding the agent. The information processing system according to claim 4.

7. A first acquisition unit acquires dialogue information relating to interaction with an avatar in one virtual space identified from multiple virtual spaces, and virtual space information relating to the said virtual space. A generation unit that generates prompts for input to the AI ​​model from at least the dialogue information and the virtual space information, A second acquisition unit acquires control information regarding the actions of an agent interacting with the avatar in the virtual space by inputting the aforementioned prompt into the AI ​​model. It comprises an output unit that outputs the aforementioned control information, The second acquisition unit inputs one or more action candidates as candidates for the agent's actions into the prompt, thereby acquiring the action candidate to be executed from among the one or more action candidates as control information. Information processing system.

8. The one or more suggested actions are set to be the same as actions that the avatar can perform in the virtual space. The second acquisition unit inputs instruction information to the prompt that instructs the AI ​​model to select an action candidate to be executed from the one or more action candidates. The information processing system according to claim 7.

9. The one or more action candidates mentioned above include speech and movement by the agent in the virtual space, The utterance is a description of a predetermined location or object in the virtual space, and the movement is movement to the predetermined location or object in the virtual space. The information processing system according to claim 7.

10. The aforementioned virtual space is a user-generated virtual space created by a user of the metaverse application that generates the virtual space, The second acquisition unit acquires the type and content of the agent's operation as control information by inputting user-related data relating to the user-generated virtual space into the prompt to the AI ​​model. The information processing system according to claim 1.

11. The user-related data includes instruction data regarding the agent settings, The second acquisition unit inputs the instruction data into the prompt and inputs it to the AI ​​model. The information processing system according to claim 10.

12. The second acquisition unit inputs the location information of a predetermined point and / or a predetermined object in the virtual space, and text information relating to the predetermined point and / or predetermined object, which have been stored in advance, into the AI ​​model as virtual space information in the prompt. The information processing system according to claim 1.

13. The virtual space information includes at least one of the following: location information, shape information, and identification information of objects placed in the virtual space; descriptive information relating to a predetermined point and / or predetermined object in the virtual space; user information relating to the user operating the avatar; and location information of avatars that exist in the virtual space and are operated by other users. The information processing system according to claim 1.

14. When the settings of the virtual space are changed, the second acquisition unit inputs some or all of the virtual space information relating to the virtual space whose settings have been changed into the AI ​​model. The information processing system according to claim 13.

15. The system further includes a configuration unit for setting up multiple agents for at least one of the virtual spaces, The configuration unit, when a user enters one of the virtual spaces, assigns at least one agent from among the multiple agents to the virtual space. The information processing system according to claim 1.

16. The second acquisition unit inputs virtual space information relating to one of the virtual spaces into the AI ​​model when the user enters one of the one or more virtual spaces. The information processing system according to claim 1.

17. The control unit, in response to a predetermined event occurring in the virtual space, switches the operation of the agent to operation by the administrator of the virtual space, or enables it to accept operation by the administrator. The information processing system according to claim 2.

18. The first acquisition unit acquires the dialogue information and the virtual space information from the metaverse application that generates the virtual space via the Application Programming Interface (API). The information processing system according to claim 1.

19. The output unit outputs the control information to the metaverse application via the API. The metaverse application controls the operation of the agent in the virtual space based on the control information. The information processing system according to claim 18.

20. A first acquisition unit acquires dialogue information relating to interaction with an avatar in one virtual space identified from multiple virtual spaces, and virtual space information relating to the said virtual space. A generation unit that generates prompts for input to the AI ​​model from at least the dialogue information and the virtual space information, A second acquisition unit acquires control information regarding the actions of an agent interacting with the avatar in the virtual space by inputting the aforementioned prompt into the AI ​​model. The metaverse application that generates the virtual space includes an output unit that outputs the control information, The system includes a display unit that displays the operation of the agent in the virtual space, which is controlled based on the control information on the metaverse application, The prompt includes instruction information to cause the AI ​​model to output an action that the agent should perform. Information processing system.

21. A unit that identifies one virtual space from multiple virtual spaces, A first acquisition unit that acquires dialogue information relating to interaction with an avatar in the virtual space and virtual space information relating to the virtual space, A generation unit that generates prompts for input to the AI ​​model from at least the dialogue information and the virtual space information, A second acquisition unit acquires control information regarding the actions of an agent interacting with the avatar in the virtual space by inputting the aforementioned prompt into the AI ​​model. The system includes a control unit that controls the operation of the agent in the virtual space based on the acquired control information, The prompt includes instruction information to cause the AI ​​model to output an action that the agent should perform. Information processing system.

22. Computers The system acquires dialogue information relating to interactions with an avatar in one virtual space identified from multiple virtual spaces, and virtual space information relating to the said virtual space. At least from the dialogue information and the virtual space information, a prompt for input to the AI ​​model is generated, By inputting the aforementioned prompt into the AI ​​model, control information regarding the actions of the agent interacting with the avatar in the virtual space is obtained. The control information is output, The prompt includes instruction information to cause the AI ​​model to output an action that the agent should perform. Information processing methods.

23. On the computer, The system obtains dialogue information relating to an interaction with an avatar in one virtual space identified from multiple virtual spaces, and virtual space information relating to the said virtual space. At least from the dialogue information and the virtual space information, prompts for input to the AI ​​model are generated. By inputting the aforementioned prompt into the AI ​​model, control information regarding the actions of the agent interacting with the avatar in the virtual space is obtained. The control information is output, The prompt includes instruction information to cause the AI ​​model to output an action that the agent should perform. program.