Information processing method and information processing device
The system addresses the challenge of generating diverse and natural avatar motions by using a large-scale language model and generative AI to regenerate and update motion banks, improving avatar interactions.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2026-04-09
AI Technical Summary
Existing systems face challenges in generating a large number of diverse and natural motions for avatars due to high expertise and time requirements, leading to monotonous and unnatural responses.
An information processing system utilizing a large-scale language model and generative AI to generate and update motion banks for avatars, ensuring diverse and natural interactions by regenerating motions using prompts.
The system prevents repetitive motions by regenerating and updating avatar motions, enhancing the naturalness and diversity of avatar responses.
Smart Images

Figure JP2025028646_09042026_PF_FP_ABST
Abstract
Description
Information Processing Method and Information Processing Apparatus
[0001] The present disclosure relates to an information processing method and an information processing apparatus.
[0002] In recent years, large language models that are being rapidly developed can return more complex responses to a wide range of tasks. Therefore, it has been considered to perform more complex interactions between a user and an avatar by controlling the avatar using a large language model.
[0003] For example, Patent Document 1 below discloses an information processing apparatus that assigns a more natural motion reflecting the character's emotion to the character's avatar by determining the emotion based on the analysis result of the character's utterance.
[0004] International Publication No. 2020 / 170441
[0005] However, since it requires high expertise and a long production time to generate motion data for controlling the motion of an avatar, it has been difficult to prepare a large number of motions of the avatar. On the other hand, when the number of motion patterns of the avatar is small, the response of the avatar becomes monotonous, which may make the user feel the unnaturalness and artificiality of the avatar.
[0006] Therefore, there has been a demand to make the avatar execute motions with fewer repetitions and more patterns.
[0007] According to the present disclosure, there is provided a computer-implemented information processing method including generating a prompt for generating a motion of an avatar that interacts with a user using a large language model, generating the motion of the avatar based on the prompt using a motion generation AI, and updating a motion bank in which the motion of the avatar is registered with the generated motion.
[0008] Furthermore, according to this disclosure, an information processing device is provided, comprising: an avatar control unit that controls the motion of an avatar interacting with a user based on an action plan generated by a large-scale language model; and a bank update unit that updates a motion bank in which the avatar's motion is registered, using the motion regenerated by a generating AI based on prompts further generated by the large-scale language model.
[0009] This is a block diagram showing the overall configuration of an information processing system according to one embodiment of the present disclosure. This is a block diagram showing an example of the configuration of an information processing device included in the information processing system. This is a sequence diagram showing a first example of operation of the information processing system according to the same embodiment. This is a sequence diagram showing a second example of operation of the information processing system according to the same embodiment. This is a sequence diagram showing a third example of operation of the information processing system according to the same embodiment. This is a block diagram showing an example of the hardware configuration of an information processing device.
[0010] Preferred embodiments of this disclosure will be described in detail below with reference to the attached drawings. In this specification and the drawings, components having substantially the same functional configuration are denoted by the same reference numerals, and redundant descriptions will be omitted.
[0011] The explanation will proceed in the following order: 1. Overall configuration 2. Configuration example 3. Operation example 3.1. First operation example 3.2. Second operation example 3.3. Third operation example 4. Hardware configuration example
[0012] <1. Overall Configuration> First, the overall configuration of an information processing system according to one embodiment of the present disclosure will be described with reference to Figure 1. Figure 1 is a block diagram showing the overall configuration of the information processing system 1 according to the present embodiment.
[0013] As shown in Figure 1, the information processing system 1 includes an information processing device 10, a large-scale language model 20, and a generative AI (Artificial Intelligence) 30. The information processing device 10, the large-scale language model 20, and the generative AI (Artificial Intelligence) 30 are connected to each other via a network 40 so that they can communicate data with one another. The network 40 is an information communication network consisting of computers connected by wired or wireless connections, such as the Internet of Things.
[0014] The information processing device 10 is an information processing device that displays an avatar AA and accepts input from the user. The avatar AA is a digital character or an AI character, and is rendered using 2D or 3D computer graphics (CG). For example, the avatar AA may be a humanoid or animal-type character whose speech and actions are controlled by AI. The user can communicate with the avatar AA in a two-way manner through the information processing device 10. For example, the user can have a conversation with the avatar AA via voice or text, or engage in non-verbal communication such as physical contact. The information processing device 10 may be, for example, a head-mounted display (HMD) device that provides the user with an XR (eXtended Reality) experience.
[0015] The large-scale language model 20 is a computer language model composed of a neural network with tens of millions to billions of parameters. The large-scale language model 20 can be trained using a vast amount of unlabeled text through self-supervised or semi-supervised learning, enabling it to output appropriate responses to diverse inputs. The large-scale language model 20 may be built, for example, on computing resources on a network 40 such as a cloud service.
[0016] The large-scale language model 20 generates an action plan, including the motion of the avatar AA, based on the context between the avatar AA and the user. The context between the avatar AA and the user includes, for example, information representing the personality and settings of the avatar AA, information representing the status of the avatar AA, information representing the environment in which the avatar AA exists, and information identifying user input to the avatar AA. By controlling the avatar AA with the action plan generated based on this information, the avatar AA can provide natural responses to the user that reflect the personality of the avatar AA.
[0017] The motions of the avatar AA included in the avatar AA's action plan are selected from, for example, a motion bank stored in the information processing device 10. The motion bank is a database in which multiple motions that can be executed by the avatar AA are registered. The large-scale language model 20 can generate an action plan for the avatar AA, including motions, by selecting motions to be executed by the avatar AA from the motions registered in the motion bank.
[0018] Furthermore, the large-scale language model 20 generates prompts for the generation AI 30, described later, to generate motions for the avatar AA. Specifically, the large-scale language model 20 may generate prompts for the generation AI 30 to regenerate motions included in the action plan of the avatar AA. When the generated prompts are input to the generation AI 30, motions included in the action plan of the avatar AA (i.e., motions executed by the avatar AA) are regenerated.
[0019] The generative AI 30 is a generative model composed of a deep neural network that generates motion for avatar AA. The large-scale language model 20 may be constructed, for example, on computing resources on a network 40 such as a cloud service. The generative AI 30 can generate motion for avatar AA based on prompts generated by the large-scale language model 20. The generated motion for avatar AA is registered in a motion bank stored in the information processing device 10, etc.
[0020] As described above, the generating AI 30 can regenerate motions performed by the avatar AA based on prompts generated by the large-scale language model 20. It is assumed that the motions regenerated by the generating AI 30 will not be identical to the motions before regeneration. Therefore, the motion bank can prevent the motions performed by the avatar AA from becoming repetitive patterns by replacing the original motions with the regenerated motions and registering them.
[0021] Furthermore, the prompts generated by the large-scale language model 20 may include information about the personality and settings of the avatar AA, and information representing the status of the avatar AA (for example, technical status related to the avatar AA's abilities or skills). By including this information in the prompts, the generating AI 30 can generate motions that are more consistent with the personality and settings of the avatar AA.
[0022] Furthermore, the large-scale language model 20 may generate a prompt for the generation AI 30 to regenerate a motion when the avatar AA's action plan includes that motion a predetermined number of times. If the generation AI 30 regenerates the motion every time the avatar AA executes a motion, the generation AI 30 may not be able to regenerate the motion in time. Therefore, the large-scale language model 20 may reduce the frequency of motion regeneration by generating a prompt for the avatar AA to regenerate a motion when that motion has been executed a predetermined number of times.
[0023] The information processing system 1 according to this embodiment can generate motions for the avatar AA using a generation AI based on prompts generated by the large-scale language model 20, and update the motion bank with the generated motions. This allows the information processing system 1 to prevent the same motion from being repeatedly executed by the avatar AA and to have the avatar AA execute newly generated motions.
[0024] <2. Configuration Example> Next, with reference to Figure 2, a configuration example of the information processing device 10 included in the information processing system 1 according to this embodiment will be described. Figure 2 is a block diagram showing a configuration example of the information processing device 10.
[0025] As shown in Figure 2, the information processing device 10 includes an information acquisition unit 110, a reflection judgment unit 120, an instruction generation unit 130, an avatar control unit 140, a bank update unit 150, and an information management unit 160. The information management unit 160 includes an AI state management unit 161, a motion bank 162, and a location information management unit 163.
[0026] The information acquisition unit 110 acquires various types of information. Specifically, the information acquisition unit 110 may acquire information that identifies an action from the user to avatar AA, information that represents the environment in which the user or avatar AA exists, and information that represents the location of the user or avatar AA. The information acquisition unit 110 may use a camera, microphone, mouse, keyboard, touch panel, button, or switch to acquire information that identifies an action from the user to avatar AA. In addition, the information acquisition unit 110 may use a camera or distance sensor to acquire information that represents the environment in which the user or avatar AA exists, and information that represents the location of the user or avatar AA.
[0027] The reflection determination unit 120 determines whether the action taken by the user to the avatar AA, as acquired by the information acquisition unit 110, is an action that responds via reflection. Reflection means controlling the movement of the avatar AA based on a pre-set action plan without waiting for the large-scale language model 20 to generate an action plan. For example, reflection means that if the avatar AA is touched from a direction it is not looking at, or if a loud noise occurs, the avatar AA will flinch or jump up.
[0028] For example, the reflection determination unit 120 may determine that an action from the user to the avatar AA is an action that responds by reflection if it is a predetermined action that responds by reflection (for example, a sound louder than a threshold, or contact from outside the avatar AA's field of view). Alternatively, the reflection determination unit 120 may determine that an action from the user to the avatar AA is an action that responds by reflection if it is not a predetermined action that generates an action plan in the large-scale language model 20 (an action within the avatar AA's field of view).
[0029] If the reflection determination unit 120 determines that an action from the user to avatar AA is an action that responds by reflection, the subsequent avatar control unit 140 may control the movement of avatar AA based on a pre-set action plan. For example, the avatar control unit 140 may control the movement of avatar AA so that it flinches, jumps up, or shouts loudly.
[0030] The instruction generation unit 130 instructs the large-scale language model 20 to generate an action plan in response to an action from the user to avatar AA. Specifically, the instruction generation unit 130 may instruct the large-scale language model 20 to generate an action plan by transmitting to the large-scale language model 20 information that identifies the action from the user to avatar AA, information that represents the environment in which the user or avatar AA exists, information that represents the location of the user or avatar AA, status information of avatar AA, the current action plan of avatar AA, and a list of motions registered in the motion bank 162.
[0031] Information identifying user actions on avatar AA, and information representing the environment in which the user or avatar AA exists, are acquired by the information acquisition unit 110. Information representing the location of the user or avatar AA is managed by the location information management unit 163, and the status information of avatar AA is managed by the AI state management unit 161. In addition, the current action plan of avatar AA is acquired from the avatar control unit 140.
[0032] Furthermore, if the instruction generation unit 130 is notified by the avatar control unit 140 that the stock of action plans for avatar AA is below a threshold, it may instruct the large-scale language model 20 to generate an action plan. In such a case, the instruction generation unit 130 may instruct the large-scale language model 20 to generate an action plan by transmitting to the large-scale language model 20 information representing the environment in which the user or avatar AA exists, information representing the location of the user or avatar AA, status information of the avatar AA, and a list of motions registered in the motion bank 162.
[0033] The avatar control unit 140 controls the movement of avatar AA according to the action plan. Specifically, the avatar control unit 140 may control the motion and speech content of avatar AA according to the action plan generated by the large-scale language model 20. Furthermore, if a new action plan is generated by the large-scale language model 20, the avatar control unit 140 may update the action plan by inserting the newly generated action plan into the current action plan, and then control the motion and speech content of avatar AA according to the updated action plan.
[0034] Furthermore, if the reflection determination unit 120 determines that an action from the user to avatar AA is an action that responds by reflection, the avatar control unit 140 may control the movement of avatar AA based on a pre-set action plan. For example, the avatar control unit 140 may control the movement of avatar AA so that it flinches, jumps up, or shouts loudly, based on a pre-set action plan.
[0035] The bank update unit 150 updates the motion bank 162 with motions generated by the generating AI 30. Specifically, the bank update unit 150 updates the motions of avatar AA included in the action plan generated by the large-scale language model 20 with motions generated by the generating AI 30 based on prompts generated by the large-scale language model 20. As a result, the bank update unit 150 can update the motions executed by avatar AA with motions regenerated by the generating AI 30. Therefore, the information processing system 1 can prevent the same motion from being repeatedly executed by avatar AA by having avatar AA execute newly generated motions.
[0036] The information management unit 160 includes an AI state management unit 161, a motion bank 162, and a location information management unit 163. The information management unit 160 may be, for example, a data storage device configured as an example of a storage unit of the information processing device 10. The information management unit 160 may be composed of, for example, a magnetic storage device such as an HDD (Hard Disk Drive), a semiconductor storage device, an optical storage device, or a magneto-optical storage device.
[0037] The information management unit 160 may be located on the information processing device 10, or on a server connected to the information processing device 10 via the network 40. When the information management unit 160 is located on the information processing device 10, data retrieval from the information management unit 160 is faster, allowing the information processing device 10 to further enhance the real-time control of the avatar AA. On the other hand, when the information management unit 160 is located on a server, multiple information processing devices can access the data from the information management unit 160, allowing the information processing device 10 to share and interact with the same avatar AA with other information processing devices.
[0038] The AI State Management Unit 161 stores and manages the status of the avatar AA. Specifically, the AI State Management Unit 161 may store and manage the emotional status of the avatar AA, such as joy, anger, sadness, or boredom; the physiological status of the avatar AA, such as hunger or sleepiness; or the technical status related to the avatar AA's abilities or skills. The emotional and physiological status of the avatar AA is updated each time the avatar AA's movements are controlled by the action plan.
[0039] Motion bank 162 registers and manages motions that avatar AA can perform. Motions registered in motion bank 162 can be updated with motions generated by generation AI 30. The motion of avatar AA may be defined, for example, by the position of each bone representing the avatar AA's skeleton and the angles of the joints connecting each bone. The facial movements of avatar AA included in the motion of avatar AA may be defined, for example, by the shape and position of the mesh objects representing the avatar AA's face and mouth.
[0040] The location information management unit 163 manages information representing the location of avatar AA. For example, the location information management unit 163 may manage information representing the location of avatar AA in the environment where the user or avatar AA exists, as acquired by the information acquisition unit 110. The location of avatar AA changes, for example, based on the movement path of avatar AA when avatar AA moves according to an action plan.
[0041] With the above configuration, the information processing device 10 can control the motion of the avatar AA according to the action plan generated by the large-scale language model 20. The information processing device 10 can have the motion executed by the avatar AA according to the action plan regenerated by the generation AI 30, and update the motions registered in the motion bank 162 with the regenerated motions. This makes it possible for the information processing device 10 to prevent the same motion from being repeatedly executed by the avatar AA.
[0042] <3. Operation Example> Next, an operation example of the information processing system 1 according to this embodiment will be described with reference to Figures 3 to 5.
[0043] (3.1. First operation example) FIG. 3 is a sequence diagram showing a first operation example of the information processing system 1 according to the present embodiment.
[0044] As shown in FIG. 3, first, the information processing device 10 detects an action from the user to the avatar AA (S101). For example, the information processing device 10 may detect an action of the user facing the avatar AA or a user's utterance including the name of the avatar AA as a call (action) from the user to the avatar AA.
[0045] Next, the information processing device 10 instructs the large language model 20 to generate an action plan for a response to the user's action (S103). For example, the information processing device 10 may send the content of the utterance from the user to the avatar AA, the status information of the avatar AA, the information representing the environment where the user and the avatar AA exist, the information representing the positions of the user and the avatar AA, the current action plan of the avatar AA, and the list of motions registered in the motion bank 162 to the large language model 20 to instruct the generation of an action plan for a response to the user's action.
[0046] As a result, an action plan for the avatar AA is generated by the large language model 20 (S105). The generated action plan for the avatar AA includes the motion of the avatar AA to be executed, the utterance content, the interval until the next motion, and the changes in the emotion and physiological status of the avatar AA due to the execution of the motion. The generated action plan for the avatar AA is sent to the information processing device 10 (S107).
[0047] Subsequently, the information processing device 10 updates the action plan of the avatar AA according to the received action plan (S109). Specifically, the information processing device 10 updates the action plan of the avatar AA by interrupting the action plan generated by the large language model 20 with the current action plan. Thereafter, the information processing device 10 controls the movement of the avatar AA according to the updated action plan (S111).
[0048] On the other hand, after generating the action plan in step S105, the large language model 20 generates a prompt for causing the generation AI 30 to regenerate the motion included in the action plan (S131). For example, the large language model 20 may generate a prompt including information representing the personality and settings of the avatar AA and information representing the motion included in the action plan. The generated prompt is transmitted to the generation AI 30 (S133).
[0049] Thereby, the generation AI 30 regenerates the motion of the avatar AA based on the prompt (S135). Since the motions generated by the generation AI 30 are not the same, the regenerated motion is different from the motion executed by the avatar AA in step S111. The regenerated motion is transmitted to the information processing device 10 (S137).
[0050] Thereafter, the information processing device 10 updates the motion registered in the motion bank 162 (that is, the motion executed by the avatar AA) with the motion regenerated by the generation AI 30 (S139). Thereby, the motion regenerated by the generation AI 30 is registered in the motion bank 162.
[0051] According to the above first operation example, the information processing system 1 can regenerate the motion executed by the avatar AA by the generation AI 30 and update the motion registered in the motion bank 162 with the regenerated motion. Therefore, the information processing system 1 can prevent the avatar AA from making the same response with the same motion to the action from the user.
[0052] (3.2. Second operation example) FIG. 4 is a sequence diagram showing a second operation example of the information processing system 1 according to the present embodiment.
[0053] As shown in Figure 4, first, the information processing device 10 detects that the stock of action plans for avatar AA is below a threshold (S201). That is, the information processing device 10 detects when there is no action from the user and the stock of action plans for avatar AA is about to run out (for example, when the stock of action plans for avatar AA is 2 or less).
[0054] Next, the information processing device 10 instructs the large-scale language model 20 to generate an action plan for the state of waiting for user action (S203). For example, the information processing device 10 may instruct the large-scale language model 20 to generate an action plan for the state of waiting for user action by sending status information of avatar AA, information representing the environment in which the user and avatar AA exist, information representing the positions of the user and avatar AA, and a list of motions registered in the motion bank 162.
[0055] As a result, the large-scale language model 20 generates an action plan for the avatar AA (S105). The generated action plan for the avatar AA includes the motions to be performed by the avatar AA, the content of the utterances, the interval until the next motion, and the changes in the avatar AA's emotions and physiological status due to the execution of the motions. The generated action plan for the avatar AA is transmitted to the information processing device 10 (S107).
[0056] Next, the information processing device 10 updates the action plan of avatar AA with the received action plan (S109). Specifically, the information processing device 10 updates the action plan of avatar AA by adding the action plan generated by the large-scale language model 20 to the end of the current action plan. After that, the information processing device 10 controls the movements of avatar AA according to the updated action plan (S111).
[0057] Meanwhile, after the action plan is generated in step S105, the large-scale language model 20 generates a prompt to cause the generation AI 30 to regenerate the motions included in the action plan (S131). For example, the large-scale language model 20 may generate a prompt that includes information representing the personality and settings of the avatar AA, and information representing the motions included in the action plan. The generated prompt is sent to the generation AI 30 (S133).
[0058] As a result, the generation AI 30 regenerates the motion of the avatar AA based on the prompt (S135). Since the motion generated by the generation AI 30 will not be the same, the regenerated motion will be different from the motion that the avatar AA performed in step S111. The regenerated motion is transmitted to the information processing device 10 (S137).
[0059] Subsequently, the information processing device 10 updates the motions registered in the motion bank 162 (i.e., the motions executed by the avatar AA) with the motions regenerated by the generating AI 30 (S139). As a result, the motions regenerated by the generating AI 30 are registered in the motion bank 162.
[0060] According to the second example of operation described above, the information processing system 1 can regenerate the motion executed by the avatar AA using the generation AI 30, and update the motions registered in the motion bank 162 with the regenerated motions. Therefore, the information processing system 1 can prevent the avatar AA, which is in a state of waiting for user action, from repeatedly executing the same motion.
[0061] (3.3. Third Operation Example) Figure 5 is a sequence diagram showing a third operation example of the information processing system 1 according to this embodiment.
[0062] As shown in Figure 5, first, the information processing device 10 detects an action from the user to avatar AA (S101). For example, the information processing device 10 may detect contact from the user to avatar AA as an action from the user to avatar AA.
[0063] Here, the information processing device 10 determines whether the detected action is an action to be responded to by reflection (S301). For example, the information processing device 10 may determine that an action is to be responded to by reflection if the action from the user to the avatar AA is a predetermined action that should be responded to by reflection (for example, a sound louder than a threshold, or contact from outside the field of view of the avatar AA). Alternatively, the information processing device 10 may determine that an action is to be responded to by reflection if the action from the user to the avatar AA is anything other than a predetermined action for which the large-scale language model 20 should generate an action plan (an action within the field of view of the avatar AA).
[0064] If the detected action is determined to be an action that responds by reflection, the information processing device 10 controls the movement of avatar AA based on a pre-set action plan and causes avatar AA to perform a pre-set motion (S303). The pre-set action plan includes, for example, motions such as flinching, jumping up, being surprised, or shouting loudly.
[0065] Subsequently, the information processing device 10 instructs the large-scale language model 20 to generate an action plan in response to the user's action (S103). For example, the information processing device 10 may instruct the large-scale language model 20 to generate an action plan in response to the user's action by transmitting the content of the action from the user to avatar AA, status information of avatar AA, information representing the environment in which the user and avatar AA exist, information representing the positions of the user and avatar AA, and a list of motions registered in the motion bank 162.
[0066] As a result, the large-scale language model 20 generates an action plan for the avatar AA (S105). The generated action plan for the avatar AA includes the motions to be performed by the avatar AA, the content of the utterances, the interval until the next motion, and the changes in the avatar AA's emotions and physiological status due to the execution of the motions. The generated action plan for the avatar AA is transmitted to the information processing device 10 (S107).
[0067] Next, the information processing device 10 updates the action plan of avatar AA with the received action plan (S109). Specifically, the information processing device 10 updates the action plan of avatar AA by adding an action plan generated by the large-scale language model 20 after the pre-set action plan for reflective responses. After that, the information processing device 10 controls the movements of avatar AA according to the updated action plan (S111).
[0068] Meanwhile, after the action plan is generated in step S105, the large-scale language model 20 generates a prompt (S331) for the generating AI 30 to regenerate the motions included in the action plan and the motions executed in step S303. For example, the large-scale language model 20 may generate a prompt that includes information representing the personality and settings of the avatar AA, the motions included in the action plan, and information representing the motions executed in step S303. The generated prompt is sent to the generating AI 30 (S133).
[0069] As a result, the generation AI 30 regenerates the motion of the avatar AA based on the prompt (S135). Since the motion generated by the generation AI 30 will not be the same, the regenerated motion will be different from the motion that the avatar AA performed in step S111. The regenerated motion is transmitted to the information processing device 10 (S137).
[0070] Subsequently, the information processing device 10 updates the motions registered in the motion bank 162 (i.e., the motions executed by the avatar AA) with the motions regenerated by the generating AI 30 (S139). As a result, the motions regenerated by the generating AI 30 are registered in the motion bank 162.
[0071] According to the third example of operation described above, the information processing system 1 can regenerate the motion executed by the avatar AA using the generation AI 30, and update the motion registered in the motion bank 162 with the regenerated motion. Therefore, even if the avatar AA responds to an action from the user by reflection, the information processing system 1 can prevent the motion of the avatar AA that responded by reflection from becoming a repetition of the same motion.
[0072] <4. Hardware Configuration Example> Next, with reference to Figure 6, the hardware configuration of the information processing device 10 included in the information processing system 1 according to this embodiment will be described. Figure 6 is a block diagram showing an example of the hardware configuration of the information processing device 10.
[0073] The functions of the information processing device 10 can be realized through the cooperation of software and the hardware described below. The functions of the reflection determination unit 120, the instruction generation unit 130, the avatar control unit 140, and the bank update unit 150 may be performed by, for example, the CPU 901. The functions of the information acquisition unit 110 may be performed by, for example, the input device 906. The functions of the information management unit 160 may be performed by, for example, the storage device 908.
[0074] As shown in Figure 6, the information processing device 10 includes a CPU (Central Processing Unit) 901, a ROM (Read Only Memory) 902, and a RAM (Random Access Memory) 903.
[0075] Furthermore, the information processing device 10 may further include a host bus 904a, a bridge 904, an external bus 904b, an interface 905, an input device 906, an output device 907, a storage device 908, a drive 909, a connection port 910, or a communication device 911. The information processing device 10 may have a processing circuit such as a DSP (Digital Signal Processor) or an ASIC (Application Specific Integrated Circuit) in place of, or together with, the CPU 901.
[0076] The CPU 901 functions as an arithmetic processing unit or control unit, and controls the operation within the information processing unit 10 according to various programs recorded on the ROM 902, RAM 903, storage device 908, or removable recording medium installed in the drive 909. The ROM 902 stores programs used by the CPU 901, as well as arithmetic parameters, etc. The RAM 903 temporarily stores programs used in the execution of the CPU 901, as well as parameters used during its execution.
[0077] The CPU 901, ROM 902, and RAM 903 are interconnected by a host bus 904a capable of high-speed data transmission. The host bus 904a is connected to an external bus 904b, such as a PCI (Peripheral Component Interconnect / Interface) bus, via a bridge 904, and the external bus 904b is connected to various components via an interface 905.
[0078] The input device 906 is, for example, a device that receives input from the user, such as a mouse, keyboard, touch panel, button, switch, or lever. The input device 906 may also be a microphone that detects the user's voice. The input device 906 may also be, for example, a remote control device that uses infrared or other radio waves, or an external connection device that corresponds to the operation of the information processing device 10.
[0079] The input device 906 further includes an input control circuit that outputs an input signal generated based on information input by the user to the CPU 901. By operating the input device 906, the user can input various types of data to the information processing device 10 or instruct it to perform processing operations.
[0080] The output device 907 is a device capable of presenting information acquired or generated by the information processing device 10 to the user visually or audibly. The output device 907 may be a display device such as an LCD (Liquid Crystal Display), PDP (Plasma Display Panel), OLED (Organic Light Emitting Diode) display, hologram, or projector, or it may be an audio output device such as a speaker or headphones, or a printing device such as a printer. The output device 907 can output the information obtained by processing in the information processing device 10 as images such as text or pictures, or as sound such as voice or sound.
[0081] The storage device 908 is a data storage device configured as an example of the storage unit of the information processing device 10. The storage device 908 may be composed of, for example, a magnetic storage device such as an HDD (Hard Disk Drive), a semiconductor storage device, an optical storage device, or a magneto-optical storage device. The storage device 908 can store programs executed by the CPU 901, various data, or various data acquired from external sources.
[0082] The drive 909 is a read or write device for removable recording media such as magnetic disks, optical disks, magneto-optical disks, or semiconductor memory, and is built into or attached to the information processing device 10. For example, the drive 909 can read information recorded on the installed removable recording media and output it to the RAM 903. The drive 909 can also write data to the installed removable recording media.
[0083] The connection port 910 is a port for directly connecting an external device to the information processing device 10. The connection port 910 may be, for example, a USB (Universal Serial Bus) port, an IEEE 1394 port, or a SCSI (Small Computer System Interface) port. Alternatively, the connection port 910 may be an RS-232C port, an optical audio terminal, or an HDMI (High-Definition Multimedia Interface) port. By connecting the connection port 910 to an external device, various types of data can be transmitted and received between the information processing device 10 and the external device.
[0084] The communication device 911 is a communication interface, for example, composed of a communication device for connecting to the network 40. The communication device 911 may be, for example, a wired or wireless LAN (Local Area Network), Wi-Fi (registered trademark), Bluetooth (registered trademark), or a WUSB (Wireless USB) communication card. Alternatively, the communication device 911 may be a router for optical communication, an ADSL (Asymmetric Digital Subscriber Line) router, or a modem for various types of communication.
[0085] The communication device 911 can send and receive signals, etc., using a predetermined protocol such as TCP / IP, for example, with the Internet or other communication devices. The network 40 connected to the communication device 911 is a network connected by wire or wireless, and may be, for example, an Internet communication network, a home LAN, an infrared communication network, a radio wave communication network, or a satellite communication network.
[0086] Furthermore, it is possible to create a program that enables the CPU 901, ROM 902, RAM 903, and other hardware built into the computer to perform functions equivalent to those of the information processing device 10 described above. A computer-readable recording medium on which this program is stored can also be provided.
[0087] While preferred embodiments of the present disclosure have been described in detail above with reference to the attached drawings, the technical scope of the present disclosure is not limited to such examples. It is clear to any person with ordinary skill in the art of the present disclosure that various modifications or alterations may be conceived within the scope of the technical ideas described in the claims, and these will naturally also fall within the technical scope of the present disclosure.
[0088] In the above embodiment, an example was shown where the information input to the large-scale language model 20 is text data, but this technology is not limited to such an example. For example, the information input to the large-scale language model 20 may be data such as images, videos, or audio. Also, in the above embodiment, an example was shown where the speech content and actions of the avatar AA are controlled by AI, but this technology is not limited to such an example. For example, the speech content and actions of the avatar AA may be controlled by another user.
[0089] Furthermore, among the processes described in the embodiments of this disclosure described above, all or part of the processes described as being performed automatically may be performed manually, or all or part of the processes described as being performed manually may be performed automatically by known methods. In addition, the processing procedures, specific names, and information including various data and parameters shown in the above document and drawings may be changed at will unless otherwise specified. For example, the various information shown in each figure is not limited to the information shown.
[0090] Furthermore, the components of each illustrated device are functionally conceptual and do not necessarily need to be physically configured as shown. In other words, the specific forms of distribution and integration of each device are not limited to those shown, and all or part of them can be functionally or physically distributed and integrated in any unit according to various loads and usage conditions.
[0091] Furthermore, the embodiments of this disclosure described above can be combined as appropriate in areas where the processing content is not contradictory. Also, the order of each step shown in the sequence diagram or flowchart of this embodiment can be changed as appropriate. For example, each step may be processed chronologically, repeatedly, or partially in parallel.
[0092] Furthermore, the effects described herein are merely descriptive or illustrative and not limiting. In other words, the technology relating to this disclosure may produce other effects that will be apparent to those skilled in the art from the description herein, in addition to or in lieu of the effects described herein.
[0093] Furthermore, the following configurations also fall within the technical scope of this disclosure: (1) A computer-based information processing method comprising: generating a prompt using a large-scale language model for generating motion for an avatar that interacts with a user; generating motion for the avatar using a generation AI based on the prompt; and updating a motion bank in which the avatar's motion is registered with the generated motion. (2) The information processing method according to (1), further comprising generating an action plan for the avatar, including motion to be performed by the avatar, using the large-scale language model based on context. (3) The information processing method according to (2), wherein the motion included in the action plan is selected from motions registered in the motion bank. (4) The information processing method according to (3), wherein the motion performed by the avatar among the motions registered in the motion bank is updated with the generated motion. (5) The information processing method according to (4), wherein the prompt is generated when the motion is performed a predetermined number of times by the avatar. (6) The information processing method according to any one of (2) to (5), wherein the prompt is a prompt for regenerating the motion included in the action plan. (7) The information processing method according to any one of (2) to (6), wherein the prompt is a prompt for regenerating the motion that the avatar reflexively performed in response to a predetermined action from the user. (8) The information processing method according to any one of (2) to (7), wherein the generation of the action plan by the large-scale language model is performed in response to an action from the user or when the stock of the avatar's action plans is below a threshold. (9) The information processing method according to any one of (2) to (8), wherein the context includes at least information about the environment in which the avatar exists or information about the status of the avatar. (10) The information processing method according to (9), wherein the context further includes information about an action from the user to the avatar.(11) The information processing method according to any one of (1) to (10), wherein the prompt includes information relating to the individuality of the avatar. (12) The information processing method according to any one of (1) to (11), wherein the avatar is displayed on a head-mounted display device worn by the user. (13) An information processing device comprising: an avatar control unit that controls the motion of an avatar interacting with a user based on an action plan generated by a large-scale language model; and a bank update unit that updates a motion bank in which the motion of the avatar is registered, based on the motion regenerated by a generating AI based on a prompt further generated by the large-scale language model.
[0094] 1. Information Processing System 10. Information Processing Device 20. Large-Scale Language Model 30. Generative AI 40. Network 110. Information Acquisition Unit 120. Reflection Judgment Unit 130. Instruction Generation Unit 140. Avatar Control Unit 150. Bank Update Unit 160. Information Management Unit 161. AI State Management Unit 162. Motion Bank 163. Location Information Management Unit AA. Avatar
Claims
1. A computer-based information processing method comprising: generating prompts using a large-scale language model for generating motions for an avatar that interacts with a user; generating motions for the avatar using a generation AI based on the prompts; and updating a motion bank in which the avatar's motions are registered with the generated motions.
2. The information processing method according to claim 1, further comprising generating an action plan for the avatar, including motions to be performed by the avatar, using the large-scale language model based on context.
3. The information processing method according to claim 2, wherein the motion included in the action plan is selected from motions registered in the motion bank.
4. The information processing method according to claim 3, wherein the motions performed by the avatar among the motions registered in the motion bank are updated with the generated motions.
5. The information processing method according to claim 4, wherein the prompt is generated when the motion is performed a predetermined number of times by the avatar.
6. The information processing method according to claim 2, wherein the prompt is a prompt for regenerating the motion included in the action plan.
7. The information processing method according to claim 2, wherein the prompt is a prompt for regenerating a motion that the avatar reflexively performed in response to a predetermined action from the user.
8. The information processing method according to claim 2, wherein the generation of the action plan by the large-scale language model is performed in response to an action from the user, or when the stock of the avatar's action plans is below a threshold.
9. The information processing method according to claim 2, wherein the context includes at least information about the environment in which the avatar exists, or information about the status of the avatar.
10. The information processing method according to claim 9, wherein the context further includes information relating to an action taken by the user on the avatar.
11. The information processing method according to claim 1, wherein the prompt includes information relating to the personality of the avatar.
12. The information processing method according to claim 1, wherein the avatar is displayed on a head-mounted display device worn by the user.
13. An information processing device comprising: an avatar control unit that controls the motion of an avatar interacting with a user based on an action plan generated by a large-scale language model; and a bank update unit that updates a motion bank in which the avatar's motion is registered, using the motion regenerated by a generating AI based on prompts further generated by the large-scale language model.
Citation Information
Patent Citations
Method, system, server device, terminal device, and program for distributing data constituting three-dimensional figure
JP2013175066A
Translation device and program
JP2021196708A
Subtitle presentation control device and program
JP2023025538A
Nonverbal information generation device, method, and program
WO2019160090A1
Information processing device, information processing method, and program
WO2020170441A1