Information processing system and information processing method
The information processing system facilitates natural non-verbal interactions in virtual reality by generating and transmitting context information for expected avatar responses, addressing the challenge of devices lacking motion detection capabilities.
Patent Information
- Application Number
- PCT/JP2025/007064
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-19
- Filing Date
- 2025-02-28
- Publication Date
- 2025-09-25
AI Technical Summary
Existing virtual reality systems struggle to accurately represent non-verbal interactions between avatars, particularly when users are using devices that cannot detect arm or face movements, leading to a lack of natural communication and excitement in virtual events.
An information processing system that includes a first acquisition unit to identify interaction partners and generate context information for expected responses, transmitting this information to devices that cannot detect user movements, allowing them to replicate intended avatar motions.
Enables natural non-verbal communication between avatars, enhancing the excitement and engagement in virtual events by ensuring devices without motion detection capabilities can react to intended movements in real-time.
Smart Images

Figure JP2025007064_25092025_PF_FP_ABST
Abstract
Description
Information processing system and information processing method
[0001] The present disclosure relates to an information processing system and an information processing method for representing an avatar in virtual reality.
[0002] Various technologies have been put to practical use that present users with three-dimensional virtual spaces constructed within computers or computer networks, such as those called metaverses. For example, by applying so-called XR technologies, such as virtual reality (VR), augmented reality (AR), and mixed reality (MR), to the representation of virtual spaces, it is possible to reflect the behavior of real users in avatars.
[0003] As an example, a technique is known for analyzing the voices and movements of users and reflecting the analyzed predicted movements in an avatar as a means of communication between users in a virtual space (see, for example, Patent Document 1).
[0004] JP 2016-48855 A
[0005] However, there is room for further improvement in the way avatars are represented in virtual spaces.
[0006] For example, in a virtual space with many participants, many users participate in events created by creators, and various forms of communication are exchanged between users. Creators want to express in the virtual space how excited avatars are at an event. Specifically, creators want to express how avatars actively communicate with each other through nonverbal means, such as cheering and high-fiving, in scenes of an exciting sports game. Participating users also expect their avatars to actively communicate with each other in the virtual space, just as they do in the real world, and to express their avatars' movements in a variety of ways.
[0007] However, users participating in virtual spaces use a variety of devices, and it is not always possible for a user to reflect the desired actions in the avatar. For example, a user using a device that cannot acquire the user's arm movements or face direction may not be able to immediately respond to the other user's avatar's high-five action.
[0008] Therefore, the present disclosure proposes an information processing system and an information processing method that can express avatars in virtual space in a variety of ways as intended by the user.
[0009] In order to solve the above problems, one form of information processing system according to the present disclosure includes a first acquisition unit that acquires information identifying a second avatar with which a first avatar interacts in a virtual space and input information that triggers the motion of the first avatar, a first generation unit that generates motion of the first avatar according to the input information and generates context information including an expected response that is an expected response from the second avatar in response to the motion of the first avatar, and a first transmission unit that transmits the context information to the second avatar as control information that controls the motion of the second avatar.
[0010] FIG. 1 is a diagram illustrating an overview of an information processing system according to an embodiment. FIG. 2 is a diagram illustrating an overview of the flow of information processing according to an embodiment. FIG. 3 is a sequence diagram illustrating the flow of information processing according to an embodiment. FIG. 4 is a diagram illustrating an example format of context information. FIG. 5 is a diagram illustrating a specific example of context information. FIG. 6 is a sequence diagram illustrating the flow of information processing including partial correction processing. FIG. 7 is a sequence diagram illustrating the flow of information processing including pre-motion. FIG. 8 is a sequence diagram illustrating the flow of information processing including a response block. FIG. 9 is a sequence diagram illustrating the flow of information processing including a plurality of response candidates. FIG. 10 is a diagram illustrating an example configuration of an information processing system according to an embodiment. FIG. 11 is a flowchart illustrating the procedure of response processing according to an embodiment. FIG. 12 is a hardware configuration diagram illustrating an example of a computer that realizes the functions of an HMD.
[0011] Hereinafter, embodiments will be described in detail with reference to the drawings. In the following embodiments, the same components are designated by the same reference numerals, and redundant description will be omitted.
[0012] The present disclosure will be described in the following order of items: 1. Embodiment 1-1. Overview of information processing system according to embodiment 1-2. Overview of information processing according to embodiment 1-3. Specific example of information processing according to embodiment 1-4. Example of context information 1-5. Variation of information processing 1-6. Configuration of information processing system according to embodiment 1-7. Procedure of response processing according to embodiment 2. Other embodiments 3. Effects of information processing system according to the present disclosure 4. Hardware configuration
[0013] (1. Embodiment) (1-1. Overview of Information Processing System According to Embodiment) An overview of an information processing system according to an embodiment will be described using Fig. 1. Fig. 1 is a diagram showing an overview of an information processing system 1 according to an embodiment.
[0014] Information processing according to the embodiment is executed by an information processing system 1 illustrated in Fig. 1. The information processing system 1 includes a head mounted display (HMD) 100, a smartphone 200, and a cloud server 300. The HMD 100, the smartphone 200, and the cloud server 300 are connected to each other via a network so as to be able to communicate with each other.
[0015] The cloud server 300 is an information processing server that provides users with a three-dimensional virtual space (hereinafter simply referred to as "virtual space") constructed within a computer or computer network, such as a metaverse.
[0016] The HMD 100 is an information processing device used by the user 10. In a virtual space 30 provided by a cloud server 300, an avatar, which is a virtual character that reflects the movements of the user 10, is displayed. For ease of distinction, the avatar corresponding to the HMD 100 will be referred to as an "HMD avatar" and the avatar corresponding to the smartphone 200 will be referred to as a "mobile avatar" below. When there is no need to distinguish between the two, they will simply be referred to as "avatars."
[0017] The HMD 100 reflects the movements of the user 10 in the HMD avatar 10A based on input information from, for example, an acceleration sensor included in the device itself, a controller operated by the user 10, or various sensors worn by the user 10, and displays the reflected movements in the virtual space. In other words, in the present disclosure, the HMD 100 is an example of a device that can estimate the behavior in the virtual space performed by the HMD avatar 10A by acquiring the behavior of the user 10, such as head and arm movements. Note that, although the embodiment exemplifies the HMD 100 as a device capable of acquiring the user's behavior, the device capable of acquiring the user's behavior is not limited to the HMD 100 and may be a VR controller or the like. Furthermore, the device capable of acquiring the user's behavior may be a device connected to a camera that captures an image of the user in real space and reflects the user's gestures and facial expressions in an avatar in the virtual space.
[0018] The smartphone 200 is an information processing device used by the user 20. For example, the smartphone 200 determines the behavior of the mobile avatar 20A operated by the user 20 based on information input to an input unit such as a touch display, and displays the mobile avatar 20A in the virtual space 30. The smartphone 200 may have a built-in acceleration sensor, but in the present disclosure, information from such sensors is not reflected in the operation of the mobile avatar 20A. That is, in the present disclosure, the smartphone 200 is an example of a device that cannot acquire (or requires time to acquire) behavior such as head or arm movements of the user 20 and cannot instantly estimate the motions performed by the mobile avatar 20A. Note that, although the embodiment exemplifies the smartphone 200 as a device that cannot acquire user behavior, the device that cannot acquire user behavior may be various devices such as a PC (Personal Computer) or a tablet terminal.
[0019] (1-2. Overview of Information Processing According to the Embodiment) When the information processing system 1 provides a virtual space 30 to users 10, 20, etc., multiple avatars (in other words, multiple users) exist simultaneously in the virtual space 30. As a result, users can communicate via their own avatars in the same way as in real space. For example, users can greet other users and converse (voice chat or text chat) with them via their avatars, thereby deepening their interactions with other users.
[0020] Virtual spaces can also serve as a place for users to interact with each other or for sales promotion, with operators of the virtual spaces and users themselves hosting events. For example, event creators may hold virtual live music concerts or sports matches in virtual spaces to attract many users. In such cases, creators hope to express the excitement of the event, such as avatars cheering and high-fiving each other as if they were in real life. This creates a fun atmosphere and may encourage more users to participate in the event. In other words, creators expect not only verbal communication such as voice chat but also non-verbal communication through avatar movements.
[0021] However, many users participating in virtual spaces use devices such as smartphones 200, which make it difficult to detect user movements in real time. As a result, even when avatars interact with each other at an event, natural, human-like motions may not be generated as intended by the user. In other words, nonverbal communication is not actively carried out throughout the event, and as a result, the event lacks excitement.
[0022] Therefore, in a virtual space, it is desirable that the behavior of the participating users be reflected in the avatars as intended, regardless of the devices they use, so that the avatars can communicate with each other naturally without using verbal language.
[0023] Therefore, the information processing system 1 according to the present disclosure realizes the above process by the following process. Specifically, the HMD 100 of the information processing system 1 identifies a second avatar (e.g., a mobile avatar 20A) with which a first avatar (e.g., an HMD avatar 10A) is to interact in a virtual space, and then acquires input information that triggers a movement of the first avatar (e.g., a movement of the user 10 raising his / her hand to perform a high-five). Furthermore, the HMD 100 generates context information including expected response information, which is information regarding a response expected from the second avatar in response to the movement of the first avatar (e.g., a high-five movement by the second avatar). The HMD 100 then transmits the generated context information to the smartphone 200, which is the second avatar side terminal. Upon receiving the context information, the smartphone 200 reflects a movement based on the expected response information included in the context information in the second avatar.
[0024] That is, the information processing system 1 transmits information for automatically generating a reaction motion of the mobile avatar 20A from the HMD 100, which reflects the movements of the user 10, to the smartphone 200, which is the user's interaction partner in the virtual space but cannot reflect the movements of the user 20. This allows the information processing system 1 to realize nonverbal communication, in which avatars high-five each other in a natural manner, even if one device, such as the smartphone 200, cannot capture the user's movements themselves. This allows the information processing system 1 to express the visual excitement of watching a sports game or a live music concert in a virtual space. Furthermore, by users sharing their experiences on social networking services (SNS), it is expected that the number of visitors to the virtual space will increase and the advertising effectiveness of the event will improve.
[0025] (1-3. Specific Example of Information Processing According to the Embodiment) The information processing according to the embodiment will be described in detail with reference to Figure 2 and subsequent figures. Figure 2 is a diagram showing an overview of the flow of information processing according to the embodiment.
[0026] 2 shows communication between an HMD avatar 10A corresponding to the HMD 100 and a mobile avatar 20A corresponding to the smartphone 200. In Fig. 2, screen displays 32A, 32B, and 32C show the HMD avatar 10A and the mobile avatar 20A displayed on the HMD 100. In addition, screen displays 34A, 34B, and 34C show the mobile avatar 20A and the HMD avatar 10A displayed on the smartphone 200.
[0027] As shown on the screen display 32A, the HMD 100 identifies the mobile avatar 20A as the communication partner of the HMD avatar 10A. For example, the HMD 100 identifies the communication partner based on the body orientation and line of sight of the HMD avatar 10A, whether a text or voice chat has been established with the HMD avatar 10A, whether the mobile avatar 20A is located in a specific position (for example, a specific seat at an event), etc.
[0028] Furthermore, as shown in screen display 34A, when mobile avatar 20A is identified by HMD avatar 10A as a communication partner, smartphone 200 outputs an image that lets user 20 know that they have been identified (such as displaying a reaction indicating that the avatars have noticed each other).
[0029] Next, as shown in screen display 32B, HMD 100 executes a motion of HMD avatar 10A that reflects the behavior of user 10. Specifically, HMD 100 acquires input information that user 10 has raised his / her hand, estimates the overall motion, which is a body expression of HMD avatar 10A, as a "high five" based on the input information, and outputs an image that reflects the high five motion.
[0030] At this time, the HMD 100 generates context information that includes information regarding a motion expected from the mobile avatar 20A in response to the motion of the HMD avatar 10A (hereinafter referred to as an "expected response") and may also include various information regarding the HMD avatar 10A. The context information includes ID information for identifying the avatar (in this example, the mobile avatar 20A) specified as the partner of the HMD avatar 10A's interaction, the motion performed by the HMD avatar 10A, and various information associated with the motion. The context information is transmitted via the cloud server 300 to the smartphone 200, which is a device corresponding to the mobile avatar 20A. In the smartphone 200 that receives the context information, the context information functions as control information for controlling, for example, the motion of the mobile avatar 20A.
[0031] The screen display 34B shows the HMD avatar 10A, with which the mobile avatar 20A is interacting, making a motion of high-fiving. When the smartphone 200 receives the context information, the mobile avatar 20A determines whether or not an automatic response is to be performed, that is, whether or not to automatically execute the expected response included in the context information (step S10).
[0032] For example, the smartphone 200 determines to make an automatic response in accordance with a rule that, when an expected response is received, a motion is automatically executed based on the expected response, based on manual settings by the user 20. Alternatively, when the smartphone 200 is set so that the user 200 determines each time whether or not to respond to an expected response, the smartphone 200 does not automatically execute a motion related to the expected response, but asks the user 20 whether or not to execute a motion.
[0033] If the automatic response determination determines that the mobile avatar 20A will automatically make the expected response, or if the smartphone 200 receives an operation to make the expected response from the user 20, the smartphone 200 causes the mobile avatar 20A to make the expected response. Specifically, the mobile avatar 20A makes a "high-five" motion, which is the expected response transmitted from the HMD avatar 10A.
[0034] As shown in the screen display 34C, when the mobile avatar 20A performs a motion, the HMD avatar 10A and the mobile avatar 20A perform a high-five with each other. This allows the user 20 to instantly react to the motion of the HMD avatar 10A and have the mobile avatar 20A perform a high-five without having to perform a motion corresponding to the high-five or perform a key operation to respond to the high-five.
[0035] Furthermore, as shown in screen display 32C, when the mobile avatar 20A performs a high-five, the HMD avatar 10A and the mobile avatar 20A high-fiving each other are also displayed on the screen of the HMD 100. This allows the user 10 to confirm that the mobile avatar 20A immediately responded to the motion of the user's HMD avatar 10A, allowing for natural, real-life communication with the user 20 in the virtual space.
[0036] The information processing shown in Fig. 2 will be described in more detail with reference to Fig. 3. Fig. 3 is a sequence diagram showing the flow of information processing according to the embodiment.
[0037] 3 shows an example in which the HMD 100 and the smartphone 200 transmit and receive information via a cloud server 300 (not shown). Furthermore, the HMD avatar 10A of the HMD 100 as viewed from the HMD 100 side is referred to as "local" on the HMD 100 side, and the mobile avatar 20A, which is the other party, is referred to as "remote." Conversely, the mobile avatar 20A of the HMD 100 as viewed from the smartphone 200 side is referred to as "local" on the smartphone 200 side, and the HMD avatar 10A, which is the other party, is referred to as "remote."
[0038] First, the HMD 100 identifies a communication partner based on the line of sight of the HMD avatar 10A, etc. (Step S30) For example, the HMD 100 identifies the mobile avatar 20A as the communication partner.
[0039] Next, the HMD 100 generates a whole motion of the HMD avatar 10A based on the movement of the user 10 (step S31). For example, the HMD 100 generates a "high five" as a whole motion of the HMD avatar 10A based on the movement of the user 10 raising their arms high.
[0040] At this time, the smartphone 200 receives information about the motion performed by the avatar 10A in the virtual space via the cloud server 300, and reproduces, on the screen display, the motion generated by the HMD 100. That is, the smartphone 200 replicates the high-five motion of the HMD avatar 10A in the HMD 100 (step S32).
[0041] Next, the HMD 100 estimates the motion of the mobile avatar 20A as an expected response to the high-five of the HMD avatar 10A (step S33). For example, the HMD 100 expects the mobile avatar 20A to return the same motion in response to the high-five of the HMD avatar 10A, and estimates the motion of the mobile avatar 20A as a "high-five."
[0042] The HMD 100 generates context information including the expected response, and transmits the generated context information to the cloud server 300 (step S34).
[0043] When the smartphone 200 receives the context information, it analyzes the expected response included in the context information, narrows down the response motions, and switches the internal state of the smartphone 200 to "automatic response mode ON" (step S35). "Automatic response mode ON" refers to, for example, an internal processing setting that returns a response automatically or by a simple operation to an expected response sent from the other party. Note that narrowing down of the response motions does not need to be performed if a motion is specified in the context information. For example, narrowing down of the response motions may be performed when there are multiple possible response motions that the mobile avatar 20A can make in response to the expected response in the context information.
[0044] The smartphone 200 starts generating a response motion automatically or when a simple operation by the user 20 (such as touching the screen) is used as a response start trigger (step S36).
[0045] When the smartphone 200 receives the response start trigger, the smartphone 200 generates a response motion corresponding to the expected response in the mobile avatar 20A (step S37). Specifically, the smartphone 200 generates a high-five motion for the mobile avatar 20A.
[0046] The HMD 100 replicates the high-five motion of the mobile avatar 20A on the smartphone 200 (step S38). As a result, a scene of the HMD avatar 10A and the mobile avatar 20A high-fiving each other is displayed on the screens of both the HMD 100 and the smartphone 200. After that, the smartphone 200 switches the automatic response mode, which was turned on upon receiving the context information, to OFF (step S39).
[0047] (1-4. Example of Context Information) Context information will be described with reference to Figures 4A and 4B. Figure 4A is a diagram showing an example of the format of context information.
[0048] The data table 40 shows examples of each piece of information that constitutes the context information, which is, for example, a command (control information) transmitted via a network.
[0049] In the example of Figure 4, "Context" is selected in the "Command" field. Furthermore, the "Address" field contains identification information (ID) and the like for specifying the destination avatar. The information for specifying the avatar may be an individual avatar ID, information specifying a group of avatars, or information specifying avatars located in a specified group of seats (block) at an event venue.
[0050] The "Body" field includes an expected response that the sending avatar expects from the other party. The expected response is, specifically, information specifying a motion of the other party. The method for specifying the expected response may be any method, and may be specified, for example, by an administrator of the virtual space. As an example, the expected response is specified as the same motion as the motion of the destination. In this case, if the destination is a "high five," the expected response is also specified as "high five." Alternatively, several expected responses may be set in advance for a certain motion, and a motion may be specified randomly from among them. For example, if the destination is a "high five," multiple motions such as "high five," "cheer," and "jump for joy" may be set as candidates for expected responses, and one motion randomly selected from among them may be specified as the expected response included in the context information.
[0051] "Options" is information that specifies various information processes that are executed in association with the expected response.
[0052] The "motion option" is an option that specifies whether the expected response is passive or active. For example, if the expected response is "passive," the avatar that received the context information will perform a motion in response to the motion of the destination avatar after it has been performed. Alternatively, if the expected response is "active," the avatar that received the context information will perform a motion before or at the same time as the motion of the destination avatar is completed.
[0053] The "UI (User Interface) Option" is information that sets whether or not to display candidate expected responses on the screen of the receiving device. For example, if the "UI Option" is set to "Hide," the receiving device automatically executes one specified motion without displaying candidate motions. For example, a device that receives a desired response based on a "high five" motion executes a motion to return a "high five" motion to the other party without displaying candidate response motions.
[0054] The "voice option" is information that specifies the sound that the avatar will emit when performing a motion. For example, if "cheers" is set in the "voice option," the avatar will be controlled to cheer in response to a high-five. This allows the information processing system 1 to create a scene in which the avatars are excited with each other at the event venue. The "voice option" may also be control information that silences (mutes) the other party or outputs other sound effects besides cheers.
[0055] The "SE (Sound Effect) Option" is information that specifies sound effects other than voice. For example, if "clap" is set in the "SE Option," the receiving device outputs the sound of clapping along with the avatar's motion. This allows the information processing system 1 to create excitement based on the behavior of the avatars, as well as acoustic excitement.
[0056] The options are not limited to the above examples and may be set arbitrarily by an administrator of the virtual space, etc. Furthermore, the context information may have security options such as a command length or a checksum value added to the command header. While the command transmission of the context information is assumed to be, for example, Internet Protocol (IP) communication using User Data Protocol (UDP) or Transmission Control Protocol (TCP), the transmission method is not limited to these examples and may be based on any rules.
[0057] A specific example of context information is shown in Fig. 4B. Fig. 4B is a diagram showing a specific example of context information.
[0058] 3 , the context information 42 is an example of the context information transmitted by the HMD 100 to the smartphone 200. As shown in FIG. 4B , the context information 42 is information in which the destination is the "mobile avatar 20A," the response expected from the mobile avatar 20A is a "high-five," and options are set to "passive, with cheers, and with applause sound effects." Upon receiving the context information 42, the smartphone 200 controls the motion, etc., of the mobile avatar 20A in accordance with the information included in the context information 42.
[0059] In the examples up to Figure 4B, a "high five" is shown as an example of a motion between avatars, but there are various possible examples of motions between avatars, and therefore various motions can be selected as the expected response accordingly.
[0060] For example, motion 44 is a hand signal that mimics the action of grinding a pepper mill. If the avatar to which the context information is sent makes the motion of grinding a pepper mill, the recipient user expects the recipient to make the same motion, and therefore the context information includes the expected response of "motion of grinding a pepper mill." Another example of avatars making the same motion like this could be a motion of clapping their fists together.
[0061] Furthermore, the motions of the avatars are not limited to being identical, but may respond to the destination user's motion and perform a different motion from that of the other user. For example, motion 46 is an example of a hand signal that forms a rock-paper-scissors relationship with the other user, and the winner is determined by the strength of that relationship (called "rock-paper-scissors" in Japan). In this case, the destination user expects the other user to perform one of the rock-paper-scissors motions, so the context information includes the expected response of "any of the rock-paper-scissors motions."
[0062] Furthermore, motions that expect a different response from the other avatar may include a relationship in which one avatar makes a certain motion and the other avatar receives the motion. For example, there may be a motion between two avatars in which the destination avatar makes a motion of hugging or kissing the other, and the other avatar accepts it.
[0063] Furthermore, the motions between avatars may include a combination of multiple signs to form a single motion. For example, motion 48 is an example in which the hand signs of each avatar form a shape with a single meaning (a "heart" in the example of FIG. 4B ). In this case, the destination user expects the recipient user to perform a motion that is paired with the recipient user, and the context information includes an expected response of "a motion to form the other half of a heart formed by two people."
[0064] Note that the motions shown in FIG. 4B are merely examples, and the motions made up of multiple people may include all kinds of motions, such as various gestures and love signs in the country.
[0065] (1-5. Variations of Information Processing) Next, variations (modified examples) of the information processing shown in FIG. 3 will be described using FIG. 5 and subsequent figures. FIG. 5 is a sequence diagram showing the flow of information processing including partial correction processing. Note that in FIG. 5 and subsequent figures, explanations of overlapping content explained in FIG. 3 and other figures will be omitted.
[0066] 5, the explanation begins with the HMD avatar 10A performing a full-body motion and the HMD 100 transmitting its context information (step S50). The smartphone 200 narrows down the response motions and turns on the automatic response mode (step S51), and generates a response motion via a response start trigger (step S52) (step S53). The HMD 100 then copies the motion of the mobile avatar 20A (step S54).
[0067] In this case, in the example of FIG. 5, the hand positions of the HMD avatar 10A and the mobile avatar 20A are misaligned, or there is a slight misalignment in their standing positions, so that the coordinates of their hands are misaligned from the specified values.
[0068] In this case, the smartphone 200 determines the difference in the positions of the hands and partially corrects the motion (step S55). Specifically, the smartphone 200 partially corrects the motion of the mobile avatar 20A so that the coordinates of the hand held out by the HMD avatar 10A and the coordinates of the hand of the mobile avatar 20A fall within specified values.
[0069] The HMD 100 replicates the partially corrected motion of the mobile avatar 20A on the HMD 100 side (step S56). This allows the HMD 100 and the smartphone 200 to output a high-five motion between the avatars as video without displaying an unnatural screen image, such as the motion being performed with the hands of the avatars out of sync. Thereafter, the smartphone 200 switches the automatic response mode to OFF (step S57).
[0070] In this way, the smartphone 200 that executes a motion related to an expected response may make partial corrections (such as shifting the hand position) to the overall motion (such as a high-five) so that the motion is expressed as expected by the user 10. This allows the information processing system 1 to realize communication between users as intended, without displaying something that looks unnatural, such as a misaligned hand or head position, when attempting to execute a motion.
[0071] Next, a process including pre-motion in addition to the whole motion will be described with reference to Fig. 6. Fig. 6 is a sequence diagram showing the flow of information processing including pre-motion. The process shown in Fig. 6 is performed when attempting to execute a motion that cannot be executed unless the distance between avatars is within a specified value, such as a high-five.
[0072] First, the HMD 100 identifies the other person of the HMD avatar 10A (step S60). Furthermore, when the user 10 raises his / her arm to make the HMD avatar 10A perform a high-five motion, the HMD 100 generates a pre-motion if it determines that the other person is at a distance greater than a certain distance (step S61).
[0073] A pre-motion is a motion that is required to complete a motion that the user 10 wants the HMD avatar 10A to execute when the motion cannot be completed for some reason. An example of a pre-motion is when the avatar moves to a distance where a high-five can be performed because the other person is far away when the user 10 is about to high-five.
[0074] When the pre-motion is generated by the HMD 100, the smartphone 200 copies the pre-motion (step S62). Specifically, a motion such as the HMD avatar 10A calling out to the mobile avatar 20A (such as a motion requesting the mobile avatar 20A to move to the vicinity of the HMD avatar 10A) is displayed on the screen of the smartphone 200.
[0075] The HMD 100 generates a pre-motion, estimates a motion expected from the other person in relation to the pre-motion (step S63), and generates context information including the expected response (step S64). Specifically, the context information includes an expected response that the other person's mobile avatar 20A moves closer to the HMD avatar 10A (i.e., the two avatars move closer to each other).
[0076] When the smartphone 200 receives the context information, it turns on the automatic response mode (step S65), triggers a response start (step S66), and generates a motion corresponding to the expected response (in this example, a movement toward the HMD avatar 10A) (step S67).
[0077] The HMD 100 replicates the movement motion of the mobile avatar 20A (step S68). This allows the avatars to approach each other close enough to perform a high-five without requiring any special user operation. The information processing system 1 then executes a process (similar to that shown in FIG. 3) to realize the "high-five" motion detected in step S61.
[0078] In this way, the information processing system 1 can automatically execute a pre-motion before a whole motion, so the user can reflect the intended motion in the avatar without having to take the trouble of explicitly moving closer to the other person. Note that the pre-motion is not limited to attracting the other person, but may also be various adjustment processes for realizing the subsequent motion, such as moving from the user toward the other person or rotating the body to face the other person. Furthermore, the pre-motion may be executed by the HMD avatar 10A or the mobile avatar 20A.
[0079] Next, the block processing will be described with reference to Fig. 7. Fig. 7 is a sequence diagram showing the flow of information processing including a response block. Fig. 7 shows an example in which the smartphone 200 receives context information including an expected response and does not execute a motion related to the expected response.
[0080] The HMD 100 identifies the person with whom the HMD avatar 10A is interacting (step S70) and generates a whole-body motion (step S71). The smartphone 200 replicates the motion of the HMD avatar 10A (step S72). The HMD 100 estimates the motion of the person corresponding to the whole-body motion (step S73), generates context information including the estimated expected response, and transmits it to the smartphone 200 (step S74).
[0081] The smartphone 200 analyzes the received context information (step S75). At this time, if the smartphone 200 determines that it will not execute a motion to respond to the other party for some reason, it keeps the automatic response mode OFF and blocks responses to the other party (step S76).
[0082] There may be various conditions under which the smartphone 200 blocks responses to other parties. For example, each avatar may be defined with a boundary that denies interaction with other parties in the virtual space. Hereinafter, the line defining such a boundary will be referred to as a personal boundary line. The personal boundary line may be formed, for example, as a circle of a specified size centered on the coordinates where the avatar is located.
[0083] For example, an avatar may block an avatar attempting to communicate with another avatar from outside the personal boundary. In this case, the avatar may not block the avatar from interacting with another avatar if there is prior communication between the avatar and the avatar, such as eye contact or chat. In other words, the avatar may block the avatar from interacting with another avatar from outside the personal boundary only if the avatar has its back turned to the avatar and no prior communication has been established.
[0084] Alternatively, if the avatar is an avatar attempting to interact from outside the personal boundary and has not established prior communication with the avatar, but is gazing at the same object, such as watching a sporting event together, the avatar may not be blocked from interacting with the avatar. Alternatively, if the avatar is an avatar attempting to interact from outside the personal boundary and is concentrating on some object (such as being in the same place for a certain period of time or looking at the same object for a certain period of time), the avatar may be blocked from interacting with the avatar.
[0085] An avatar may also block others for other reasons, regardless of personal boundaries. For example, an avatar may automatically block interactions with others who are on a block list created by the avatar. An avatar may also automatically block interactions with others if a setting to block interactions from others is configured in the virtual space, such as a dialogue blocking function. An avatar may also block interactions if the virtual space is equipped with a function to notify the user when an interaction request is made, and as long as the user does not take an action to allow the interaction.
[0086] Furthermore, the avatar may determine the block based on profile information, which is information such as attributes that are individually set for the avatar or the user.
[0087] The profile information includes, for example, a name for user search, the appearance of the avatar (appearance such as a gamer or a suit), a profile photo or image, follower information, and functions for connecting with other users (follow, follow back, unfollow, linking with SNS), etc. Note that profile information can be set to public or private, and some users may have multiple profiles set up (for family, friends, community, etc.).
[0088] Furthermore, the profile information may also include user status information (the status of the user controlling the avatar, such as whether they are participating in an event, free, or busy), registration information for friends and groups, a list of friend locations, a trust rank indicating the trustworthiness of the user, etc. The profile information may also include a free text field for introducing the user or avatar and writing any content, link information they would like to introduce to others, their own supported languages, etc.
[0089] Furthermore, the profile information may include the degree of intimacy between avatars or users, the user's level of liking for other users, and attributes (gender, age, etc.) set for the avatar.
[0090] The information processing system 1 or a user of the information processing system 1 can create a block determination condition using the above-mentioned various profile information.
[0091] For example, if the intimacy between avatars or users is low, the avatar may block expected responses from the other party. The avatar may also block expected responses from avatars other than friends. Alternatively, if the genders set for the avatars are opposite genders, the avatar may block expected responses that involve touching the other party. Blocking may be determined based on a combination of profile information and motion characteristics.
[0092] Next, an example in which there are multiple possible response candidates for an expected response will be shown with reference to Fig. 8. Fig. 8 is a sequence diagram showing the flow of information processing including multiple possible responses.
[0093] 8, the explanation begins with the HMD avatar 10A performing a whole body motion and the HMD 100 transmitting its context information (step S80). The smartphone 200 turns on the automatic response mode in response to the expected response (step S81).
[0094] At this time, if there are multiple candidates for the expected response, the smartphone 200 waits for an input from the user 20, for example, and selects a response from the multiple candidates according to an operation of the user 20 (step S82). For example, if the expected response is a motion requesting mutual joy (such as a "high five," "raising both hands," or "cheer") and the user 20 needs to select one of the candidates, the smartphone 200 waits for an input from the user 20. As another example, if the expected response is "rock-paper-scissors" and the user 20 needs to select one of the candidates in a rock-paper-scissors game, the smartphone 200 waits for an input from the user 20.
[0095] When the user 20 inputs, the smartphone 200 starts a response triggered by the input (step S83) and generates a response motion (step S84). The HMD 100 replicates the motion of the mobile avatar 20A (step S85). Thereafter, the smartphone 200 switches the automatic response mode to OFF (step S86).
[0096] In this way, when there are multiple possible responses, the information processing system 1 may select and adopt one of the responses and reflect the adopted response motion in the avatar. This method also allows smooth communication between the user 10 of the HMD 100 and the user 20 of a device that is difficult to reflect movements of, such as the smartphone 200.
[0097] The processes described with reference to FIGS. 3 to 8 may be modified in various ways other than the processes shown.
[0098] For example, the motion initiation side is not limited to a device capable of user tracking such as the HMD 100, but may also be a side with no user tracking function such as the smartphone 200 or a desktop PC. In this case, the user 20 of the smartphone 200 can specify the motion to be transmitted by performing various operations to specify the motion (such as text or voice input, bone estimation using a camera, etc.).
[0099] Furthermore, although the example has been shown in which the expected response included in the context information is set by the transmitting device estimating the motion of the other party, the expected response may also be set by the receiving device. For example, when the smartphone 200 detects a high-five motion of the HMD avatar 10A, it estimates a "high-five" as a response motion corresponding to the motion. In this case, when the smartphone 200 receives the context information, it may execute a "high-five" estimated by itself as a response motion, regardless of the expected response in the context information. Alternatively, the HMD 100 may receive information regarding the response motion estimated by the smartphone 200 and, based on the received information, set the response motion estimated by the smartphone 200 as an expected response to be included in the context information.
[0100] Furthermore, the context information may be generated not by motion estimation but by UI operations using a UI displayed by the HMD 100, voice input, etc. This allows even a device without a motion estimation function to generate context information and transmit it to the other party.
[0101] Furthermore, context information may be transmitted and received between multiple avatars, not only in one-to-one communication but also in one-to-many communication. For example, in a virtual live venue, a performer avatar may transmit context information including expected responses, such as cheering or raising both hands, to multiple avatars present in the venue. This allows the information processing system 1 to realize a production that creates a sense of unity, such as many spectators performing the same or similar motions in the virtual live venue.
[0102] (1-6. Configuration of Information Processing System According to Embodiment) Next, the configuration of the information processing system 1 will be described. FIG. 9 is a diagram showing an example of the configuration of the information processing system 1 according to the embodiment. The information processing system 1 transmits and receives information to and from the HMD 100, the smartphone 200, the cloud server 300, or an external device (not shown) via a network N. Note that the information processing system 1 may take various forms other than the example of FIG. 9 as long as it includes a device that transmits context information, a device that receives context information, and a server device that constructs and operates a virtual space.
[0103] First, an HMD 100 will be described as an example of a device that transmits context information. As shown in Fig. 9, the HMD 100 includes a communication unit 110, a storage unit 120, and a control unit 130. Although not shown in the figure, the HMD 100 may also include a sensor unit for detecting a user's movement, a display unit for displaying an image, and an input unit (such as a touch panel) for receiving various operations from the user.
[0104] The communication unit 110 is realized by, for example, a network interface card (NIC) or a network interface controller. The communication unit 110 is connected to a network N by wire or wirelessly, and transmits and receives information to and from the cloud server 300, etc., via the network N. The network N is realized by, for example, a wireless communication standard or method such as Bluetooth (registered trademark), the Internet, Wi-Fi (registered trademark), UWB (Ultra Wide Band), or LPWA (Low Power Wide Area).
[0105] The storage unit 120 is realized by, for example, a semiconductor memory element such as a random access memory (RAM) or a flash memory, or a storage device such as a hard disk or an optical disk.
[0106] The storage unit 120 may have, for example, learning data that defines the relationship between a user's actions and motions, a trained model trained based on the learning data, etc. For example, the trained model is a neural network model that receives an input of a user's action of raising their hand and outputs an output of causing an avatar to perform a motion of high-fiving.
[0107] The storage unit 120 is not limited to storing learning data, and may store, for example, a data table in which user actions (gestures) are paired with avatar motions. The storage unit 120 may also store motion data (animation of bone structure and facial expressions, dramatic effects, etc.) for moving a 3D avatar model in accordance with user actions.
[0108] The control unit 130 is realized by, for example, a central processing unit (CPU), a micro processing unit (MPU), a GPU, etc. executing a program (for example, an information processing program according to the present disclosure) stored inside the HMD 100 using a RAM, etc. as a work area. The control unit 130 is a controller, and may be realized by, for example, an integrated circuit such as an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA).
[0109] As shown in FIG. 9, the control unit 130 includes an acquisition unit 131 , an analysis unit 132 , a generation unit 133 , a replication unit 134 , and a transmission unit 135 .
[0110] The acquisition unit 131 acquires various types of information. For example, the acquisition unit 131 acquires information related to the user's movements via various sensors. Specifically, the acquisition unit 131 tracks the body movements of the user 10 using various sensors that cooperate with the HMD 100, thereby acquiring input information to be reflected in the motion of the avatar.
[0111] The acquisition unit 131 also acquires information for identifying the mobile avatar 20A with whom the HMD avatar 10A is to interact in the virtual space. For example, the acquisition unit 131 identifies the mobile avatar 20A with whom the HMD avatar 10A is to interact based on various information such as the line of sight and body direction of the HMD avatar 10A, whether text chat or voice chat has been established, whether the mobile avatar 20A is within a personal boundary line, etc.
[0112] The analysis unit 132 analyzes the information acquired by the acquisition unit 131. For example, the analysis unit 132 analyzes a motion to be reflected in the HMD avatar 10A based on the movements, voice input, UI operations, etc. of the user 10. For example, the acquisition unit 131 acquires input information using a trained model trained based on data that learns the relationship between information tracked on the body movements of the user 10 and the motion to be reflected in the HMD avatar 10A using the tracked information. The analysis unit 132 analyzes a motion to be reflected in the HMD avatar 10A based on the input information acquired by the acquisition unit 131 (for example, body movements such as gestures of the user 10).
[0113] The generation unit 133 generates a motion of the avatar based on the information analyzed by the analysis unit 132. Specifically, the generation unit 133 generates a motion of the HMD avatar 10A in response to input information based on the movement of the user 10.
[0114] The generator 133 also generates context information including an expected response, which is a response expected from the mobile avatar 20A, the other party in communication, as a reaction to the motion of the HMD avatar 10A. The context information is information that the user 10 or the HMD avatar 10A uses to request a motion from another avatar, i.e., information that triggers a motion from another avatar.
[0115] The duplication unit 134 duplicates the motions, sounds, etc. of avatars other than the HMD avatar 10A corresponding to the user 10 of the device itself. The motions duplicated by the duplication unit 134 are output as images to the display unit of the HMD 100, etc.
[0116] The transmission unit 135 transmits the motion and context information of the HMD 100 generated by the generation unit 133. For example, the transmission unit 135 transmits the motion and context information to the cloud server 300, and transmits the motion and context information to the smartphone 200 as the destination via the cloud server 300.
[0117] Next, a description will be given of the configuration of the smartphone 200. Note that the communication unit 210 corresponds to the communication unit 110, and the storage unit 220 has a function corresponding to the storage unit 120, and therefore a description thereof will be omitted.
[0118] As shown in FIG. 9, the control unit 230 includes an acquisition unit 231 , an analysis unit 232 , a correction unit 233 , a determination unit 234 , a generation unit 235 , a replication unit 236 , and a transmission unit 237 .
[0119] The acquisition unit 231 acquires various types of information. For example, the acquisition unit 231 acquires the motion of the HMD avatar 10A generated by the HMD 100 and context information generated by the HMD 100.
[0120] The analysis unit 232 analyzes the context information acquired by the acquisition unit 231. For example, the analysis unit 232 analyzes information such as the type of user from which the context information was sent and the type of expected response included in the context information.
[0121] The correction unit 233 performs a predetermined correction process after the generation unit 235 (described later) generates a motion or before executing the motion. For example, the correction unit 233 causes the mobile avatar 20A to execute a motion of moving closer to the other person before executing an originally intended motion such as a high-five. Alternatively, the correction unit 233 fine-tunes the position of the motion or the position of the avatar so that the hands of the two people overlap when executing a high-five motion. More specifically, when executing a fixed motion (such as a high-five) determined based on context information transmitted from the other person, the correction unit 233 corrects the motion by overwriting the hand positions, facial orientation, and the like to match the actual situation between the avatars.
[0122] The determination unit 234 determines a motion to be executed as an expected response to the received context information. For example, the determination unit 234 determines a motion of the mobile avatar 20A based on a UI operation or a voice input by the user 20.
[0123] The generation unit 235 generates a motion of the mobile avatar 20A through processing by the correction unit 233 and the determination unit 234. Specifically, the generation unit 235 generates a motion to be executed by the mobile avatar 20A in response to the motion of the HMD avatar 10A as a response expected by the user 10 of the HMD 100. Note that although the correction unit 233, the determination unit 234, and the generation unit 235 are described separately in the embodiment, these processing units may be integrated.
[0124] Note that, when the generation unit 235 determines that the context information or the user 10 who transmitted the context information meets the specified blocking condition, the generation unit 235 may not generate a motion of the mobile avatar 20A corresponding to the expected response. In other words, the generation unit 235 may block the expectation from the other party without necessarily generating a response motion.
[0125] The copying unit 236 copies the motions, sounds, etc. of avatars other than the mobile avatar 20A corresponding to the user 20 of the own device. The motions copied by the copying unit 236 are output as video to the display unit, etc. of the smartphone 200.
[0126] The transmission unit 237 transmits various information such as the motion of the mobile avatar 20A. For example, the transmission unit 237 transmits information such as the motion to the cloud server 300, and transmits the motion and context information to the smartphone 200, which is the other party, via the cloud server 300.
[0127] (1-7. Procedure of response processing according to embodiment) Next, an example of the procedure of the response processing according to the embodiment will be described with reference to Fig. 10. Fig. 10 is a flowchart showing the procedure of the response processing according to the embodiment.
[0128] As shown in FIG. 10, the smartphone 200 logs in to the virtual space provided by the cloud server 300 in accordance with the operation of the user 20 (step S101).
[0129] The smartphone 200 determines whether or not context information has been received from a communication partner in the virtual space (step S102). If the context information has been received (step S102; Yes), the smartphone 200 analyzes what information is included in the context information (step S103).
[0130] Next, the smartphone 200 determines whether to execute a motion process in accordance with the expected response included in the context information (step S104). If it is determined that the motion process is to be executed (step S104; Yes), the smartphone 200 further determines whether a correction operation, such as adjusting the position of the avatar's hand, is required (step S105).
[0131] If a correction operation is necessary (step S105; Yes), the smartphone 200 executes a correction process to shift the hand position and perform an appropriate motion (step S106). On the other hand, if a correction operation is not necessary (step S105; No), the smartphone 200 skips the correction process and executes a motion process (step S107).
[0132] Note that if the smartphone 200 does not execute the motion processing according to the expected response by, for example, blocking the response (step S104; No), or if it has not received the context information in the first place (step S102; No), the smartphone 200 does not execute the motion processing according to the expected response. In this case, the smartphone 200 determines whether it has received an operation command based on another user's manual operation or the like (step S108).
[0133] If there is an operation command (step S108; Yes), the smartphone 200 executes a predetermined process in accordance with the operation command (step S109). If there is no operation command (step S108; No), or if the execution of the motion process has been completed, the smartphone 200 determines that the process for the avatar has been completed.
[0134] The smartphone 200 then determines whether the user has performed an operation to log out of the virtual space (step S110), and if the user has not logged out, the smartphone 200 remains in the virtual space (step S110; No), and if the user has logged out (step S110; Yes), the smartphone 200 completes the information processing.
[0135] (2. Other Embodiments) The processing according to each of the above-described embodiments may be implemented in various different forms other than the above-described embodiments.
[0136] For example, the data stored in the storage unit 120 and the storage unit 220 shown in FIG. 9 may be stored in the cloud server 300. Specifically, the cloud server 300 may store learning data that has learned the correspondence between the user's movements and the avatar's motion. In this case, the cloud server 300 may acquire input information from the HMD 100 that tracks the movements of the user 10, input the acquired input information into a learned model, and output a motion to be applied to the HMD avatar 10A. In other words, the various processes described in the embodiments may not necessarily be executed by the HMD 100 or the smartphone 200, but may include processes executed by the cloud server 300.
[0137] Furthermore, among the processes described in the above embodiments, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically using known methods. Furthermore, the information, including the processing procedures, specific names, various data, and parameters shown in the above documents and drawings, can be changed as desired unless otherwise specified. For example, the various information shown in each drawing is not limited to the information shown in the drawings.
[0138] Furthermore, the components of each device shown in the figure are conceptual functional components and do not necessarily have to be physically configured as shown in the figure. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of them can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc.
[0139] Furthermore, the above-described embodiments and modifications can be combined as appropriate within the scope of not causing any contradiction in the processing content.
[0140] Furthermore, the effects described in this specification are merely examples and are not limiting, and other effects may also be present.
[0141] (3. Effects of the Information Processing System According to the Present Disclosure) As described above, the information processing system according to the present disclosure (information processing system 1 in the embodiment) includes a first acquisition unit (acquisition unit 131 in the embodiment), a first generation unit (generation unit 133 in the embodiment), and a first transmission unit (transmission unit 135 in the embodiment). The first acquisition unit acquires information identifying a second avatar (mobile avatar 20A in the embodiment) with which a first avatar (HMD avatar 10A in the embodiment) interacts in virtual space, and input information that triggers a motion of the first avatar. The first generation unit generates a motion of the first avatar in accordance with the input information, and generates context information including an expected response that is a response expected from the second avatar in response to the motion of the first avatar. The first transmission unit transmits the context information to the second avatar (smartphone 200 in the embodiment) as control information for controlling the motion of the second avatar.
[0142] In this way, the information processing system according to the present disclosure transmits context information including information about the desired motion to a person interacting with the virtual space. This allows the information processing system to accurately convey the intention of the sender to the destination and to express the intended communication through the avatar. This allows the information processing system to generate the desired motion at the destination even if the destination device is unable to perform motion tracking and cannot instantly generate a motion, thereby creating excitement in the virtual space.
[0143] The information processing system includes a first terminal (HMD 100 in this embodiment) having a first acquisition unit, a first generation unit, and a first transmission unit, and a second terminal (HMD 100 in this embodiment) having a second acquisition unit (acquisition unit 231 in this embodiment), a second generation unit (correction unit 233, determination unit 234, and generation unit 235 in this embodiment), and a second transmission unit (transmission unit 237 in this embodiment). In this case, the first acquisition unit acquires input information by tracking the body movement of the user of the first terminal. In other words, the first terminal is a device that can track the user's movement and reflect the tracked movement in motion.
[0144] In this way, the information processing system is composed of one device that can be tracked and another device. In other words, the information processing system allows for a variety of representations of avatars in a virtual space, regardless of the environment of each user.
[0145] The second acquisition unit acquires context information. The second generation unit generates a motion for the second avatar based on the expected response included in the context information, and generates a correction motion required to complete the motion of the second avatar. The second transmission unit transmits the motion for the second avatar and the correction motion to the first avatar.
[0146] In this way, the information processing system may execute a motion that partially corrects the hand position of the executed overall motion, etc. This makes it possible for the information processing system to prevent unnaturalness, such as misalignment of a high-five, or execution of an unrealistic motion, thereby enhancing the user's sense of immersion.
[0147] 6, the first acquisition unit acquires information regarding the presence or absence of a pre-motion, which is a preliminary motion for completing a motion of the first avatar. If the first generation unit determines that a pre-motion is required based on the information regarding the presence or absence of a pre-motion, the first generation unit generates a pre-motion for the first avatar and also generates second context information including a second expected response, which is an expected response in response to the pre-motion. The first transmission unit transmits the second context information to the second avatar prior to transmitting the context information.
[0148] In this way, when a motion is requested that cannot be executed and completed unless the user gets close to the other person, the information processing system executes a pre-motion and induces each avatar to move, etc. This allows the information processing system to have the avatar execute the motion naturally without requiring the user to perform any cumbersome operations.
[0149] The second acquisition unit acquires second context information. The second generation unit generates a motion for moving or rotating the second avatar as a motion corresponding to the pre-motion, based on the second expected response included in the second context information.
[0150] In this way, the information processing system can achieve natural motion in the subsequent main motion processing by moving or rotating the avatar to align the body orientation as a motion corresponding to the pre-motion.
[0151] The second acquisition unit acquires context information. When the second generation unit determines that the context information or the user of the first terminal that transmitted the context information satisfies a blocking condition defined in the second terminal, the second generation unit may not generate a motion of the second avatar corresponding to the expected response included in the context information.
[0152] In this way, the information processing system can block the other party's expectations without necessarily generating a motion that matches the expected response, thereby preventing the information processing system from arbitrarily executing a response motion that the user does not want.
[0153] For example, if the context information is transmitted from outside the personal boundary defined around the second avatar, the second generation unit may not need to generate a motion of the second avatar corresponding to the expected response contained in the context information.
[0154] Alternatively, if the context information is transmitted from outside the personal boundary defined around the second avatar and is transmitted from a user who has not previously interacted with the second avatar, a user who is not identified by the second avatar, or a user whose gaze is different from that of the second avatar, the second generation unit may not need to generate a motion of the second avatar corresponding to the expected response contained in the context information.
[0155] In addition, the second generation unit may not need to generate a motion of the second avatar corresponding to the expected response included in the context information if the context information is sent from a user included in a block list set by the user of the second terminal, or if the user of the second terminal has disabled interaction with other users or does not allow interaction.
[0156] In addition, the second generation unit may determine whether or not to generate a motion of the second avatar corresponding to the expected response included in the context information based on profile information, which is information about the user who sent the context information or the first avatar.
[0157] In this way, the information processing system can set various blocking conditions and automatically block responses based on profile information set in the avatar, etc. This allows the information processing system to realize communication as intended by the user, not only regarding the execution of motions but also regarding the non-execution of motions.
[0158] The first acquisition unit may also acquire input information through a user interface (UI) operation or voice input by the user of the first terminal.
[0159] In this way, the information processing system can acquire input information through various inputs, not limited to body tracking. That is, the sender of the context information does not necessarily have to be a tracking-capable device such as the HMD 100, but may be a device that is not suitable for tracking, such as the smartphone 200.
[0160] The second acquisition unit acquires context information and a motion of the first avatar. The second generation unit estimates a motion of the second avatar as a response expected from the second avatar in response to the motion of the first avatar, and generates the estimated motion of the second avatar as a response to the expected response included in the context information.
[0161] In this way, the information processing system may adopt motions estimated not only by the sender of context information but also by the receiver as expected responses. This allows the information processing system to realize active communication initiated by both parties in a virtual space, rather than necessarily passive communication.
[0162] The first generator generates context information based on a user interface (UI) operation or voice input by a user corresponding to the first avatar.
[0163] In this way, the information processing system may generate context information based on any input from the user, not necessarily based on information estimated from motion. For example, the user can optionally include various information in the context information to produce a variety of expressions in communication.
[0164] The second acquisition unit acquires context information, and the second generation unit estimates a plurality of candidates for the second avatar's motion based on the expected response included in the context information, and generates a motion from one of the estimated candidates.
[0165] In this case, the second generation unit may generate one of the estimated motion candidates in accordance with a selection made by the user of the second terminal.
[0166] Alternatively, the second generation unit may generate, based on the expected response included in the context information, a motion for the second avatar that is different from the motion of the first avatar but can be combined with the motion of the first avatar.
[0167] In this way, the information processing system can provide users with entertaining communication, such as a game of rock-paper-scissors that requires multiple choices as the other person's response motion, or a motion that combines the two people, such as a love sign formed with each other's hands.
[0168] In addition, the first acquisition unit acquires input information based on data learned about the relationship between information tracking the body movements of the user of the first terminal and motions that reflect the tracked information in the first avatar.
[0169] In this way, the information processing system can learn the motions to be reflected in the avatar based on the user's gestures, etc., and can cause the avatar to properly execute the motions intended by the user.
[0170] The first acquisition unit acquires information identifying the second avatars. The first generation unit generates context information including expected responses that are responses expected from the second avatars. The first transmission unit transmits the context information to the second avatars as control information for controlling the motions of the second avatars.
[0171] In this way, the information processing system may use context information not only in one-to-one communication but also in communication between multiple avatars, thereby enabling the information processing system to create a sense of unity, such as when many spectators perform the same or similar motions at a virtual live concert venue.
[0172] (4. Hardware Configuration) Information devices such as the HMD 100, smartphone 200, and cloud server 300 according to the above-described embodiments are realized by a computer 1000 configured as shown in FIG. 11 , for example. The following description will be given using the HMD 100 as an example. FIG. 11 is a hardware configuration diagram showing an example of a computer 1000 that realizes the functions of the HMD 100. The computer 1000 has a CPU 1100, a RAM 1200, a ROM (Read Only Memory) 1300, a HDD (Hard Disk Drive) 1400, a communication interface 1500, and an input / output interface 1600. The components of the computer 1000 are connected by a bus 1050.
[0173] The CPU 1100 operates and controls each component based on programs stored in the ROM 1300 or the HDD 1400. For example, the CPU 1100 loads the programs stored in the ROM 1300 or the HDD 1400 into the RAM 1200 and executes processing corresponding to the various programs.
[0174] The ROM 1300 stores boot programs such as a Basic Input Output System (BIOS) that is executed by the CPU 1100 when the computer 1000 is started, and programs that depend on the hardware of the computer 1000 .
[0175] HDD 1400 is a computer-readable recording medium that non-temporarily records programs executed by CPU 1100 and data used by such programs. Specifically, HDD 1400 is a recording medium that records an information processing program according to the present disclosure, which is an example of program data 1450.
[0176] The communication interface 1500 is an interface for connecting the computer 1000 to an external network 1550 (e.g., the Internet). For example, the CPU 1100 receives data from other devices and transmits data generated by the CPU 1100 to other devices via the communication interface 1500.
[0177] The input / output interface 1600 is an interface for connecting the input / output device 1650 and the computer 1000. For example, the CPU 1100 receives data from an input device such as a keyboard or a mouse via the input / output interface 1600. The CPU 1100 also transmits data to an output device such as a display, a speaker, or a printer via the input / output interface 1600. The input / output interface 1600 may also function as a media interface for reading programs and the like recorded on a predetermined recording medium. Examples of media include optical recording media such as a DVD (Digital Versatile Disc) or a PD (Phase Change Rewritable Disc), magneto-optical recording media such as an MO (Magneto-Optical Disk), tape media, magnetic recording media, and semiconductor memories.
[0178] For example, when the computer 1000 functions as the HMD 100 according to the embodiment, the CPU 1100 of the computer 1000 executes an information processing program loaded onto the RAM 1200 to realize functions of the control unit 130, etc. The information processing program according to the present disclosure and data in the storage unit 120 are stored in the HDD 1400. The CPU 1100 reads and executes program data 1450 from the HDD 1400, but as another example, the CPU 1100 may obtain these programs from another device via an external network 1550.
[0179] The present technology may also be configured as follows: (1) An information processing system comprising: a first acquisition unit that acquires information identifying a second avatar with which a first avatar interacts in a virtual space and input information that triggers a motion of the first avatar; a first generation unit that generates a motion of the first avatar in accordance with the input information and generates context information including an expected response that is a response expected from the second avatar in response to the motion of the first avatar; and a first transmission unit that transmits the context information to the second avatar as control information for controlling the motion of the second avatar. (2) The information processing system according to (1), includes: a first terminal having the first acquisition unit, the first generation unit, and the first transmission unit; and a second terminal having a second acquisition unit, the second generation unit, and the second transmission unit, wherein the first acquisition unit acquires the input information by tracking a body motion of a user of the first terminal. (3) The information processing system according to (2), wherein the second acquisition unit acquires the context information, the second generation unit generates a motion of the second avatar based on an expected response included in the context information and generates a correction motion required to complete the motion of the second avatar, and the second transmission unit transmits the motion of the second avatar and the correction motion to the first avatar. (4) The information processing system according to (2) or (3), wherein the first acquisition unit acquires information regarding the presence or absence of a pre-motion that is a preliminary motion for completing the motion of the first avatar, and when the first generation unit determines that a pre-motion is required based on the information regarding the presence or absence of the pre-motion, generates a pre-motion of the first avatar and generates second context information including a second expected response that is an expected response in response to the pre-motion, and the first transmission unit transmits the second context information to the second avatar prior to transmitting the context information.(5) The information processing system according to (4), wherein the second acquisition unit acquires the second context information, and the second generation unit generates a motion of moving or rotating the second avatar as a motion corresponding to the pre-motion based on the second expected response included in the second context information. (6) The information processing system according to any one of (2) to (5), wherein the second acquisition unit acquires the context information, and when the second generation unit determines that the context information or the user of the first terminal that transmitted the context information satisfies a blocking condition defined in the second terminal, does not generate a motion of the second avatar corresponding to the expected response included in the context information. (7) The information processing system according to (6), wherein the second generation unit does not generate a motion of the second avatar corresponding to the expected response included in the context information when the context information is transmitted from outside a personal boundary defined around the second avatar. (8) The information processing system according to (6) or (7), wherein the second generation unit does not generate a motion of the second avatar corresponding to the expected response included in the context information when the context information is transmitted from outside a personal boundary defined around the second avatar and is transmitted from a user who has not previously interacted with the second avatar, a user who is not identified by the second avatar, or a user who has a different gaze target from the second avatar. (9) The information processing system according to any one of (6) to (8), wherein the second generation unit does not generate a motion of the second avatar corresponding to the expected response included in the context information when the context information is transmitted from a user included in a block list set by a user of the second terminal, or when the user of the second terminal has disabled interaction with other users or does not allow interaction.(10) The information processing system according to any one of (2) to (9), wherein the second acquisition unit acquires the context information, and the second generation unit determines whether to generate a motion of the second avatar corresponding to an expected response included in the context information, based on profile information that is information about a user who transmitted the context information or the first avatar. (11) The information processing system according to any one of (1) and (3) to (10), wherein the information processing system includes a first terminal having the first acquisition unit, the first generation unit, and the first transmission unit, and a second terminal having a second acquisition unit, the second generation unit, and a second transmission unit, and the first acquisition unit acquires the input information by a UI (User Interface) operation or voice input by a user of the first terminal. (12) The information processing system according to any one of (2) to (11), wherein the second acquisition unit acquires the context information and a motion of the first avatar, and the second generation unit estimates a motion of the second avatar as a response expected of the second avatar to the motion of the first avatar, and generates the estimated motion of the second avatar as a response to the expected response included in the context information. (13) The information processing system according to any one of (1) to (12), wherein the first generation unit generates the context information based on a UI (User Interface) operation or voice input by a user corresponding to the first avatar. (14) The information processing system according to any one of (2) to (13), wherein the second acquisition unit acquires the context information, and the second generation unit estimates a plurality of candidates for the motion of the second avatar based on the expected response included in the context information, and generates a motion of one of the estimated candidates. (15) The information processing system according to (14), wherein the second generation unit generates one of the estimated motion candidates in accordance with a selection by a user of the second terminal.(16) The information processing system according to any one of (2) to (15), wherein the second acquisition unit acquires the context information, and the second generation unit generates, as the motion of the second avatar, a motion that is different from the motion of the first avatar and can be combined with the motion of the first avatar, based on an expected response included in the context information. (17) The information processing system according to any one of (2) to (16), wherein the first acquisition unit acquires the input information based on data obtained by learning a relationship between information obtained by tracking a body movement of a user of the first terminal and a motion that reflects the tracked information on the first avatar. (18) The information processing system according to any one of (1) to (17), wherein the first acquisition unit acquires information identifying the plurality of second avatars, the first generation unit generates context information including expected responses that are responses expected from the plurality of second avatars, and the first transmission unit transmits the context information to the plurality of second avatars as control information for controlling motions of the plurality of second avatars. (19) An information processing system comprising: a second acquisition unit that acquires, from the first avatar, context information including expected responses that are responses expected from a second avatar with which the first avatar interacts in a virtual space, a second generation unit that generates a motion of the second avatar based on the expected responses included in the context information, and a second transmission unit that transmits the generated motion of the second avatar to the first avatar. (20) An information processing method including: a computer acquiring information identifying a second avatar with which a first avatar interacts in a virtual space and input information that triggers the motion of the first avatar; generating motion of the first avatar according to the input information; and generating context information including an expected response that is an expected response from the second avatar in response to the motion of the first avatar; and transmitting the context information to the second avatar as control information for controlling the motion of the second avatar.(21) An information processing program for causing a computer to function as an information processing system comprising: a first acquisition unit that acquires information identifying a second avatar with which a first avatar interacts in a virtual space and input information that triggers the motion of the first avatar; a first generation unit that generates motion of the first avatar according to the input information and generates context information including an expected response that is an expected response from the second avatar in response to the motion of the first avatar; and a first transmission unit that transmits the context information to the second avatar as control information for controlling the motion of the second avatar.
[0180] REFERENCE SIGNS LIST 10 User 10A HMD avatar 20 User 20A Mobile avatar 100 HMD 110 Communication unit 120 Memory unit 130 Control unit 131 Acquisition unit 132 Analysis unit 133 Generation unit 134 Replication unit 135 Transmission unit 200 Smartphone 230 Control unit 231 Acquisition unit 232 Analysis unit 233 Correction unit 234 Determination unit 235 Generation unit 236 Replication unit 237 Transmission unit 300 Cloud server
Claims
1. An information processing system comprising: a first acquisition unit that acquires information identifying a second avatar with which a first avatar interacts in a virtual space and input information that triggers the motion of the first avatar; a first generation unit that generates motion of the first avatar according to the input information and generates context information including an expected response that is an expected response from the second avatar in response to the motion of the first avatar; and a first transmission unit that transmits the context information to the second avatar as control information for controlling the motion of the second avatar.
2. The information processing system according to claim 1, comprising: a first terminal having the first acquisition unit, the first generation unit, and the first transmission unit; and a second terminal having a second acquisition unit, the second generation unit, and a second transmission unit, wherein the first acquisition unit acquires the input information by tracking the physical movements of a user of the first terminal.
3. The information processing system described in claim 2, wherein the second acquisition unit acquires the context information, the second generation unit generates a motion of the second avatar based on the expected response included in the context information and generates a correction motion required to complete the motion of the second avatar, and the second transmission unit transmits the motion of the second avatar and the correction motion to the first avatar.
4. The information processing system described in claim 2, wherein the first acquisition unit acquires information regarding the presence or absence of a pre-motion, which is a preliminary motion for completing the motion of the first avatar; the first generation unit, when determining that a pre-motion is required based on the information regarding the presence or absence of a pre-motion, generates a pre-motion for the first avatar and also generates second context information including a second expected response, which is an expected response in response to the pre-motion; and the first transmission unit transmits the second context information to the second avatar prior to transmitting the context information.
5. The information processing system described in claim 4, wherein the second acquisition unit acquires the second context information, and the second generation unit generates a motion to move or rotate the second avatar as a motion corresponding to the pre-motion based on the second expected response included in the second context information.
6. The information processing system of claim 2, wherein the second acquisition unit acquires the context information, and the second generation unit, when determining that the context information or the user of the first terminal that sent the context information meets a blocking condition specified on the second terminal, does not generate a motion of the second avatar corresponding to the expected response included in the context information.
7. The information processing system of claim 6, wherein the second generation unit does not generate a motion of the second avatar corresponding to the expected response contained in the context information when the context information is transmitted from outside a personal boundary defined around the second avatar.
8. The information processing system of claim 6, wherein the second generation unit does not generate a motion of the second avatar corresponding to the expected response contained in the context information when the context information is transmitted from outside the personal boundary defined around the second avatar and is transmitted from a user who has not previously interacted with the second avatar, a user who is not identified by the second avatar, or a user whose gaze is different from that of the second avatar.
9. The information processing system of claim 6, wherein the second generation unit does not generate a motion of the second avatar corresponding to the expected response included in the context information if the context information is sent from a user included in a block list set by the user of the second terminal, or if the user of the second terminal has disabled interaction with other users or does not allow interaction.
10. The information processing system of claim 2, wherein the second acquisition unit acquires the context information, and the second generation unit determines whether or not to generate a motion of the second avatar corresponding to the expected response included in the context information based on profile information which is information about the user who sent the context information or the first avatar.
11. The information processing system according to claim 1, comprising: a first terminal having the first acquisition unit, the first generation unit, and the first transmission unit; and a second terminal having a second acquisition unit, the second generation unit, and a second transmission unit; and the first acquisition unit acquires the input information by a UI (User Interface) operation or voice input by a user of the first terminal.
12. The information processing system of claim 2, wherein the second acquisition unit acquires the context information and the motion of the first avatar, and the second generation unit estimates the motion of the second avatar as a response expected from the second avatar to the motion of the first avatar, and generates the estimated motion of the second avatar as a response to the expected response included in the context information.
13. The information processing system according to claim 1, wherein the first generation unit generates the context information based on a UI (User Interface) operation or voice input by a user corresponding to the first avatar.
14. The information processing system of claim 2, wherein the second acquisition unit acquires the context information, and the second generation unit estimates multiple candidates for the motion of the second avatar based on the expected response included in the context information, and generates a motion of one of the multiple estimated candidates.
15. The information processing system according to claim 14, wherein the second generation unit generates one of the estimated motion candidates in accordance with a selection made by the user of the second terminal.
16. The information processing system of claim 2, wherein the second acquisition unit acquires the context information, and the second generation unit generates, based on the expected response included in the context information, a motion for the second avatar that is different from the motion of the first avatar and can be combined with the motion of the first avatar.
17. The information processing system of claim 2, wherein the first acquisition unit acquires the input information based on data learned from tracking information of the body movements of the user of the first terminal and the relationship between the tracked information and the motion that reflects the tracked information in the first avatar.
18. The information processing system of claim 1, wherein the first acquisition unit acquires information identifying the plurality of second avatars; the first generation unit generates context information including expected responses that are responses expected from the plurality of second avatars; and the first transmission unit transmits the context information to the plurality of second avatars as control information for controlling the motion of the plurality of second avatars.
19. An information processing system comprising: a second acquisition unit that acquires context information from a first avatar, the context information including an expected response that is an expected response from a second avatar with which the first avatar interacts in a virtual space; a second generation unit that generates a motion of the second avatar based on the expected response included in the context information; and a second transmission unit that transmits the generated motion of the second avatar to the first avatar.
20. An information processing method comprising: a computer acquiring information identifying a second avatar with which a first avatar interacts in a virtual space, and input information that triggers the motion of the first avatar; generating motion of the first avatar in accordance with the input information; generating context information including an expected response that is an expected response from the second avatar in response to the motion of the first avatar; and transmitting the context information to the second avatar as control information for controlling the motion of the second avatar.
Citation Information
Patent Citations
Information processing system, information processing method, and storage medium
JP2024018864A
Information processing system, information processing method, and storage medium
JP2024021028A