Information processing device, method and program

By using an information processing apparatus that generates reflected motion and sound for a second avatar in a virtual space based on user data, the processing load is reduced while maintaining user engagement and excitement, addressing the challenges faced by existing technologies.

JP2025090132APending Publication Date: 2025-06-17SONY GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023205170
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-05
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

Existing technologies face challenges in reducing the processing load of information processing apparatuses that provide virtual spaces, while maintaining the excitement and unity among users, especially when a large number of users participate in events.

Method used

The implementation of an information processing apparatus that generates a virtual space with a first avatar operated by a user and a second avatar, where the control unit generates reflected motion and sound to be reflected in the second avatar based on user data, thereby reducing processing load while maintaining user engagement.

Benefits of technology

This solution effectively reduces the processing load of the information processing apparatus while creating a more immersive and engaging virtual space experience for users, even with a large number of participants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025090132000001_ABST
    Figure 2025090132000001_ABST
Patent Text Reader

Abstract

To reduce the processing load when placing a user avatar that reflects a motion in a virtual space.SOLUTION: An information processing device comprises a control unit that controls the generation of a virtual space, where the virtual space includes a first avatar operated by a user and a second avatar different from the first avatar, and the control unit controls the generation of a reflection motion to be reflected in the second avatar based on user data obtained from the user.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an information processing apparatus, method, and program.

Background Art

[0002] In recent years, technologies for providing virtual spaces in which users' avatars are placed have become widespread. For example, there is a provision of a virtual space in which an avatar of a user who performs at an event such as a music live or a play and avatars of a plurality of users who watch the performance by the user are placed in an event venue in the virtual space. Each avatar reflects the user's motion. As a result, the user can obtain an experience as if participating in the event with other users.

[0003] An information processing apparatus that provides such a virtual space exchanges data with terminals used by each user. Therefore, when a large number of users participate in an event, the information processing apparatus is burdened with a processing load. Thus, technologies for reducing the processing load of an information processing apparatus that provides a virtual space have been developed. For example, Patent Document 1 below discloses a technology for reducing the processing load of an information processing apparatus that provides a virtual space by not receiving user motion data from some of the connected terminals.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, if the motion is not reflected in the avatars of some users who did not receive the motion data, it becomes difficult to create the heat generated by the motions of a large number of spectators in the real event venue in the virtual space.

[0006] Therefore, the present disclosure proposes a novel and improved technology capable of reducing the processing load when arranging user avatars reflecting motion in a virtual space.

Means for Solving the Problem

[0007] According to the present disclosure, there is provided an information processing apparatus including a control unit that performs control to generate a virtual space, the virtual space including a first avatar operated by a user and a second avatar different from the first avatar, and the control unit performing control to generate a reflected motion to be reflected in the second avatar based on user data obtained from the user.

[0008] Also, according to another aspect of the present invention for solving the above problem, there is provided an information processing apparatus including a control unit that performs control to generate a virtual space, the virtual space including a first avatar operated by a user and a second avatar different from the first avatar, and the control unit performing control to generate sound to be reflected in a partial space where the second avatar exists among one or more partial spaces in the virtual space based on user data obtained from the user.

[0009] Also, according to another aspect of the present invention for solving the above problem, there is provided a method executed by a processor, including performing control to generate a virtual space, the virtual space including a first avatar operated by a user and a second avatar different from the first avatar, and performing control to generate a reflected motion to be reflected in the second avatar based on user data obtained from the user.

[0010] Also, according to another aspect of the present invention to solve the above problems, a computer is operated as a control unit that controls the generation of a virtual space. The virtual space includes a first avatar operated by a user and a second avatar different from the first avatar. The control unit performs control to generate a reflection motion to be reflected on the second avatar based on user data obtained from the user. A program is provided.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Mode for Carrying Out the Invention

[0012] Hereinafter, preferred embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. In the present specification and drawings, components having substantially the same functional configuration are denoted by the same reference numerals, and redundant description is omitted.

[0013] Note that the description will be made in the following order. 1. Overview 2. Configuration Example of Server 10 3. Operation Processing Example 4. Modification Examples 5. Hardware Configuration 6. Supplementary

[0014] <1. Overview> First, the overall configuration of the information processing system according to an embodiment of the present disclosure will be described with reference to FIG. 1. FIG. 1 is a diagram for explaining the overall configuration of the information processing system 1 according to an embodiment of the present disclosure.

[0015] As shown in FIG. 1, an information processing system 1 according to an embodiment of the present disclosure includes a server 10 and user terminals 20 (user terminals 20A, 20B, 20C...). As shown in FIG. 1, the server 10 and the user terminals 20 are configured to be communicable via a network 30.

[0016] (Server 10) The server 10 has a function of providing a virtual space to a user who uses the user terminal 20. In the virtual space, concerts, plays, games, and various other events can be held. In this embodiment, as an example, an example in which the event is a concert will be mainly described.

[0017] A plurality of users participate in the event. In this embodiment, each user participates in the event either as a performer who performs a performance in the event or as an audience who watches the performance performed by the performer. Hereinafter, a user who is a performer is also referred to as a "performer user". Also, a user who is an audience is also referred to as an "audience user".

[0018] Note that the number of users participating in the event (the number of performer users and the number of audience users) is not limited. For example, the number of users participating in the event, particularly the number of audience users, may be limited according to the processing load of the server 10.

[0019] The server 10 generates a virtual space in which user avatars, which are avatars corresponding to each user, are arranged. More specifically, the user avatars include performer avatars corresponding to performer users and audience avatars corresponding to audience users. Each user avatar may be a 3D realistic avatar or a 3D character avatar.

[0020] Each user avatar is operated by the corresponding user. For example, each user avatar can move within the virtual space based on the operation of the user terminal 20 by the corresponding user. More specifically, the performer avatar may move on the stage provided in the concert venue of the virtual space according to the user operation. Also, the audience avatar may move within the audience area outside the stage of the concert venue according to the user operation.

[0021] In addition, the motion obtained from the corresponding user is reflected in real time in each user avatar. The motion data of the motion obtained from each user is transmitted from the user terminal 20 to the server 10 in real time.

[0022] Also, the server 10 acquires sound from each user in order to provide the acoustics of the virtual space. Hereinafter, the sound acquired from each user is also referred to as "user voice".

[0023] Among the user voices, the sound acquired from the performer user includes the sound emitted by the performer user. The sound emitted by the performer user may include the singing voice of the performer user, the sound of the musical instrument played by the performer user, the spoken voice of the performer user, and the conversation voice between the performer users. Hereinafter, the sound emitted by the performer user is also referred to as "performer voice".

[0024] Furthermore, the server 10 may acquire the sound source of the music used in the event from the performer user. Hereinafter, the sound source of the music used in the event is also simply referred to as the "sound source". The sound source may be, for example, the sound source of a music piece that does not include the singing part by the performer user. Hereinafter, the sound source of the music piece excluding the singing part by the performer user is also referred to as the "off-vocal sound source". By superimposing and playing the real-time singing voice of the performer user on the off-vocal sound source, a concert in the virtual space may be realized. Also, the sound source may be the sound source of a music piece that includes the singing part by the performer user or the like. The server 10 may acquire the performer voice of the performer user singing while playing the off-vocal sound source. That is, the server 10 may acquire the sound in which the off-vocal sound source and the performer voice are superimposed. Also, the sound source may include the singing part by the pre-recorded performer user or other users.

[0025] Also, among the user voices, the sound acquired from the spectator user includes the voices of the spectator user cheering for the performance by the performer user, the sound of finger snapping, the sound of clapping, and the sound emitted by props such as musical instruments held by the spectator user. Hereinafter, the sound emitted by the spectator user is also referred to as the "spectator voice".

[0026] The server 10 transmits virtual space content data including data for rendering the video of the virtual space and data for playing the sound of the virtual space to the user terminal 20.

[0027] The data for rendering the video of the virtual space is data related to virtual objects such as user avatars arranged in the virtual space, map information of the virtual space, position information of each user avatar in the virtual space, and the like. In the present embodiment, it is assumed that the position information of the user avatar is information on the coordinates of the user avatar in the virtual space. Hereinafter, the coordinates of the user avatar in the virtual space are referred to as "user avatar coordinates". Also, the coordinates of the performer avatar in the virtual space are referred to as "performer avatar coordinates". Also, the coordinates of the spectator avatar in the virtual space are referred to as "spectator avatar coordinates".

[0028] The data for reproducing the sound in the virtual space is, for example, the user voice of each user and the sound data of the sound source of the music used in the event.

[0029] (User terminal 20) The user terminal 20 is an information processing terminal used by the user. The user in this embodiment corresponds to either a performer user or an audience user. Each user participates in an event held on the virtual space generated by the server 10.

[0030] In FIG. 1, an example is shown in which the user terminal 20 is a PC (Personal Computer), and an HMD (Head Mounted Display) 21 (21A, 21B, 21C ···) is connected to each user terminal 20. The user terminal 20 controls to display the video of the virtual space on the connected HMD 21.

[0031] Note that the configuration of the user terminal 20 is not limited to the example shown in FIG. 1, and may be various devices alone such as a tablet terminal, a smartphone, an HMD, or a game terminal. Or, the user terminal 20 may be configured by a combination of the above various devices. Or, the user terminal 20 may be configured by connecting various terminals such as a display, a wearable device, a motion capture device, a camera, a microphone, and a speaker to the above various devices.

[0032] The user terminal 20 exchanges various data with the server 10 through the network 30. The user terminal 20 transmits user data obtained from the user to the server 10. The user data may include, for example, operation data representing the operation content of the user regarding the virtual space.

[0033] The user terminal 20 acquires operations by the user regarding the virtual space. For example, the user terminal 20 acquires operations on a user avatar placed in the virtual space. Operations on the user avatar are, for example, operations for controlling the position of the user avatar, the facial expression, gestures (emotional expressions using hands or the whole body), and motion states (sitting, standing, jumping, running, etc.). Operations on the user avatar may be detected, for example, from a controller operated by the user. Also, the operation information of the viewer avatar may be input from various operation input units such as a keyboard, a mouse, and a touch pad.

[0034] As an operation by the user, the user terminal 20 may acquire a motion of the user in which the position and posture of the body part, etc. of the user are represented by coordinates, rotation angles, etc. The motion data of the user is an example of operation data. The motion of the user may be detected by a motion capture device or may be detected by a camera. The detected motion of the user is reflected in the user avatar in real time. Note that the above-described operations on the user avatar can also be reflected in the avatar in real time. Hereinafter, when referring to "the motion data of the user", it may include not only the motion information of the user obtained by sensing but also the operation information regarding the movement of the avatar obtained by the controller.

[0035] The user terminal 20 outputs the acquired motion data of the user to the server 10. For example, the user terminal 20 outputs the motion data reflected as the motion of the user avatar to the server 10 in real time.

[0036] Also, the user data may include audio data of the user voice. The user terminal 20 may acquire the user voice by a voice input device such as a microphone. The user terminal 20 outputs the audio data of the user voice to the server 10 in real time.

[0037] In particular, in the user terminal 20 especially used by the performer user, the user terminal 20 may acquire the performer voice of the performer user who sings while playing a sound source. In this case, the user terminal 20 may output sound data including the performer voice and the sound source to the server 10. Note that the user terminal 20 may upload the sound data of the sound source to the server 10 in advance at the stage of event preparation.

[0038] Also, the user terminal 20 receives content data of the virtual space from the server 10. The user terminal 20 outputs to the user by combining the video and sound of the virtual space according to the content data received from the server 10. Thereby, the user can enjoy the event held on the virtual space.

[0039] Specifically, the user terminal 20 arranges virtual objects based on the map information of the virtual space and draws the video of the virtual space from the user's perspective. The user's perspective may be the perspective of the user avatar or an overhead perspective including the user avatar in the field of view. Note that the server 10 may generate the video from an arbitrary perspective in the virtual space and transmit it to the user terminal 20. The arbitrary perspective may be, for example, the user's perspective or an overhead perspective from which the entire event venue can be seen.

[0040] Also, the user terminal 20 superimposes the user voices of each user and plays the sound of the virtual space. For example, the user terminal 20 may play the sound of the virtual space according to the coordinates of the user avatar. For example, the user terminal 20 may play the sound of the virtual space so that the closer the distance between the coordinates of the user avatar of the user using the user terminal 20 and the coordinates of the user avatar is to the spectator voice of the spectator user, the louder the superimposed volume becomes.

[0041] In addition, for the voice corresponding to the user avatar of a user (hereinafter also referred to as "User B") who is a friend of the user (hereinafter also referred to as "User A") using the user terminal 20, or for the voice of an artist (i.e., a performer user) that User A likes, which is so-called "favorite", it is assumed that there is a need to superimpose it with a larger volume. Therefore, for these voices, even when the distance between User A's user avatar and the user avatar of User B or the above artist is far, a larger volume may be superimposed. For example, in such a case, the voice corresponding to the user avatar of User B or the above artist may be superimposed at a certain volume that is easy for User A to hear regardless of the above distance, or the volume may be adjusted so that it is superimposed at a larger volume than other users existing at the same distance. Whether User A and User B are friends may be specified from the user attribute information. The user attribute information is, for example, user information (such as age, gender, height, grouping information with specific users other than oneself, donation information, information on the user's preferences, etc.) set by the user from an application or the like prior to an event. The artist that User A likes may be specified from the above-mentioned user attribute information (such as donation information, etc.).

[0042] Also, regardless of the user avatar coordinates of the performer user, the user terminal 20 may superimpose the performer voice so that the user can easily hear the performer voice. The server 10 may generate the acoustics of the virtual space (for example, the acoustics according to the user avatar coordinates) and transmit the sound data of the acoustics to the user terminal 20.

[0043] (Sorting out the problems) Here, as described above, when the server 10 transmits and receives the motion data of a large number of users and the sound data of user voices in real time to and from each user terminal 20, a problem of processing load may occur. Therefore, it is conceivable to reduce the processing load by suppressing the amount of data transmitted and received between the server 10 and the user terminal 20 by limiting the number of users participating in the event to a small number.

[0044] However, when the number of users participating in the event is limited to a small number, it becomes difficult to create the excitement generated by the motions and cheers of a large number of spectators in the virtual space. Usually, in an event, for example, when a large number of spectators move in unison to the singer's performance or send cheers, a sense of unity among the spectators is created. And when the spectators who feel such a sense of unity get even more excited and move or send cheers, excitement is generated.

[0045] FIG. 2 is a diagram for explaining a comparative example with the present embodiment of the video and sound of the virtual space generated when the number of users participating in the event is limited to a small number. Here, an example where the number of spectator users participating in the event is limited to a small number is shown. The video 210 of the virtual space is a video in which the performer avatars P (P1, P2, and P3) and the spectator avatars O (O1, O2, etc.) are arranged in the concert venue V of the virtual space. In the example described with reference to FIG. 2, the sound 211 of the virtual space generated by superimposing the sounds emitted by the users corresponding to each of the performer avatar P and the spectator avatar O is combined with the video 210 of the virtual space.

[0046] In the example shown in FIG. 2, since the number of spectator avatars O is small compared to the size of the concert venue V in the virtual space, it is difficult for spectator users to feel a sense of unity with the spectator avatars O of other spectator users existing around their own spectator avatar O. Also, although it is conceivable to generate the virtual space by adjusting the size of the concert venue V in the virtual space according to the number of spectator avatars O, it is also conceivable that due to the small number of spectator avatars O, the compelling force of the cheering is not strong enough to feel a sense of unity.

[0047] Therefore, in addition to the spectator avatar O operated by the spectator user, the server 10 according to the present embodiment generates a spectator avatar not operated by the spectator user by the server 10 and arranges it in the virtual space. Hereinafter, in order to distinguish between the spectator avatar O operated by the spectator user and the spectator avatar not operated by the spectator user, the spectator avatar O operated by the spectator user is also referred to as the "operated spectator avatar O". The operated spectator avatar O is the first avatar according to the present embodiment. Further, hereinafter, the spectator avatar generated by the server 10 and not operated by the spectator user is also referred to as the "generated spectator avatar". The generated spectator avatar is the second avatar according to the present embodiment.

[0048] Hereinafter, the video and sound of the virtual space generated by the server 10 according to the present embodiment will be described with reference to FIGS. 3 and 4.

[0049] FIG. 3 is a diagram for explaining an example of the video and sound of the virtual space generated by the server 10 according to the present embodiment. In FIG. 3, a video 212 of the virtual space is shown. The video 212 of the virtual space is a video of the virtual space drawn by the user terminal 20 based on the content data of the virtual space generated by the server 10.

[0050] Similar to the video 210 of the virtual space described with reference to FIG. 2, the video 212 of the virtual space is a video of the concert venue V of the virtual space in which the performer avatars P (P1, P2, and P3) and the operated spectator avatars O (O1, O2, etc.) are arranged. Then, in the virtual space of the video 212 of the virtual space, generated spectator avatars G (G1, G2, etc.) are further arranged.

[0051] FIG. 4 is a diagram for explaining an example of generating a virtual space by the server 10 according to the present embodiment. As shown in FIG. 4, the server 10 generates a reflected motion 2010 based on the user data 2001. The user data 2001 includes, as described above, the motion data of each user, user voices (audience voices and performer voices), sound data such as the sound source of the music used in the event, and the user avatar coordinates in the virtual space.

[0052] The reflected motion 2010 is a motion that is reflected on the generated audience avatar G. The server 10 may generate the reflected motion 2010 for each of the plurality of generated audience avatars G existing at different positions on the virtual space. The server 10 generates content data by including, in the data for drawing the video of the virtual space, the appearance data of the generated audience avatar G and the data about the generated audience avatar G such as the reflected motion 2010.

[0053] Thereby, the user terminal 20 can reproduce the video 212 of the virtual space in which the user avatars (performer avatar P and manipulated audience avatar O) reflecting the motions of each user generated based on the user data 2001 and the generated audience avatar G reflecting the reflected motion 2010 are merged. Such a video 212 of the virtual space may be generated by the server 10, transmitted to the user terminal 20 as content data, and reproduced by the HMD 21. Details of the generation of the reflected motion 2010 will be described later.

[0054] According to the server 10 according to the present embodiment, as shown in FIG. 3, compared with the example described with reference to FIG. 2, it is possible to realize the generation of a virtual space in which a larger number of audiences participate while reducing the processing load of the server 10. As a result, it is possible to expect to create the enthusiasm of the audience on the virtual space, and thus it is possible to improve the entertainment property of the event held on the virtual space.

[0055] Also, as shown in FIG. 4, the server 10 according to the present embodiment generates area sound 2020 based on user data 2001.

[0056] The area sound 2020 is sound that is reflected in the audience area where the generated spectator avatar G exists in the virtual space. The audience area is an area set by pre-dividing an area where the manipulated spectator avatar O and the generated spectator avatar G can exist in the virtual space. The audience area is an example of a partial space that is a part of the virtual space. For example, in the example shown in FIG. 3, an example where audience areas A1 to A4 are set in the virtual concert hall V is shown. Hereinafter, when the audience areas A1 to A4 are not particularly distinguished, they are referred to as "audience area A".

[0057] The manipulated spectator avatar O may move within the audience area A and between a plurality of audience areas A based on a user operation. Also, the generated spectator avatar G may move within the audience area A and between a plurality of audience areas A according to the control by the server 10. However, seats may be provided in advance in the virtual space, and the positions of the manipulated spectator avatar O and the generated spectator avatar G may be fixed.

[0058] Details of the area sound 2020 will be described later, but the area sound 2020 may include sounds such as cheers. The server 10 generates content data by including the sound data of the area sound 2020 in the data for playing the sound of the virtual space. The user terminal 20 plays back the area sound 2020 superimposed on the user voice of each user when playing back the sound of the virtual space.

[0059] The user terminal 20 may superimpose only the area sound 2020 reflected in the spectator area A where the operated spectator avatar O of the user using the user terminal 20 exists on the user voice. Further, the user terminal 20 may superimpose a plurality of area sounds 2020 on the user voice based on the relationship between the user avatar coordinates of the user using the user terminal 20 and the location of the spectator area A. For example, the user terminal 20 may superimpose each area sound 2020 on the user voice so that the volume of the area sound 2020 to be reproduced increases as the distance between the user avatar coordinates and the location in the spectator area A is closer.

[0060] In the examples shown in FIGS. 3 and 4, the sound 213 of the virtual space in which the area sound 2020 is merged with the user voice is combined with the video 212 of the virtual space. As a result, since the types and volumes of cheering and the like in the sound of the virtual space increase, it is possible to further improve the entertainment of the event performed on the virtual space.

[0061] Such a sound 213 of the virtual space may be generated by the server 10, transmitted to the user terminal 20 as content data, and combined with the video 212 of the virtual space by the HMD 21 and reproduced. Further, the video 212 and the sound 213 of the virtual space may be combined by the server 10 and transmitted to the user terminal 20 as content data, and reproduced by the HMD 21. Note that the user terminal 20 may continuously acquire the content data in which the video 212 and the sound 213 of the virtual space are combined from the server 10 in a streaming manner and reproduce it by the HMD 21.

[0062] Subsequently, the specific configuration of each device included in the information processing system 1 according to the present embodiment will be described with reference to the drawings.

[0063] <2. Configuration example of server 10> First, a configuration example of the server 10 will be described with reference to FIG. 5. FIG. 5 is a block diagram showing an example of the configuration of the server 10 according to the present embodiment.

[0064] As shown in FIG. 5, the server 10 includes a communication unit 110, a control unit 120, and a storage unit 130.

[0065] (Communication Unit 110) The communication unit 110 communicates with the user terminal 20 by wire or wirelessly to transmit and receive data. The communication unit 110 can perform communication using, for example, a wired / wireless LAN, Wi-Fi (registered trademark), Bluetooth (registered trademark), infrared communication, or a mobile communication network (4G (4th generation mobile communication system), 5G (5th generation mobile communication system)).

[0066] The communication unit 110 receives user data 2001 of each user from each user terminal 20, for example. Further, the communication unit 110 transmits content data to each user terminal 20.

[0067] (Control Unit 120) The control unit 120 functions as an arithmetic processing device and a control device, and controls the overall operation within the server 10 according to various programs. The control unit 120 is realized by an electronic circuit such as a CPU (Central Processing Unit) or a microprocessor, for example. Further, the control unit 120 may include a ROM (Read Only Memory) that stores programs and arithmetic parameters to be used, and a RAM (Random Access Memory) that temporarily stores parameters that change as appropriate.

[0068] The control unit 120 performs appropriate processing based on the data received from an external device, and performs storage control to the storage unit 130, data transmission control to an external device, and the like.

[0069] Further, the control unit 120 also functions as a motion generation unit 121, an acoustic generation unit 122, and a virtual space generation unit 123.

[0070] (Motion Generation Unit 121) Based on each user data 2001 received by the communication unit 110, the motion generation unit 121 generates a reflection motion 2010 to be reflected on the generated spectator avatar G. The motion generation unit 121 generates the reflection motion 2010 so that the generated spectator avatar G performs a motion corresponding to the situation of the event.

[0071] The motion generation unit 121 may generate, for example, a reflection motion 2010 to be reflected in a predetermined number of frames. The predetermined number is not particularly limited, and may be set, for example, so that the reflection motion 2010 reproduced in 1 second to several seconds is generated. The motion generation unit 121 repeatedly executes the generation process of the reflection motion 2010. For example, by repeating the generation process during the event, the reflection motion 2010 may be continuously reflected on the generated spectator avatar G during the event.

[0072] The motion generation unit 121 may generate the reflection motion 2010 using a machine learning model generated using a machine learning technique such as a DNN (Deep Neural Network). FIG. 6 is a diagram for explaining an example in which the motion generation unit 121 generates the reflection motion 2010 using a machine learning model.

[0073] As shown in FIG. 6, the reflection motion 2010 is generated using a pre-seed motion generation model M1 and a reflection motion generation model M2. The pre-seed motion generation model M1 is the first machine learning model according to the present embodiment. The reflection motion generation model M2 is the second machine learning model according to the present embodiment.

[0074] The pre-seed motion generation model M1 is a machine learning model that outputs the pre-seed motion 2011 using the sound data of the sound source B as input data. The reflection motion generation model M2 is a machine learning model that generates the reflection motion 2010 using the pre-seed motion 2011, user data 2001, and generated avatar coordinates 1001 as input data. The user data 2001 includes the performer data 2001A, which is the user data 2001 of the performer user, and the audience data 2001B, which is the user data 2001 of the audience user.

[0075] The pre-seed motion 2011 is the (seed) motion that serves as the basis when the reflection motion generation model M2 outputs the reflection motion 2010. The motion generation unit 121 inputs the pre-seed motion 2011 output from the pre-seed motion generation model M1 into the reflection motion generation model M2. Then, based on the pre-seed motion 2011, the reflection motion generation model M2 modifies the pre-seed motion 2011 based on the user data 2001 and the generated avatar coordinates 1001 to generate the reflection motion 2010.

[0076] The motion generation unit 121 may input the sound data of the sound source B obtained in real time from the user terminal 20 used by the performer user into the pre-seed motion generation model M1. As described above, the sound source B may or may not include a singing part. That is, the motion generation unit 121 may input, as the sound source B, the sound data in which the singing voice of the performer user is superimposed on the off-vocal sound source. Also, the motion generation unit 121 may input the sound source B obtained in advance from the user terminal 20 at the stage of event preparation into the pre-seed motion generation model M1. In this case, the sound source B may include a pre-recorded singing part. When the singing part is included in the sound source B, the pre-seed motion generation model M1 may generate the pre-seed motion 2011 according to the content of the lyrics. Thereby, it is possible to make the generated spectator avatar G take the motion generated along with the meaning of the music sung by the performer user.

[0077] The pre-seed motion generation model M1 may be a machine learning model that generates the pre-seed motion 2011 according to the melody of the sound source B.

[0078] As an example, the pre-seed motion generation model M1 may generate the pre-seed motion 2011 such that the pre-seed motion 2011 becomes a motion that moves in accordance with the rhythm of the sound source B.

[0079] Also, as another example, the pre-seed motion generation model M1 may generate the pre-seed motion 2011 in accordance with the melody for each part of the sound source B. For example, the pre-seed motion 2011 may be generated such that the motion in the intro part of the sound source B is smaller compared to other parts. Also, the pre-seed motion generation model M1 may generate the pre-seed motion 2011 such that the motion becomes larger, for example, over the chorus part of the sound source B.

[0080] When singing is performed at an event, it is common for the audience to move in accordance with the music being sung. Therefore, by generating the pre-seed motion 2011 that serves as the basis for the reflected motion 2010 based on the sound data of the sound source B, it is possible to generate a reflected motion 2010 suitable for the situation of the event.

[0081] Note that the pre-seed motion generation model M1 may be learned in any way as long as it can output a pre-seed motion 2011 appropriate as a motion for the singing using the sound source B at the event based on the sound source B.

[0082] By using the pre-seed motion generation model M1, the motion generation unit 121 can generate the reflected motion 2010 by making changes to the pre-seed motion 2011 generated in accordance with the melody of the sound source B. Thereby, it is possible to cause the spectator avatar G to take a motion appropriate for the music sung by the performer user.

[0083] The motion generation unit 121 inputs the preliminary seed motion 2011 output from the preliminary seed motion generation model M1 into the reflected motion generation model M2. Further, the motion generation unit 121 inputs the user data 2001 and the generated avatar coordinates 1001 into the reflected motion generation model M2 together with the preliminary seed motion 2011. The generated avatar coordinates 1001 are coordinates representing the position of the generated spectator avatar G in the virtual space.

[0084] The motion generation unit 121 inputs the user data 2001 obtained from a predetermined time ago to the present into the reflected motion generation model M2. The predetermined time may be, for example, 1 second to several seconds.

[0085] The motion generation unit 121 may input the user data 2001 of all users participating in the event into the reflected motion generation model M2.

[0086] Further, the motion generation unit 121 may sample a part of the user data 2001 from the user data 2001 of all users participating in the event and input it into the reflected motion generation model M2. For example, the motion generation unit 121 may sample a part of the spectator data 2001B from the spectator data 2001B of all spectator users participating in the event and input it into the reflected motion generation model M2.

[0087] Whether to sample the spectator data 2001B may be determined based on the processing load of the server 10. The processing load of the server 10 may be represented by, for example, the amount of memory used or the amount of power consumed by the GPU (Graphics Processing Unit) or CPU (Central Processing Unit) of the server 10. For example, when the processing load of the server 10 is greater than the threshold, the motion generation unit 121 may sample the spectator data 2001B and generate the reflected motion 2010.

[0088] Sampling of the spectator data 2001B may be performed, for example, by randomly selecting a predetermined proportion of the manipulated spectator avatars O for each spectator area A.

[0089] The generated avatar coordinates 1001 are the coordinates of the generated spectator avatar G in the virtual space. The generated avatar coordinates 1001 are an example of the position of the generated spectator avatar G in the virtual space.

[0090] The reflection motion generation model M2 may generate the reflection motion 2010 such that, for example, the closer the user avatar coordinates are to the generated avatar coordinates 1001, the greater the influence of the user data 2001 corresponding to the user avatar coordinates on the generation of the reflection motion 2010. The reflection motion generation model M2 may generate the reflection motion 2010 using only the user data 2001 of users whose distance between the generated avatar coordinates 1001 and the user avatar coordinates is equal to or less than a predetermined distance. The reflection motion generation model M2 may generate the reflection motion 2010 by weighting each user data 2001 according to the distance between the generated avatar coordinates 1001 and the user avatar coordinates.

[0091] It is assumed that the spectator takes motion under the stronger influence of the performers and spectators existing near the spectator himself / herself than the influence of the performers and spectators existing far away. Therefore, by generating the reflection motion 2010 based on the generated avatar coordinates 1001 and the user avatar coordinates, it is possible to cause the generated spectator avatar G to perform a motion appropriate to the situation of the event.

[0092] When the generated avatar coordinates 1001 change dynamically, the motion generation unit 121 acquires the generated avatar coordinates 1001 at each generation process and generates the reflection motion 2010. Here, the generated avatar coordinates 1001 may change according to the reflection motion 2010. On the other hand, when the generated avatar coordinates 1001 are fixed throughout the event, the motion generation unit 121 may use the same coordinates for the generation process of the reflection motion 2010 that is continuously executed during the event.

[0093] In the above description, an example was given in which each user data 2001 is weighted according to the distance between the generated avatar coordinates 1001 and the user avatar coordinates to generate the reflected motion 2010. However, the example of weighting the user data 2001 is not limited to this. For example, the reflected motion generation model M2 may weight each user data 2001 according to the feature information of the preset generated spectator avatar G to generate the reflected motion 2010. The feature information of the generated spectator avatar G may be, for example, age, gender, height, grouping information with a specific user, or information on a favorite artist, which are set as features of the generated spectator avatar G.

[0094] The grouping information with a specific user may be, for example, information indicating that the generated avatar and the specific user are friends. In this case, based on the operation of the user terminal 20 by the user, the generated avatar and the user may be set as friends.

[0095] As a more specific example, the reflected motion generation model M2 may weight each user data 2001 according to the user attribute information of the user who is in a friendship relationship with the generated spectator avatar G. For example, the reflected motion generation model M2 may weight the performer data 2001A of the performer user who matches the favorite artist of the user who is in a friendship relationship with the generated spectator avatar G so that the influence on the generation of the reflected motion 2010 becomes greater. Also, the reflected motion generation model M2 may weight the user data 2001 of the user who is in a friendship relationship with the generated spectator avatar G and the users existing around that user so that the influence on the generation of the reflected motion 2010 becomes greater.

[0096] As another example, the reflection motion generation model M2 may calculate the degree of match between the artist preferred by the generated spectator avatar G and each performer user, and weight the performer data 2001A according to the degree of match. Then, the reflection motion generation model M2 may generate the reflection motion 2010 such that the reflection motion 2010 is more greatly affected by the performer data 2001A of the performer user with a higher degree of match.

[0097] Subsequently, the generation of the reflection motion 2010 based on various data included in the user data 2001 by the reflection motion generation model M2 will be described. The reflection motion generation model M2 generates the reflection motion 2010 based on the performer data 2001A included in the user data 2001. For example, the reflection motion generation model M2 generates the reflection motion 2010 based on the performer motion data 2002A among the performer data 2001A. The reflection motion generation model M2 may generate the reflection motion 2010 such that the generated spectator avatar G moves more greatly as the motion represented by the performer motion data 2002A is larger.

[0098] Also, the reflection motion generation model M2 may generate the reflection motion 2010 so as to resemble the motion represented by the performer motion data 2002A and the reflection motion 2010. In particular, when the motion represented by the performer motion data 2002A is a specific motion (such as a motion of clapping hands or a motion of waving hands at a certain rhythm), the reflection motion generation model M2 may generate the reflection motion 2010 such that the generated spectator avatar G performs a specific motion.

[0099] It is assumed that in an event, the audience is affected by the motion of the performer and moves. Therefore, by generating the reflection motion 2010 based on the performer motion data 2002A, it is possible to cause the generated spectator avatar G to perform a motion suitable for the situation of the event.

[0100] Further, the reflection motion generation model M2 may generate a reflection motion 2010 based on the speaker voice data 2003A. For example, the reflection motion generation model M2 may generate the reflection motion 2010 such that the generated viewer avatar G moves more greatly as the volume of the speaker voice represented by the speaker voice data 2003A is larger.

[0101] Also, when the speaker voice data 2003A includes the voice of a specific word spoken by the speaker user, the reflection motion generation model M2 may generate the reflection motion 2010 such that the generated viewer avatar G performs a specific motion. For example, when the speaker voice data 2003A includes the voice of the word "clap" spoken by the speaker user, the reflection motion generation model M2 may generate the reflection motion 2010 such that the generated viewer avatar G claps.

[0102] In an event, it is assumed that the viewers are affected by the sound emitted by the speaker and move. Therefore, by generating the reflection motion 2010 based on the speaker voice data 2003A, the generated viewer avatar G can be made to perform a motion corresponding to the situation of the event.

[0103] Further, the reflection motion generation model M2 generates the reflection motion 2010 based on the viewer data 2001B among the user data 2001.

[0104] For example, the reflection motion generation model M2 generates a reflection motion 2010 based on the viewer motion data 2002B included in the viewer data 2001B. The reflection motion generation model M2 may generate the reflection motion 2010 such that the generated viewer avatar G moves more greatly as the motion represented by the viewer motion data 2002B is greater. Also, the reflection motion generation model M2 may generate the reflection motion 2010 so as to resemble the motion represented by the viewer motion data 2002B and the reflection motion 2010. In particular, when the motion represented by the viewer motion data 2002B is a specific motion (such as a motion of clapping hands or a motion of waving hands at a certain rhythm), the reflection motion 2010 may be generated such that the generated viewer avatar G performs the specific motion.

[0105] In an event, it is assumed that a viewer takes a motion under the influence of the motions of other viewers. Therefore, by generating the reflection motion 2010 based on the viewer motion data 2002B, it is possible to cause the generated viewer avatar G to perform a motion corresponding to the situation of the event.

[0106] Also, the reflection motion generation model M2 may generate the reflection motion 2010 based on the viewer voice data 2003B. For example, the reflection motion generation model M2 may generate the reflection motion 2010 such that the generated viewer avatar G moves more greatly as the volume of the viewer voice represented by the viewer voice data 2003B is greater.

[0107] In an event, it is assumed that a viewer takes a motion under the influence of the sounds made by other viewers. Therefore, by generating the reflection motion 2010 based on the viewer voice data 2003B, it is possible to cause the generated viewer avatar G to perform a motion corresponding to the situation of the event.

[0108] Although an example in which the user data 2001 is input to the reflection motion generation model M2 has been described so far, the virtual space information including the video and sound of the virtual space generated based on the user data 2001 may be input to the reflection motion generation model M2 together with the user data 2001 or instead of the user data 2001.

[0109] The video and sound of the virtual space reproduced on the user terminal 20 are generated based on the content data managed by the virtual space generation unit 123 described later. Therefore, the motion generation unit 121 may reproduce the video and sound of the virtual space generated on each user terminal 20 based on the content data managed by the virtual space generation unit 123. Then, the motion generation unit 121 may generate the reflection motion 2010 based on the reproduced video and sound of the virtual space. However, the video and sound of the virtual space generated on each user terminal 20 may be acquired again by the server 10.

[0110] It is assumed that the viewer takes motion based on the motion of the viewer and the performer reflected in his own eyes and the sound entering his own ears. Therefore, by generating the reflection motion 2010 based on the video and sound of the virtual space reproduced on the user terminal 20 of the user participating in the event, particularly the viewer user, it is possible to cause the generated viewer avatar G to perform a motion suitable for the situation of the event.

[0111] Further, the reflection motion generation model M2 may generate the reflection motion 2010 based on the user attribute information 131 stored in the storage unit 130 described later. The user attribute information 131 is an example of the user data 2001. The user attribute information 131 is information representing the attributes of each user. The user attribute information 131 may be, for example, information regarding the age, gender, or preference regarding performance of each user. The information regarding the preference regarding performance may be, for example, information such as the music the user likes and the music videos the user likes. The user attribute information 131 is acquired in advance by the communication unit 110 based on the operation of the user terminal 20 by the user.

[0112] The motions of the audience may tend to vary depending on factors such as age, gender, or preferences regarding the performance. Therefore, by generating the reflected motion 2010 based on the user attribute information 131 of each user, it is possible to cause the generated spectator avatar G to perform a motion that matches the users participating in the event.

[0113] Also, the reflected motion generation model M2 may generate the reflected motion 2010 according to the preset characteristic information of the generated spectator avatar G. For example, the reflected motion generation model M2 may generate the reflected motion 2010 such that the motion of a user in a friendship relationship is more strongly influenced than that of a user not in a friendship relationship with the generated spectator avatar G.

[0114] As another example, when the distance between the generated avatar coordinates 1001 and the performer avatar coordinates of a performer user who matches the favorite artist of the generated spectator avatar G is shorter than a predetermined distance, the reflected motion generation model M2 may generate the reflected motion 2010 such that a larger movement is performed.

[0115] In this way, by generating the reflected motion 2010 according to the characteristic information of the generated spectator avatar G by the reflected motion generation model M2, the generated spectator avatar G can be given individuality, so a more realistic audience can be reproduced.

[0116] Although it has been described so far that various data included in the user data 2001 are used for the reflected motion generation model M2 to generate the reflected motion 2010, there is no particular limitation on which data is used to generate the reflected motion 2010.

[0117] For example, the reflection motion generation model M2 may generate the reflection motion 2010 using only a part of the data among the performer data 2001A and the audience data 2001B. In particular, the performer data 2001A may be prioritized over the audience data 2001B and used to generate the reflection motion 2010. This is because it is assumed that the audience participates in the event with the greatest interest in the performer's motion and voice.

[0118] Also, the performer motion data 2002A and the audience motion data 2002B may be prioritized over the performer voice data 2003A and the audience voice data 2003B and used to generate the reflection motion 2010. This is because it is assumed that the audience's motion is most affected by the motion of the performer and other audience members.

[0119] When generating the reflection motion 2010 using a plurality of the data included in the user data 2001, the reflection motion 2010 may be generated by weighting each data. For example, each data may be weighted so that the performer motion data 2002A is most reflected in the reflection motion 2010. For example, priorities may be set for each of the plurality of data included in the user data 2001. The priority may be set as an integer value from 0 to 10, for example, with 0 being the minimum and 10 being the maximum. The reflection motion 2010 may be generated such that the data with a higher priority has a stronger influence on the reflection motion 2010. The priority may be automatically set on the system side in advance, or may be appropriately set or changed by the user. Note that, depending on the network situation, the resources of each processing device (such as the server 10 and the user terminal 20), the remaining battery level, etc., which data among the data included in the user data 2001 is used to generate the reflection motion 2010 may be dynamically changed. For example, when the network situation is good, the reflection motion 2010 may be generated based on the data with priorities of 5 to 10 among the data included in the user data 2001, and when the network situation is not good, the reflection motion 2010 may be generated based only on the data with priorities of 9 to 10.

[0120] The generation of the reflection motion 2010 based on various data included in the user data 2001 by the reflection motion generation model M2 has been described above. However, the reflection motion generation model M2 is not limited to the examples described so far as long as it can output a reflection motion 2010 appropriate to the situation of the event under the influence of the speaker avatar P and the manipulated spectator avatar O in the event, and it may be learned in any way. Note that the reflection motion 2010 generated by the motion generation unit 121 may be used for the learning of the machine learning model.

[0121] As described above, by using a machine learning model for the generation of the reflection motion 2010, the provider of the virtual space can generate a motion appropriate to the situation of the event without performing a complicated design. Note that a machine learning model may not be used for the generation of the reflection motion 2010. For example, the motion generation unit 121 may generate the reflection motion 2010 based on predetermined rules for the user data 2001 and the reflection motion 2010.

[0122] The motion generation unit 121 may generate, for each of the plurality of generated spectator avatars G, a reflection motion 2010 that reflects the preliminary motion 2011 by inputting the preliminary motion 2011 into the reflection motion generation model M2. In this case, the reflection motion generation model M2 may output a plurality of reflection motions 2010 corresponding to the respective generated avatar coordinates 1001, using the plurality of generated avatar coordinates 1001 as input data. Further, the reflection motion generation model M2 may output, for each of the plurality of generated avatar coordinates 1001, a reflection motion 2010 corresponding thereto, using each of the plurality of generated avatar coordinates 1001 as input data.

[0123] Also, although an example of generating the reflected motion 2010 using two machine learning models has been described here with reference to FIG. 6, only one machine learning model may be used in generating the reflected motion 2010. More specifically, only the reflected motion generation model M2 may be used to generate the reflected motion 2010. Here, the reflected motion generation model M2 may generate the reflected motion 2010 by using the sound source B as input data instead of the pre-seed motion 2011.

[0124] (Acoustic generation unit 122) The acoustic generation unit 122 generates the area sound 2020 based on each user data 2001 received by the communication unit 110. The acoustic generation unit 122 generates the area sound 2020 to be a sound corresponding to the situation of the event.

[0125] The acoustic generation unit 122 generates the area sound 2020 as the sound reflected in the spectator area A where the generated spectator avatar G exists in the virtual space. More specifically, the acoustic generation unit 122 generates the area sound 2020 as the sound emitted by the generated spectator avatar G existing in the spectator area A. For example, the acoustic generation unit 122 may generate sounds that spectators may make during the event. Sounds that spectators may make during the event include, for example, cheering, the sound of clapping hands, the sound of applause, and the sounds emitted by props such as musical instruments held by spectators.

[0126] The acoustic generation unit 122 may generate, for example, the area sound 2020 to be reflected in a predetermined number of frames. The predetermined number is not particularly limited, and may be set, for example, so that the area sound 2020 played in 1 second to several seconds is generated. The acoustic generation unit 122 repeatedly executes the generation process of the area sound 2020. For example, by repeatedly performing the generation process during the event, the area sound 2020 may be continuously reflected in the spectator area A during the event.

[0127] The sound generation unit 122 may generate the area sound 2020 using a machine learning model generated using machine learning techniques such as DNN. FIG. 7 is a diagram for explaining an example in which the sound generation unit 122 generates the area sound 2020 using a machine learning model.

[0128] As shown in FIG. 7, the area sound 2020 is generated using the sound generation model M3. The sound generation model M3 is a machine learning model that outputs the area sound 2020 using the user data 2001 and the data indicating the location of the generated avatar area Ar as input data. The sound generation model M3 outputs the area sound 2020 using the user data 2001 of a plurality of users.

[0129] The sound generation unit 122 inputs the user data 2001 obtained from a predetermined time ago to the present to the sound generation model M3. The predetermined time may be, for example, 1 second to several seconds.

[0130] The sound generation unit 122 may input the user data 2001 of all users participating in the event to the sound generation model M3.

[0131] Further, the sound generation unit 122 may sample some of the user data 2001 from the user data 2001 of all users participating in the event and input it to the sound generation model M3. The sampling method of the user data 2001 is the same as the method described in the motion generation unit 121.

[0132] The generated avatar area Ar is the spectator area A in the virtual space where the generated avatar exists, that is, the spectator area A where the generated area sound 2020 is reflected. The sound generation model M3 generates the area sound 2020 based on the relationship between the user avatar coordinates of each user and the location of the generated avatar area Ar.

[0133] For example, the acoustic generation model M3 may generate the area sound 2020 such that the closer the user avatar coordinates are to the location of the generated avatar area Ar, the greater the influence of the user data 2001 corresponding to the user avatar coordinates on the generation of the area sound 2020. More specifically, the acoustic generation model M3 may generate the area sound 2020 such that the influence of the spectator data 2001B of the spectator avatar coordinates included in the generated avatar area Ar on the generation of the area sound 2020 is greater than that of other spectator data 2001B.

[0134] Further, the acoustic generation model M3 may generate the area sound 2020 without using the spectator data 2001B for the manipulated spectator avatar O existing in the spectator area A other than the generated avatar area Ar. Also, the acoustic generation model M3 may generate the area sound 2020 by weighting the spectator data 2001B based on the distance between the location of the generated avatar area Ar and the spectator avatar coordinates.

[0135] It is assumed that the spectator is more influenced by the performers and spectators existing closer to the spectator themselves than by those existing far away and makes a sound. Therefore, by generating the area sound 2020 based on the location of the generated avatar area Ar, the sound corresponding to the situation of the event can be superimposed on the sound in the virtual space.

[0136] In the above, an example of generating the area sound 2020 by weighting the spectator data 2001B based on the distance between the location of the generated avatar area Ar and the spectator avatar coordinates has been described, but the example of weighting is not limited to this. For example, the acoustic generation model M3 may generate the area sound 2020 by weighting each user data 2001 according to the feature information of the generated spectator avatar G existing in the preset generated avatar area Ar.

[0137] As a more specific example, the acoustic generation model M3 may weight each user data 2001 according to the user attribute information of a user who has a friendship relationship with the generated audience avatar G existing in the generation avatar area Ar. For example, the acoustic generation model M3 may weight the performer data 2001A of a performer user that matches the favorite artist of a user who has a friendship relationship with the generated audience avatar G so that the influence on the generation of the area acoustic 2020 becomes greater. Also, the acoustic generation model M3 may weight the user data 2001 of a user who has a friendship relationship with the generated audience avatar G and the users existing around the user so that the influence on the generation of the area acoustic 2020 becomes greater.

[0138] As another example, the acoustic generation model M3 may calculate the degree of match between the favorite artist of the generated audience avatar G and each performer user, and weight the performer data 2001A according to the degree of match. Then, the acoustic generation model M3 may generate the area acoustic 2020 so that the influence on the area acoustic 2020 becomes greater for the performer data 2001A of a performer user with a higher degree of match.

[0139] Subsequently, the generation of the area acoustic 2020 based on various data included in the user data 2001 by the acoustic generation model M3 will be described. The acoustic generation model M3 generates the area acoustic 2020 based on the performer data 2001A included in the user data 2001.

[0140] For example, the acoustic generation model M3 generates the area acoustics 2020 based on the performer voice data 2003A among the performer data 2001A. The acoustic generation model M3 may generate the area acoustics 2020 such that the volume of the area acoustics 2020 increases as the volume of the performer voice represented by the performer voice data 2003A increases. Further, when the performer voice data 2003A includes a sound of a specific word uttered by the performer user, the acoustic generation model M3 may generate the area acoustics 2020 such that the area acoustics 2020 includes the specific sound. For example, when the performer voice data 2003A includes a voice uttering the word "clap" by the performer user, the acoustic generation model M3 may generate the area acoustics 2020 to include the sound of the clap.

[0141] In an event, it is assumed that the audience makes a sound under the influence of the sound made by the performer. Therefore, by generating the area acoustics 2020 based on the performer voice data 2003A, acoustics corresponding to the situation of the event can be superimposed on the acoustics of the virtual space.

[0142] In addition, the acoustic generation model M3 generates area acoustics 2020 based on the performer motion data 2002A. For example, the acoustic generation model M3 may generate the area acoustics 2020 such that the volume of the area acoustics 2020 increases as the motion represented by the performer motion data 2002A becomes larger. Here, the volume of the area acoustics 2020 may be changed linearly or non-linearly according to the magnitude of the motion represented by the performer motion data 2002A. Further, when the motion represented by the performer motion data 2002A is a specific motion (such as a motion of clapping hands or a motion of waving hands at a certain rhythm), the acoustic generation model M3 may generate the area acoustics 2020 to include a specific sound. For example, when the motion represented by the performer motion data 2002A is a motion of clapping hands, the acoustic generation model M3 may generate the area acoustics 2020 to include the sound of clapping hands. Also, in games or the like that utilize a virtual space such as a metaverse, an "emote" for causing an avatar to perform a preset motion may be set. The "emote" may be included in the user data 2001. The acoustic generation model M3 may generate the area acoustics 2020 using the "emote" preset by the performer user included in the performer data 2001A. The acoustic generation model M3 may generate the area acoustics 2020 based on the motion represented by the emote in the same manner as when using the performer motion data 2002A.

[0143] In the event, it is assumed that the audience is affected by the motion of the performer and makes a sound. Therefore, by generating the area acoustics 2020 based on the performer motion data 2002A, acoustics corresponding to the situation of the event can be superimposed on the acoustics of the virtual space.

[0144] In addition, the acoustic generation model M3 generates the area acoustics 2020 based on the audience data 2001B among the user data 2001.

[0145] For example, the acoustic generation model M3 generates the area acoustics 2020 based on the spectator voice data 2003B among the spectator data 2001B. For example, the acoustic generation model M3 may generate the area acoustics 2020 such that the volume of the area acoustics 2020 increases as the volume of the spectator voice represented by the spectator voice data 2003B increases. Further, the acoustic generation model M3 may generate the area acoustics 2020 so as to resemble the spectator voice represented by the spectator voice data 2003B and the area acoustics 2020.

[0146] In an event, it is assumed that a spectator emits a sound under the influence of the sounds emitted by other spectators. Therefore, by generating the area acoustics 2020 based on the spectator motion data 2002B, acoustics corresponding to the situation of the event can be superimposed on the acoustics of the virtual space.

[0147] Also, the acoustic generation model M3 generates area acoustics 2020 based on the audience motion data 2002B among the audience data 2001B. For example, the acoustic generation model M3 may generate the area acoustics 2020 such that the louder the motion represented by the audience motion data 2002B is, the louder the volume of the area acoustics 2020 becomes. Here, the volume of the area acoustics 2020 may be linearly or non-linearly changed according to the magnitude of the motion represented by the audience motion data 2002B. Also, when the motion represented by the audience motion data 2002B is a specific motion (such as a motion of clapping hands, a motion of waving hands at a certain rhythm, etc.), the acoustic generation model M3 may generate the area acoustics 2020 to include a specific sound. For example, when the motion represented by the audience motion data 2002B is a motion of clapping hands, the acoustic generation model M3 may generate the area acoustics 2020 to include the sound of clapping hands. Also, in games or the like that utilize a virtual space such as a metaverse, there may be set an "emote" in which an avatar performs a preset motion. The acoustic generation model M3 may generate the area acoustics 2020 using the "emote" preset by the audience user included in the audience data 2001B. The acoustic generation model M3 may generate the area acoustics 2020 based on the motion represented by the emote in the same manner as when using the audience motion data 2002B.

[0148] In an event, it is assumed that the audience is affected by the motions of other audiences and makes sounds. Therefore, by generating the area acoustics 2020 based on the audience motion data 2002B, acoustics corresponding to the situation of the event can be superimposed on the acoustics of the virtual space.

[0149] Note that although the example in which the user data 2001 is input to the acoustic generation model M3 has been described so far, the virtual space information including the video and acoustics of the virtual space generated based on the user data 2001 may be input to the acoustic generation model M3 together with the user data 2001 or instead of the user data 2001.

[0150] The audience is assumed to emit sounds based on the motions of the audience and performers that they see with their own eyes and the voices that they hear with their own ears. Therefore, by generating the area sound 2020 based on the video and sound in the virtual space generated on the user terminal 20 of the user, particularly the audience user, it is possible to superimpose a sound appropriate for the event situation on the sound in the virtual space.

[0151] Further, the sound generation model M3 may generate the area sound 2020 based on the user attribute information 131 as well. The sounds emitted by the audience may tend to vary depending on factors such as age, gender, or preferences regarding the performance. Therefore, by generating the area sound 2020 based on the user attribute information 131 of each user, it is possible to superimpose a sound that matches the manipulated audience avatar O participating in the event on the sound in the virtual space.

[0152] Further, the sound generation model M3 may generate the reflection motion 2010 according to the feature information of the generated audience avatar G. For example, the area sound 2020 may be generated such that the motion of a user in a friendship relationship with the generated audience avatar G is more strongly influenced than that of a user not in a friendship relationship.

[0153] As another example, when the distance between the generated avatar coordinates 1001 and the performer avatar coordinates of the performer user that matches the favorite artist of the generated audience avatar G is shorter than a predetermined distance, the sound generation model M3 may generate the area sound 2020 such that the volume becomes louder.

[0154] In this way, by generating the area sound 2020 according to the feature information of the generated audience avatar G by the sound generation model M3, it is possible to give personality to the generated audience avatar G, and thus it is possible to reproduce a more realistic sound of the audience.

[0155] Note that although it has been described so far that various data included in the user data 2001 can be used for the sound generation model M3 to generate the area sound 2020, the data used to generate the area sound 2020 is not particularly limited.

[0156] For example, the acoustic generation model M3 may generate the area acoustics 2020 using only a part of the data among the performer data 2001A and the audience data 2001B. In particular, the performer data 2001A may be given priority over the audience data 2001B and used for generating the area acoustics 2020. This is because it is assumed that the audience participates in the event with the greatest interest in the performer's motion and voice.

[0157] Also, the performer voice data 2003A and the audience voice data 2003B may be given priority over the performer motion data 2002A and the audience motion data 2002B and used for generating the area acoustics 2020. This is because it is assumed that the sounds emitted by the audience are most affected by the sounds emitted by the performer and other audiences.

[0158] In addition, when generating the area acoustics 2020 using a plurality of the data included in the user data 2001, the area acoustics 2020 may be generated by weighting each data. For example, each data may be weighted so that the performer voice data 2003A is most reflected in the area acoustics 2020. For example, priorities may be set for each of the plurality of data included in the user data 2001. The priority may be set as an integer value from 0 to 10, for example, with 0 being the minimum and 10 being the maximum. The area acoustics 2020 may be generated and reflected so that the data with a higher priority has a stronger influence on the area acoustics 2020. The priority may be automatically set on the system side in advance, or may be appropriately set or changed by the user. Note that, depending on the network situation, the resources of each processing device (such as the server 10 and the user terminal 20), the remaining battery level, etc., which data among the data included in the user data 2001 is used to generate the area acoustics 2020 may be dynamically changed. For example, when the network situation is good, the area acoustics 2020 is generated based on the data with priorities of 5 to 10 among the data included in the user data 2001, but when the network situation is not good, the area acoustics 2020 may be generated based only on the data with priorities of 9 to 10.

[0159] The generation of the area sound 2020 based on various data included in the user data 2001 and the location of the generated avatar area Ar by the above sound generation model M3 has been described. The sound generation model M3 may further use the motion data of the reflection motion 2010 generated by the motion generation unit 121 as input data to generate the area sound 2020.

[0160] The sound generation model M3 may generate the area sound 2020 so as to include a sound appropriate as the sound emitted by the generated spectator avatar G in which the reflection motion 2010 is reflected. For example, when the reflection motion 2010 is a motion of clapping hands, the sound generation model M3 may generate the area sound 2020 including the sound of clapping hands.

[0161] The area sound 2020 generated using the reflection motion 2010 by the sound generation model M3 is reflected into the virtual space simultaneously when the reflection motion 2010 is reflected in the generated spectator avatar G. Thereby, since the area sound 2020 appropriate as the sound emitted by the generated spectator avatar G can be generated, a virtual space where a large number of spectators get excited can be produced more realistically.

[0162] Whether to use the reflection motion 2010 or not in the generation of the area sound 2020 can be set as appropriate.

[0163] Here, the difference in time taken from the collection of the user data 2001 to the reflection into the virtual space when the sound generation model M3 uses the reflection motion 2010 and when it does not use it in the generation of the area sound 2020 will be described with reference to FIGS. 8 and 9.

[0164] FIG. 8 is a diagram for explaining the time taken from the collection of user data 2001 to the reflection in the virtual space when the reflection motion 2010 is used for the generation of the area sound 2020. FIG. 9 is a diagram for explaining the time taken from the collection of user data 2001 to the reflection in the virtual space when the reflection motion 2010 is not used for the generation of the area sound 2020.

[0165] In FIGS. 8 and 9, the processing from the collection of user data 2001 from the user terminal 20 to the reflection of the reflection motion 2010 and the area sound 2020 generated based on the user data 2001 in the virtual space is shown in chronological order.

[0166] When the reflection motion 2010 is used for the generation of the area sound 2020, as shown in FIG. 8, first, the user data 2001 is collected by time T0 (user data collection 1a). Next, by time T1, the reflection motion 2010 is generated based on the user data 2001 (reflection motion generation 1a).

[0167] Subsequently, by time T2, the area sound 2020 is generated based on the user data 2001 and the reflection motion 2010 (area sound generation 1a). And by time T3, the generated reflection motion 2010 and area sound 2020 are reflected in the virtual space (reflection 1a).

[0168] In parallel with the above processing, new user data 2001 is collected by time T1 (user data collection 2a), the reflection motion 2010 is generated using the user data 2001 (reflection motion generation 2a), the area sound 2020 is generated based on the user data 2001 and the reflection motion 2010 (area sound generation 2a), and the reflection motion 2010 and area sound 2020 are reflected in the virtual space (reflection 2a). Thus, the reflection of the reflection motion 2010 and the area sound 2020 in the virtual space from the collection of the user data 2001 is continuously performed.

[0169] On the one hand, when the reflection motion 2010 is not used for generating the area sound 2020, as shown in FIG. 9, first, user data 2001 is collected by time T0 in the same manner as the example shown in FIG. 8 (user data collection 1b).

[0170] Next, by time T1, the reflection motion 2010 is generated based on the user data 2001 (reflection motion generation 1b). Also, in parallel with the process of reflection motion generation 1b, by time T2, the area sound 2020 is generated based on the user data 2001 (area sound generation 1b). That is, in the example shown in FIG. 9, the generation of the reflection motion 2010 and the area sound 2020 is performed simultaneously. Hereinafter, the simultaneous generation of the reflection motion 2010 and the area sound 2020 is also referred to as "simultaneous generation". In this specification, "simultaneous" includes substantially simultaneous. Also, "simultaneous generation" may mean that at least a part of each generation timing overlaps.

[0171] Then, by time T2, the generated reflection motion 2010 and area sound 2020 are reflected in the virtual space (reflection 1b). In parallel with the above process, new user data 2001 is collected by time T1, and the reflection motion 2010 and area sound 2020 generated using the user data 2001 are reflected in the virtual space (user data collection 2b to reflection 2b). In this way, the reflection motion 2010 and the area sound 2020 are continuously reflected in the virtual space.

[0172] As described so far, in the example shown in FIG. 9, the reflected motion 2010 and the area sound 2020 generated based on the user data 2001 collected by time T0 are reflected by time T2. On the other hand, in the example shown in FIG. 8, the reflected motion 2010 and the area sound 2020 generated based on the user data 2001 collected by the same time T0 are reflected by time T3. Thus, when generating the area sound 2020 without using the reflected motion 2010, the time from the collection of the user data 2001 to the reflection in the virtual space is shortened. Also, when generating the area sound 2020 without using the reflected motion 2010, the computational amount by the server 10 also decreases.

[0173] Whether to use the reflected motion 2010 in generating the area sound 2020 may be set by the provider of the virtual space or may be determined based on the processing load of the server 10. The processing load of the server 10 may be represented by, for example, the amount of memory used or the amount of power consumed by the GPU or CPU that the server 10 has. For example, when the processing load of the server 10 is greater than a threshold, the acoustic generation unit 122 may generate the area sound 2020 without using the reflected motion 2010.

[0174] As described above, the generation of the area sound 2020 using the acoustic generation model M3 by the acoustic generation unit 122 has been described. However, the acoustic generation model M3 is not limited to the examples described so far as long as it can output an area sound 2020 appropriate for the situation of the event under the influence of the performer avatar P and the manipulated spectator avatar O in the event, and it may be learned in any way. Note that the area sound 2020 generated by the acoustic generation unit 122 may be used for the learning of the machine learning model.

[0175] Note that the sound generation unit 122 may generate area sounds 2020 corresponding to a plurality of generation avatar areas Ar. In this case, the sound generation model M3 may take data indicating the locations of the plurality of generation avatar areas Ar as input data and output a plurality of area sounds 2020 corresponding to the respective generation avatar areas Ar. Alternatively, the sound generation model M3 may take each of the data indicating the locations of the plurality of generation avatar areas Ar as input data and output the corresponding area sound 2020.

[0176] In addition, although an example in which the area sound 2020 is generated using a machine learning model has been described here with reference to FIG. 7, a machine learning model may not be used for generating the area sound 2020. For example, the sound generation unit 122 may generate the area sound 2020 based on a predefined rule regarding the user data 2001 and the area sound 2020. However, by using a machine learning model for generating the area sound 2020, the provider of the virtual space can generate appropriate sounds according to the event situation without performing complicated design.

[0177] (Virtual space generation unit 123) The virtual space generation unit 123 generates a virtual space. Specifically, as the generation of the virtual space, the virtual space generation unit 123 performs generation, acquisition, addition, update, etc. of various data included in the content data. The various data are, for example, data for drawing an image of the virtual space and data for playing back the sound of the virtual space.

[0178] The data for drawing an image of the virtual space includes data regarding the user avatars of each user arranged in the virtual space and data regarding the generated spectator avatar G. The data regarding the user avatar is, for example, motion data of the user reflected in the user avatar, appearance data of the user avatar, user avatar coordinates, and the like.

[0179] Data related to the generated spectator avatar G includes, for example, data on the reflected motion 2010 reflected in the generated spectator avatar G, appearance data of the generated spectator avatar G, generated avatar coordinates 1001, and the like.

[0180] The appearance data of the generated spectator avatar G may be generated by the provider of the virtual space by presetting facial parts, clothing to be worn, etc. in advance, or may be generated to match the user avatar of the user participating in the event based on the user data 2001.

[0181] The generated avatar coordinates 1001 may be set in advance by the provider of the virtual space at the start of the event. Then, the generated avatar coordinates 1001 may move according to the continuously generated reflected motion 2010 during the event. Also, the generated avatar coordinates 1001 may be fixed to the coordinates preset by the provider of the virtual space throughout the event. Further, the virtual space generation unit 123 may extract a location with few manipulated spectator avatars O and set it as the generated avatar coordinates 1001.

[0182] Note that the virtual space generation unit 123 may set the number of generated spectator avatars G to be arranged in the virtual space. The number of generated spectator avatars G to be arranged in the virtual space may be appropriately adjusted during the event, for example, according to the processing load of the server 10.

[0183] Data for reproducing the acoustics of the virtual space includes, for example, performer voice data 2003A, spectator voice data 2003B, and sound data of the area acoustics 2020.

[0184] Note that the virtual space generation unit 123 may generate the video of the virtual space from each user's perspective and generate the acoustics of the virtual space, and transmit them to the user terminal 20 as content data, and perform control to play them on the user terminal 20.

[0185] The virtual space generation unit 123 controls the communication unit 110 to appropriately transmit the content data to the user terminal 20. For example, the virtual space generation unit 123 may control the communication unit 110 to transmit the content data of the virtual space for each frame to the user terminal 20 as update information of the virtual space.

[0186] (Memory unit 130) The memory unit 130 is realized by a ROM that stores programs, arithmetic parameters, etc. used in the processing of the control unit 120, and a RAM that temporarily stores parameters that change as appropriate. For example, the memory unit 130 stores the user attribute information 131 of each user obtained from each user terminal 20.

[0187] <3. Example of operation processing> Subsequently, an example of the operation processing flow of the server 10 according to the present embodiment will be described with reference to FIG. 10. FIG. 10 is a flowchart showing an example of the operation processing flow of the server 10 according to the present embodiment.

[0188] First, the communication unit 110 acquires the user data 2001 of each user from each user terminal 20 (S101). Subsequently, the control unit 120 determines whether the setting for simultaneous generation of the reflection motion 2010 and the area sound 2020 is ON (S102).

[0189] If the setting for simultaneous generation is ON (S102 / YES), the process proceeds to S108. If the setting for simultaneous generation is OFF (S102 / NO), the control unit 120 determines whether the processing load of the server 10 is equal to or greater than a first threshold value (S103). The first threshold value is a threshold value set to determine whether to execute the simultaneous generation of the reflection motion 2010 and the area sound 2020.

[0190] When the processing load of server 10 is less than the first threshold value (S103 / NO), the control unit 120 does not execute simultaneous generation, and sequentially generates the reflection motion 2010 and the area sound 2020. That is, first, the motion generation unit 121 generates the reflection motion 2010 to be reflected on each generated spectator avatar G (S104). Next, the sound generation unit 122 generates the area sound 2020 using the reflection motion 2010 generated by the motion generation unit 121 (S105).

[0191] On the other hand, when the processing load of server 10 is equal to or greater than the first threshold value (S103 / YES), the control unit 120 further determines whether the processing load of server 10 is equal to or greater than the second threshold value (S106). The second threshold value is a threshold value set to determine whether to sample the spectator data 2001B among the user data 2001 of the users participating in the event when generating the reflection motion 2010 and the area sound 2020. The second threshold value is a threshold value different from the first threshold value, and a value higher than the first threshold value is set.

[0192] When the processing load of server 10 is equal to or greater than the second threshold value (S106 / YES), the control unit 120 samples some of the spectator data 2001B to be used for generating the reflection motion 2010 and the area sound 2020 from the plurality of spectator data 2001B (S107). Subsequently, based on the sampled spectator data 2001B, the motion generation unit 121 and the sound generation unit 122 simultaneously generate the reflection motion 2010 and the area sound 2020 (that is, each executes the generation process of each data in parallel) (S108). By executing the generation process of each data using the spectator data 2001B sampled here, the processing load of server 10 is reduced.

[0193] On the other hand, when the processing load of the server 10 is less than the second threshold value (S106 / NO), each of the reflection motion 2010 and the area sound 2020 is simultaneously generated by the motion generation unit 121 and the sound generation unit 122 (S108). At this time, the motion generation unit 121 and the sound generation unit 122 may generate the reflection motion 2010 and the area sound 2020 respectively by inputting all the viewer data 2001B into the machine learning model.

[0194] The virtual space generation unit 123 generates a virtual space by generating content data including the generated data of the reflection motion 2010 and the area sound 2020. The communication unit 110 outputs the generated content data to the user terminal 20 (S109).

[0195] Then, the virtual space generation unit 123 determines whether to end the generation of the virtual space (S110). The generation of the virtual space may be ended based on an end operation by the provider of the virtual space. When the virtual space generation unit 123 does not end the generation of the virtual space (S110 / NO), the process returns to S101, and the processes of S101 to S109 are repeated, so that the virtual space is continuously generated. On the other hand, when the generation of the virtual space is ended (S110 / YES), the process ends.

[0196] <4. Modification Example> Subsequently, a modification example according to the present embodiment will be described.

[0197] (First Modification Example) First, as a first modification example, a description will be given of how each generated viewer avatar G is drawn so that the load of the drawing process for some of the generated viewer avatars G is reduced.

[0198] In the first modification example, the virtual space generation unit 123 generates a virtual space such that the rendering load for some of the generated spectator avatars G is reduced. Here, an example of the virtual space generated by the virtual space generation unit 123 in the first modification example will be described with reference to FIG. 11. FIG. 11 is a diagram for explaining an example of the virtual space generated by the virtual space generation unit 123 in the first modification example. In FIG. 11, a video 215 of the virtual space is shown. The video 215 of the virtual space is a video of the virtual space rendered by the server 10 or the user terminal 20 based on the content data of the virtual space generated by the virtual space generation unit 123.

[0199] As shown in the video 215 of the virtual space, the virtual concert hall V may have different brightness levels depending on the location. For example, the spectator area A1 is brighter compared to the other spectator areas A2 to A4. The generated spectator avatar G rendered in the spectator area A with a darker brightness is less visible compared to the generated spectator avatar G rendered in the spectator area A with a brighter brightness.

[0200] Therefore, the virtual space generation unit 123 determines the brightness level for each spectator area A on the virtual space. Then, for the generated spectator avatar G existing in the spectator area A where the brightness is lower than the threshold value, the virtual space generation unit 123 generates a virtual space such that the rendering load is reduced.

[0201] The virtual space generation unit 123 may generate content data such that, for example, the appearance of the generated spectator avatar G existing in the spectator area A where the brightness is lower than the threshold value is replaced with a simple appearance and then rendered. For example, the virtual space generation unit 123 may generate content data such that the appearance of the generated spectator avatar G existing in the spectator area A is replaced with a silhouette in which the entire generated spectator avatar G is filled with a single color such as black or gray and then rendered.

[0202] Further, the virtual space generation unit 123 may generate content data such that the reflection motion 2010 generated for the generated spectator avatar G existing in the spectator area A where the brightness is lower than the threshold is not reflected on the generated spectator avatar G. For example, the virtual space generation unit 123 may generate content data such that the generated spectator avatar G existing in the spectator area A stands upright.

[0203] In the example shown in FIG. 11, assume that the virtual space generation unit 123 determines that the brightness of the spectator area A1 is equal to or higher than the threshold, and determines that the brightnesses of the spectator areas A2 to A4 are lower than the threshold. In this case, for example, the appearance and the reflection motion 2010 are reflected on the generated spectator avatars G1 and G2 existing in the spectator area A1. Hereinafter, the generation in which the appearance and the reflection motion 2010 are reflected on the generated spectator avatar G is also referred to as "normal generation".

[0204] On the other hand, the generated spectator avatar G3 etc. existing in the spectator area A2 are generated by replacing the appearance with a silhouette and changing so that the reflection motion 2010 is not reflected. Hereinafter, the generation in which the generated spectator avatar G is changed to reduce the load of the drawing process is also referred to as "changed generation". By performing the changed generation in this way, it is possible to reduce the load of the drawing process of the virtual space without giving a sense of incongruity to the user who views the video of the virtual space.

[0205] Whether or not to perform the changed generation according to the brightness of the virtual space may be determined based on whether or not it is set in advance by the provider of the virtual space to perform the changed generation.

[0206] The brightness for each spectator area A may change depending on the change in the lighting environment of the virtual concert hall V. Therefore, the virtual space generation unit 123 may continuously (for example, for each frame) determine the brightness for each spectator area A on the virtual space.

[0207] (Second Modification Example) From here, as a second modification example, the adjustment and output of the volume of each area sound 2020 will be described.

[0208] The virtual space generation unit 123 according to the second modification example determines the volume of the area sound 2020 for each spectator area A according to the excitement level of the event. The excitement level of the event may be detected based on, for example, the magnitude of the motion indicated by the spectator motion data 2002B or the volume of the sound indicated by the spectator voice data 2003B. The excitement level of the event may be detected, for example, in two levels of "low" and "high". Also, the excitement level of the event may be represented by a numerical value.

[0209] The virtual space generation unit 123 may generate content data such that, for example, the higher the excitement level, the greater the volume of the area sound 2020 for each spectator area A.

[0210] Also, when the excitement level is low, the virtual space generation unit 123 may generate content data such that the area sound 2020 is not reflected in some of the spectator areas A.

[0211] FIG. 12 is a diagram for explaining an example of the reflection of the area sound 2020 for each of a plurality of spectator areas A when the excitement level is "low" and when the excitement level is "high". As shown on the left side of FIG. 12, when the excitement level is "low", content data may be generated such that the area sound 2020 is reflected only in the spectator areas A5, A8, and A9 among the plurality of spectator areas A5 to A10. Here, as an example, it has been described that the area sound 2020 is reflected only in the spectator areas A5, A8, and A9, but any part of the spectator areas A may be sufficient, and the spectator areas to be reflected are not limited to the example shown on the left side of FIG. 12. For example, the virtual space generation unit 123 may select non-adjacent spectator areas A as some of the spectator areas A, or may select only the central spectator area A, or may select only the spectator area A near the stage, or may select only the rear spectator area A. Also, the spectator area A designated by the user may be selected.

[0212] Also, as shown on the right side of FIG. 12, when the degree of excitement is "high", the content data may be generated so that the area sound 2020 is reflected in all of the plurality of spectator areas A5 to A10. Here, as an example, it is assumed that the area sound 2020 is reflected in all spectator areas A, but it is not limited to this. The area sound 2020 may be reflected in a majority of spectator areas A according to the degree of excitement, or may be reflected in 80% of the spectator areas A.

[0213] In addition, when the degree of excitement is "high", an area sound 2020 with a louder volume than when the degree of excitement is "low" may be reflected. For example, in the example shown in FIG. 12, in the spectator areas A5, A8, and A9 where the area sound 2020 is commonly reflected in both the case of "low" and "high" degrees of excitement, an area sound 2020 with a louder volume than when the degree of excitement is "low" may be reflected when the degree of excitement is "high".

[0214] By adjusting the volume of each spectator area A, the area sound 2020 with a louder volume is superimposed as the degree of excitement increases. Therefore, it can be expected to create a further increase when the event is lively. On the other hand, when the degree of excitement is low, it is assumed that the spectator user wants to quietly listen to the performance by the performer user. According to the above configuration, since the area sound 2020 with a small volume is superimposed when the degree of excitement is low, it is possible to produce a virtual space that meets such needs.

[0215] Whether or not to adjust the volume of each area sound 2020 may be determined based on whether it is set in advance by the provider of the virtual space to adjust the volume.

[0216] (Operation processing examples for the first modification example and the second modification example) Next, an example of the operation process flow of the server 10 when the first modification example and the second modification example are applied will be described with reference to FIG. 13. FIG. 13 is a flowchart showing an example of the operation process flow of the server 10 when the first modification example and the second modification example are applied.

[0217] The process shown in FIG. 13 is executed following the process of S105 or S108 shown in FIG. 10.

[0218] First, the virtual space generation unit 123 determines whether an adjustment setting indicating that the generation of the change of the generated spectator avatar G set by the provider of the virtual space is to be performed is ON (S201).

[0219] When the adjustment setting of the generated spectator avatar G is OFF (S201 / NO), the virtual space generation unit 123 determines to normally generate all the generated spectator avatars G (S202). That is, the virtual space generation unit 123 generates content data so that the appearance and the reflected motion 2010 are reflected in all the generated spectator avatars G.

[0220] When the adjustment setting of the generated spectator avatar G is ON (S201 / YES), the virtual space generation unit 123 determines the brightness of each spectator area A in the virtual space (S203). Then, the virtual space generation unit 123 determines the generated spectator avatar G to be changed and generated (S204). Specifically, the virtual space generation unit 123 may determine to change and generate the generated spectator avatar G existing in the spectator area A whose brightness is smaller than the threshold value.

[0221] Subsequently, the virtual space generation unit 123 determines whether an adjustment setting indicating that the volume of the area sound 2020 is to be adjusted, set by the provider of the virtual space, is ON (S205).

[0222] When the adjustment setting of the volume of the area sound 2020 is OFF (S205 / NO), the virtual space generation unit 123 determines to superimpose the area sounds 2020 of all the spectator areas A without adjusting the volume (S206).

[0223] When the adjustment setting of the volume of the area sound 2020 is ON (S205 / YES), the virtual space generation unit 123 determines the degree of excitement of the virtual space (S207). Then, the virtual space generation unit 123 determines the volume of the area sound 2020 for each spectator area A (S208). For example, the virtual space generation unit 123 may determine the volume so that the higher the degree of excitement of the virtual space, the higher the volume of each area sound 2020.

[0224] Subsequently, the process proceeds to S109 shown in FIG. 10. In S109, the appearance of the generated spectator avatar G determined to be changed and generated in S204 may be replaced with a silhouette, and the virtual space may be generated without reflecting the reflection motion 2010 on the generated spectator avatar G. Further, in S109, the virtual space is generated by reflecting the volume for each spectator area A determined in S209.

[0225] <5. Hardware Configuration> The embodiments and modified examples according to the present disclosure have been described above. Next, with reference to FIG. 14, a hardware configuration example of the server 10 and the user terminal 20 according to the embodiment of the present disclosure will be described.

[0226] The processing by the server 10 and the user terminal 20 described above can be realized by one or more information processing apparatuses. FIG. 14 is a block diagram showing a hardware configuration example of an information processing apparatus 900 that realizes the server 10 and the user terminal 20 according to the embodiment of the present disclosure. Note that the information processing apparatus 900 does not necessarily have all of the hardware configurations shown in FIG. 14. Also, a part of the hardware configuration shown in FIG. 14 may not exist in the server 10 or the user terminal 20.

[0227] As shown in FIG. 14, the information processing apparatus 900 includes a CPU 901, a ROM (Read Only Memory) 903, and a RAM 905. Further, the information processing apparatus 900 may include a host bus 907, a bridge 909, an external bus 911, an interface 913, an input device 915, an output device 917, a storage device 919, a drive 921, a connection port 923, and a communication device 925. The information processing apparatus 900 may have a processing circuit such as a GPU (Graphics Processing Unit), a DSP (Digital Signal Processor), or an ASIC (Application Specific Integrated Circuit) instead of or together with the CPU 901.

[0228] The CPU 901 functions as an arithmetic processing unit and a control unit, and controls all or part of the operations within the information processing apparatus 900 according to various programs recorded in the ROM 903, the RAM 905, the storage device 919, or the removable recording medium 927. The ROM 903 stores programs, arithmetic parameters, etc. used by the CPU 901. The RAM 905 temporarily stores programs used in the execution of the CPU 901 and parameters that change as appropriate during the execution. The CPU 901, the ROM 903, and the RAM 905 are interconnected by a host bus 907 constituted by an internal bus such as a CPU bus. Further, the host bus 907 is connected to an external bus 911 such as a PCI (Peripheral Component Interconnect / Interface) bus via the bridge 909.

[0229] The input device 915 is a device operated by a user, such as a button. The input device 915 may include a mouse, a keyboard, a touch panel, switches, levers, etc. Further, the input device 915 may include a microphone that detects the user's voice. The input device 915 may be, for example, a remote control device using infrared rays or other radio waves, or an external connection device 929 such as a mobile phone corresponding to the operation of the information processing device 900. The input device 915 includes an input control circuit that generates an input signal based on the information input by the user and outputs it to the CPU 901. By operating this input device 915, the user inputs various data to the information processing device 900 or instructs processing operations.

[0230] Further, the input device 915 may include an imaging device and a sensor. The imaging device is a device that images real space using various members such as an imaging element such as a CCD (Charge Coupled Device) or a CMOS (Complementary Metal Oxide Semiconductor), and a lens for controlling the formation of a subject image on the imaging element, and generates an imaging image. The imaging device may capture still images or may capture moving images.

[0231] The sensor is various sensors such as a distance measuring sensor, an acceleration sensor, a gyro sensor, a geomagnetic sensor, a vibration sensor, a light sensor, and a sound sensor. The sensor acquires information regarding the state of the information processing device 900 itself, such as the posture of the housing of the information processing device 900, and information regarding the surrounding environment of the information processing device 900, such as the brightness and noise around the information processing device 900. Further, the sensor may include a GPS (Global Positioning System) sensor that receives a GPS signal and measures the latitude, longitude, and altitude of the device.

[0232] The output device 917 is composed of a device capable of visually or auditorily notifying the user of the acquired information. The output device 917 can be, for example, a display device such as an LCD (Liquid Crystal Display) or an organic EL (Electro-Luminescence) display, a sound output device such as a speaker and headphones. Also, the output device 917 may include a PDP (Plasma Display Panel), a projector, a hologram, a printer device, etc. The output device 917 outputs the result obtained by the processing of the information processing device 900 as video such as text or an image, or as sound such as voice or sound. Also, the output device 917 may include an illumination device that brightens the surroundings.

[0233] The storage device 919 is a data storage device configured as an example of the storage unit of the information processing device 900. The storage device 919 is composed of, for example, a magnetic storage device such as an HDD (Hard Disk Drive), a semiconductor storage device, an optical storage device, or a magneto-optical storage device. This storage device 919 stores the programs executed by the CPU 901, various data, and various data acquired from the outside.

[0234] The drive 921 is a reader / writer for a removable recording medium 927 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, and is built into or externally attached to the information processing device 900. The drive 921 reads the information recorded on the mounted removable recording medium 927 and outputs it to the RAM 905. Also, the drive 921 writes a record to the mounted removable recording medium 927.

[0235] The connection port 923 is a port for directly connecting a device to the information processing apparatus 900. The connection port 923 can be, for example, a USB (Universal Serial Bus) port, an IEEE1394 port, a SCSI (Small Computer System Interface) port, or the like. Also, the connection port 923 may be an RS-232C port, an optical audio terminal, an HDMI (registered trademark) (High-Definition Multimedia Interface) port, or the like. By connecting the external connection device 929 to the connection port 923, various data can be exchanged between the information processing apparatus 900 and the external connection device 929.

[0236] The communication device 925 is, for example, a communication interface configured by a communication device for connecting to the network 30 or the like. The communication device 925 can be, for example, a communication card for a wired or wireless LAN (Local Area Network), Bluetooth (registered trademark), Wi-Fi (registered trademark), or WUSB (Wireless USB). Also, the communication device 925 may be an optical communication router, an ADSL (Asymmetric Digital Subscriber Line) router, or a modem for various communications. The communication device 925 transmits and receives signals and the like using a predetermined protocol such as TCP / IP, for example, between the Internet and other communication devices. Also, the network 30 connected to the communication device 925 is a network connected by wire or wirelessly, and is, for example, the Internet, a home LAN, infrared communication, radio wave communication, or satellite communication.

[0237] <6. Supplementary> As described above, the preferred embodiments of the present disclosure have been described in detail with reference to the accompanying drawings, but the technical scope of the present disclosure is not limited to such examples. It is obvious that a person having ordinary knowledge in the technical field of the present disclosure can conceive of various modification examples or correction examples within the scope of the technical idea described in the claims, and it is naturally understood that these also belong to the technical scope of the present disclosure.

[0238] For example, the user data 2001 described in the above embodiment may further include data different from the data described above. For example, the user data 2001 may include biometric information such as the heart rate or body temperature of each user.

[0239] In addition, a computer program for causing hardware such as a CPU, ROM, and RAM incorporated in the server 10 or the user terminal 20 described above to exhibit the functions of the server 10 or the user terminal 20 can also be created. Further, a computer-readable storage medium storing the computer program is also provided.

[0240] Also, the effects described in this specification are merely illustrative or exemplary and not limiting. That is, the technology according to the present disclosure can exhibit other effects that are apparent to those skilled in the art from the description of this specification, together with or instead of the above effects.

[0241] Note that the following configurations also belong to the technical scope of the present disclosure. (1) An information processing apparatus including a control unit that performs control to generate a virtual space, wherein the virtual space includes a first avatar operated by a user and a second avatar different from the first avatar, and the control unit performs control to generate a reflection motion to be reflected on the second avatar based on user data obtained from the user. Information processing apparatus. (2) The information processing apparatus according to (1) above, wherein the control unit performs control to generate sound to be reflected in a partial space in the virtual space where the second avatar exists, using the user data or using the user data and the reflection motion. (3) The user includes a performer user who is a performer participating in an event performed in the virtual space. The user data includes sound data of music used in the event. The control unit performs control to generate the reflection motion based on the sound data, and the information processing apparatus according to (1) or (2) above. (4) The user data includes at least any one of the motion data of the performer user and the sound data of the sound emitted by the performer user, and the information processing apparatus according to (3) above. (5) The user includes an audience user who is an audience of the event and operates a first avatar, The user data includes at least any one of the motion data of the audience user and the sound data of the sound emitted by the audience user, and the information processing apparatus according to (3) or (4) above. (6) The user data includes the position of the first avatar in the virtual space, The control unit performs control to generate the reflection motion further based on the position of the second avatar in the virtual space, and the information processing apparatus according to any one of (1) to (5) above. (7) The control unit generates the reflection motion based on a plurality of the user data, Among the plurality of user data, the closer the position of the first avatar of the user corresponding to the user data in the virtual space is to the position of the second avatar, the greater the influence on the generation of the reflection motion, and generates the reflection motion, The information processing apparatus according to (6) above. (8) The user data includes the position of the first avatar in the virtual space, The control unit performs control to generate the sound based on the relationship between the position of the first avatar and the location of the partial space where the second avatar exists, and the information processing apparatus according to (2) above. (9) The control unit inputs the motion output by the first machine learning model that outputs a motion using the sound data of the music used in the event occurring in the virtual space as input data, the user data or the virtual space information generated based on the user data, and the position of the second avatar in the virtual space into a second machine learning model, thereby obtaining the reflected motion. The information processing apparatus according to any one of (1) to (8) above. (10) When the processing load of the information processing apparatus is greater than a threshold value, the control unit generates the sound without using the reflected motion. The information processing apparatus according to (2) or (8) above. (11) When the processing load of the information processing apparatus is greater than a threshold value, the control unit reduces the number of user data used when generating the reflected motion as compared with the case where the processing load is less than the threshold value. The information processing apparatus according to any one of (1) to (10) above. (12) The control unit performs control to determine the brightness for each partial space on the virtual space, and performs control to generate the virtual space without reflecting the reflected motion generated for the second avatar existing in the partial space where the brightness is lower than the threshold value. The information processing apparatus according to any one of (1) to (11) above. The information processing apparatus according to any one of (1) to (11) above. (13) The control unit generates the motion of each of the plurality of second avatars existing at different positions on the virtual space. The information processing apparatus according to any one of (1) to (12) above. (14) The control unit performs control to generate the sound to be reflected in each of the plurality of partial spaces where the second avatar exists at different positions on the virtual space. The information processing apparatus according to (2), (8), and (10) above. (15) The control unit performs control to obtain the excitement level of the event occurring in the virtual space, performing control to determine the volume of sound reflected in the partial space for each partial space according to the degree of swelling; The information processing apparatus according to (14) above. (16) The information processing apparatus according to any one of (1) to (15) above, wherein the user data includes information regarding the age, gender, or preferences regarding performance of the user. (17) comprising a control unit that performs control to generate a virtual space; the virtual space includes a first avatar operated by a user and a second avatar different from the first avatar; the control unit performs control to generate sound reflected in the partial space where the second avatar exists among one or more partial spaces in the virtual space based on user data obtained from the user; Information processing apparatus. (18) the user data includes the position of the first avatar in the virtual space; the control unit, generates the sound based on a plurality of the user data, generates the sound such that the user data corresponding to the first avatar closer to the location of the partial space has a greater influence on the generation of the sound; The information processing apparatus according to (17) above. (19) performing control to generate a virtual space, including, the virtual space includes a first avatar operated by a user and a second avatar different from the first avatar; performing control to generate a reflection motion reflected on the second avatar based on user data obtained from the user; A method executed by a processor, including. (20) operating a computer as a control unit that performs control to generate a virtual space, The virtual space includes a first avatar operated by a user and a second avatar different from the first avatar. The control unit performs control to generate a reflection motion to be reflected on the second avatar based on user data obtained from the user. Program.

Explanation of Signs

[0242] 10 Server 110 Communication Unit 120 Control Unit 121 Motion Generation Unit 122 Sound Generation Unit 123 Virtual Space Generation Unit 130 Storage Unit 131 User Attribute Information 20 User Terminal 21 HMD 30 Network

Claims

1. An information processing apparatus comprising a control unit that performs control to generate a virtual space, wherein the virtual space includes a first avatar operated by a user and a second avatar different from the first avatar, and the control unit performs control to generate a reflection motion to be reflected on the second avatar based on user data obtained from the user. Information processing apparatus.

2. The information processing apparatus according to claim 1, wherein the control unit performs control to generate sound to be reflected in a partial space where the second avatar exists among one or more partial spaces in the virtual space, using the user data or using the user data and the reflection motion.

3. The user includes a performer user who is a performer participating in an event performed in the virtual space, the user data includes sound data of music used in the event, and the control unit performs control to generate the reflection motion based on the sound data. The information processing apparatus according to claim 1.

4. The information processing apparatus according to claim 3, wherein the user data includes at least one of motion data of the performer user and sound data of sound emitted by the performer user.

5. The user includes a spectator user who is a spectator of the event and operates the first avatar, and the user data includes at least one of motion data of the spectator user and sound data of sound emitted by the spectator user. The information processing apparatus according to claim 3.

6. The user data includes the position of the first avatar in the virtual space, and the control unit performs control to generate the reflection motion based on the position of the second avatar in the virtual space. The information processing apparatus according to claim 1.

7. The control unit generates the reflection motion based on the plurality of pieces of user data, and generates the reflection motion such that, among the plurality of pieces of user data, user data in which the position of the first avatar of the user corresponding to the user data in the virtual space is closer to the position of the second avatar has a greater influence on the generation of the reflection motion. The information processing apparatus according to claim 6.

8. The user data includes the position of the first avatar in the virtual space, and the control unit performs control to generate the sound based on the relationship between the position of the first avatar and the location of the partial space where the second avatar exists. The information processing apparatus according to claim 2.

9. The control unit inputs, to a second machine learning model, the motion output by a first machine learning model that outputs a motion using sound data of a music piece used in an event performed in the virtual space as input data, the user data or virtual space information generated based on the user data, and the position of the second avatar in the virtual space, thereby obtaining the reflection motion. The information processing apparatus according to claim 1.

10. When the processing load of the information processing apparatus is greater than a threshold value, the control unit generates the sound without using the reflection motion. The information processing apparatus according to claim 2.

11. When the processing load of the information processing apparatus is greater than a threshold value, the control unit reduces the number of pieces of user data used when generating the reflection motion as compared with the case where the processing load is less than the threshold value. The information processing apparatus according to claim 1.

12. The control unit performs control to determine the brightness for each partial space on the virtual space, Performing control to generate the virtual space without reflecting the reflected motion generated for the second avatar existing in the partial space where the brightness is lower than the threshold value. The information processing apparatus according to claim 1.

13. The information processing apparatus according to claim 1, wherein the control unit generates the motion of each of the plurality of second avatars existing at different positions on the virtual space.

14. The information processing apparatus according to claim 2, wherein the control unit performs control to generate sound that is reflected in each of the plurality of partial spaces where the second avatar exists at different positions on the virtual space.

15. The control unit performs control to obtain the degree of excitement of an event occurring in the virtual space, and performs control to determine the volume of the sound reflected in the partial space for each partial space according to the degree of excitement. The information processing apparatus according to claim 14.

16. The information processing apparatus according to claim 1, wherein the user data includes information regarding the user's age, gender, or preferences regarding performance.

17. Comprising a control unit that performs control to generate a virtual space, the virtual space includes a first avatar operated by a user and a second avatar different from the first avatar, the control unit performs control to generate sound that is reflected in the partial space where the second avatar exists among one or more partial spaces in the virtual space based on user data obtained from the user. Information processing apparatus.

18. The user data includes the position of the first avatar in the virtual space, the control unit, generates the sound based on a plurality of the user data, The closer the user data corresponding to the first avatar is to the location of the partial space, the greater the influence on the generation of the sound, and the sound is generated accordingly. The information processing apparatus according to claim 17.

19. Performing control to generate a virtual space, The virtual space includes a first avatar operated by a user and a second avatar different from the first avatar. Performing control to generate a reflection motion reflected on the second avatar based on user data obtained from the user. A method executed by a processor, including the above.

20. Operating a computer as a control unit for performing control to generate a virtual space, The virtual space includes a first avatar operated by a user and a second avatar different from the first avatar. The control unit performs control to generate a reflection motion reflected on the second avatar based on user data obtained from the user. A program.

Citation Information

Patent Citations

  • Video application program, video object rendering method, video distribution system, video distribution server, and video distribution method

    JP2021152785A