Information processing apparatus, method, and program

By using machine learning models to generate reflection motions and sound effects for a second avatar in a virtual space, the system addresses the increased processing load issue, creating a more realistic and exciting virtual experience for large audiences.

WO2025121219A1PCT designated stage expired Publication Date: 2025-06-12SONY GROUP CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/041948
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-05
Filing Date
2024-11-27
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Existing information processing systems that create virtual spaces for events face increased processing loads when many users participate, leading to difficulties in reflecting motions of all users, which affects the excitement and realism of the virtual experience.

Method used

The system generates a virtual space with a first avatar operated by a user and a second avatar, where the system creates reflection motions and sound effects for the second avatar based on user data, reducing the processing load by using machine learning models to generate these elements.

Benefits of technology

This approach allows for a more realistic and exciting virtual experience with a large audience participation, while reducing the processing load on the system, thus enhancing the entertainment value of events held in virtual spaces.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024041948_12062025_PF_FP_ABST
    Figure JP2024041948_12062025_PF_FP_ABST
Patent Text Reader

Abstract

Provided is an information processing apparatus including circuitry configured to generate a virtual space, and initiate display of the virtual space, wherein the virtual space includes a first avatar operated by a user and a second avatar different from the first avatar, and wherein the circuitry is further configured to generate a reflection motion to be reflected in the second avatar, based on user data obtained from the user.
Need to check novelty before this filing date? Find Prior Art

Description

INFORMATION PROCESSING APPARATUS, METHOD, AND PROGRAMCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of Japanese Priority Patent Application JP 2023-205170 filed December 5, 2023, the entire contents of which are incorporated herein by reference.

[0002] The present disclosure relates to an information processing apparatus, a method, and a program.

[0003] In recent years, there has been widespread use of technologies to provide virtual spaces in which avatars of users are arranged. For example, there is a technology to provide a virtual space in which avatars of users who perform a performance at an event such as a music concert or a theatrical play, and avatars of multiple users who see the performance by the users are arranged at an event venue in a virtual space. Each avatar reflects motions of a user. As a result, the users can have an experience as if they are participating in the event together with other users.

[0004] An information processing apparatus that provides such a virtual space exchanges data with a terminal used by each user. Accordingly, in a case where many users participate in an event, the processing load of the information processing apparatus increases undesirably. In view of this, technologies for reducing the processing load of an information processing apparatus that provides a virtual space have been developed. For example, PTL 1 described below discloses a technology to reduce the processing load of an information processing apparatus that provides a virtual space, by not receiving user motion data from some terminals in connected terminals.

[0005] JP 2021-152785ASummary

[0006] However, in a case where motions are not reflected in avatars of some users from which motion data has not been received, it becomes difficult to produce, in a virtual space, an excitement produced by motions of a large audience at a real event venue, undesirably.

[0007] In view of this, the present disclosure proposes a novel and improved technology that can reduce a processing load when user avatars reflecting motions are arranged in a virtual space.

[0008] According to an embodiment of the present disclosure, there is provided an information processing apparatus including circuitry configured to generate a virtual space, and initiate display of the virtual space, wherein the virtual space includes a first avatar operated by a user and a second avatar different from the first avatar, and wherein the circuitry is further configured to generate a reflection motion to be reflected in the second avatar, based on user data obtained from the user.

[0009] In addition, according to another embodiment of the present disclosure, there is provided an information processing apparatus including circuitry configured to generate a virtual space, and initiate display of the generated virtual space, wherein the displayed virtual space includes a first avatar operated by a user and a second avatar different from the first avatar, and wherein the circuitry is further configured to generate a sound effect to be reflected in a subspace where the second avatar is located among one or more subspaces in the virtual space, based on user data obtained from the user.

[0010] Further, according to another embodiment of the present disclosure, there is provided a method executed by a processor, the method including generating a virtual space, the virtual space including a first avatar operated by a user and a second avatar different from the first avatar, displaying the generated virtual space, and generating a reflection motion to be reflected in the second avatar, based on user data obtained from the user.

[0011] Furthermore, according to another embodiment of the present disclosure, there is provided a non-transitory computer-readable storage medium having embodied thereon a program which when executed by a computer causes the computer to execute a method, the method including generating a virtual space, the virtual space including a first avatar operated by a user and a second avatar different from the first avatar, displaying the generated virtual space, and generating a reflection motion to be reflected in the second avatar, based on user data obtained from the user.

[0012] FIG. 1 is a figure for explaining the overall configuration of an information processing system 1 according to an embodiment of the present disclosure.FIG. 2 is a figure for explaining an example to be compared with the present embodiment regarding virtual space videos and sound effects to be generated in a case where the number of users to participate in an event is restricted to a small number.FIG. 3 is a figure for explaining an example of virtual space videos and sound effects generated by a server 10 according to the present embodiment.FIG. 4 is a figure for explaining an example of generation of a virtual space by the server 10 according to the present embodiment.FIG. 5 is a block diagram depicting an example of the configuration of the server 10 according to the present embodiment.FIG. 6 is a figure for explaining an example of generation of reflection motions 2010 with use of machine learning models by a motion generating section 121.FIG. 7 is a figure for explaining an example of generation of area sound effects 2020 with use of machine learning models by a sound effect generating section 122.FIG. 8 is a figure for explaining time required from collection of user data 2001 to reflection in a virtual space in a case where reflection motions 2010 are used for generation of area sound effects 2020.FIG. 9 is a figure for explaining time required from collection of user data 2001 to reflection in a virtual space in a case where reflection motions 2010 are not used for generation of area sound effects 2020.FIG. 10 is a flowchart depicting an example of the procedure of an operation process performed by the server 10 according to the present embodiment.FIG. 11 is a figure for explaining an example of the virtual space generated by a virtual space generating section 123 in a first modification example.FIG. 12 is a figure for explaining examples of reflection of area sound effects 2020 in each of multiple audience areas A in a case where the degree of excitement is "low" and in a case where the degree of excitement is "high."FIG. 13 is a flowchart depicting an example of the procedure of an operation process performed by the server 10 in a case where the first modification example and a second modification example are applied.FIG. 14 is a block diagram depicting a hardware configuration example of an information processing apparatus 900 that realizes the server 10 and user terminals 20 according to an embodiment of the present disclosure.Description of Embodiment

[0013] Hereinbelow, a preferred embodiment of the present disclosure is explained in detail with reference to the attached figures. Note that, in the present specification and the figures, constituent elements having substantially identical functional configurations are given identical reference signs, and overlapping explanations thereof are omitted.

[0014] Note that the explanation is given in the following order. 1. Overview 2. Configuration Example of Server 10 3. Operation Processing Example 4. Modification Examples 5. Hardware Configuration 6. Supplementary Notes

[0015] <1. Overview> First, the overall configuration of an information processing system according to an embodiment of the present disclosure is explained using FIG. 1. FIG. 1 is a figure for explaining the overall configuration of an information processing system 1 according to an embodiment of the present disclosure.

[0016] As depicted in FIG. 1, the information processing system 1 according to an embodiment of the present disclosure includes a server 10 and user terminals 20 (user terminals 20A, 20B, 20C …). As depicted in FIG. 1, the server 10 and the user terminals 20 are configured to be capable of communication via a network 30.

[0017] (Server 10) The server 10 has a function to provide a virtual space to users who use the user terminals 20. In the virtual space, a concert, a theatrical play, a sport game, or any of various other types of entertainment (events) can be held. In an example mainly explained in the present embodiment, a concert is held as an event, as an example.

[0018] Multiple users participate in the event. In the present embodiment, each user participates in the event either as a performer who performs a performance in the event or as an audience member who sees the performance performed by the performer. Hereinbelow, a user who is a performer is also referred to as a "performer user." In addition, a user who is an audience member is also referred to as an "audience user."

[0019] Note that the number of participants as users of the event (the number of participants as performer users, and the number of participants as audience users) is not limited. For example, it is sufficient if the number of participants as users, in particular the number of participants as audience users, is restricted according to the processing load of the server 10.

[0020] The server 10 generates the virtual space in which a user avatar which is an avatar corresponding to each user is arranged. More specifically, the user avatars include performer avatars corresponding to performer users and audience avatars corresponding to audience users. Each user avatar may be a 3D live-action avatar or may be a 3D character avatar.

[0021] Each user avatar is operated by a corresponding user. For example, each user avatar can move in the virtual space on the basis of operation of a user terminal 20 by a corresponding user. More specifically, according to user operation, a performer avatar may move on a stage provided at a virtual space concert venue. In addition, according to user operation, an audience avatar may move in an audience area outside the stage of the concert venue.

[0022] In addition, each user avatar reflects, in real time, motions acquired from a corresponding user. Motion data of motions acquired from each user is transmitted, in real time, from a user terminal 20 to the server 10.

[0023] In addition, the server 10 acquires sounds from each user in order to provide virtual space sound effects. Hereinbelow, sounds acquired from each user are also referred to as "user sounds."

[0024] Sounds acquired from performer users in user sounds include sounds generated by the performer users. Sounds generated by performer users may include sounds of singing by a performer user; sounds of a musical instrument played by a performer user; sounds of talking by a performer user; sounds of a conversation between performer users; and the like. Hereinbelow, sounds generated by performer users are also referred to as "performer sounds."

[0025] Furthermore, the server 10 may acquire a sound recording of music used at the event from a performer user. Hereinbelow, sound recordings of music used at the event are also referred to as "sound recordings" simply. For example, sound recordings may be sound recordings of music not including singing parts of performer users. Hereinbelow, sound recordings of music excluding singing parts of performer users are also called "backing tracks." The concert in the virtual space may be realized by reproducing a sound of real-time singing by a performer user as a sound superimposed on a backing track. In addition, sound recordings may be sound recordings of music including singing parts of a performer user or the like. The server 10 may acquire a performer sound of a performer user who is singing while a backing track is being reproduced. That is, the server 10 may acquire a sound including a backing track and a performer sound that are superimposed one on another. In addition, a sound recording may include a pre-recorded singing part of a performer user or another user.

[0026] In addition, sounds acquired from audience users in user sounds include sounds generated by audience users such as sounds of cheers from audience users for a performance by a performer user; sounds of handclapping; sounds of applauding; or sounds generated by sound-generating tools such as musical instruments held by audience users. Hereinbelow, sounds generated by audience users are also referred to as "audience sounds."

[0027] The server 10 transmits, to the user terminals 20, content data of the virtual space including data for rendering virtual space videos, and data for reproducing virtual space sound effects.

[0028] The data for rendering virtual space videos is data related to virtual objects such as user avatars to be arranged in the virtual space, map information regarding the virtual space, positional information regarding each user avatar in the virtual space, and the like. It is assumed in the present embodiment that the positional information regarding user avatars is information regarding the coordinates of the user avatars in the virtual space. Hereinbelow, the coordinates of user avatars in the virtual space are referred to as "user avatar coordinates." In addition, the coordinates of performer avatars in the virtual space are referred to as "performer avatar coordinates." In addition, the coordinates of audience avatars in the virtual space are referred to as "audience avatar coordinates."

[0029] For example, the data for reproducing virtual space sound effects is sound data of user sounds of each user, and sound recordings of music used at the event.

[0030] (User Terminals 20) The user terminals 20 are information processing terminals used by users. Each user in the present embodiment is either a performer user or an audience user. Each user participates in the event held in the virtual space generated by the server 10.

[0031] FIG. 1 depicts an example in which the user terminals 20 are PCs (Personal Computers), and an HMD (Head Mounted Display) 21 (21A, 21B, 21C …) is connected to each user terminal 20. The user terminals 20 perform control such that virtual space videos are displayed on the connected HMDs 21.

[0032] Note that the configuration of the user terminals 20 is not limited to the example depicted in FIG. 1, and each user terminal 20 may be configured as a single body which is any of various types of apparatuses such as a tablet terminal, a smartphone, an HMD, or a game terminal. Alternatively, each user terminal 20 may be configured as a combination of the various types of apparatuses described above. Alternatively, each user terminal 20 may be configured by connecting any of the various types of apparatuses described above with any of various types of terminals such as a display, a wearable device, a motion capturing apparatus, a camera, a microphone, or a speaker.

[0033] The user terminals 20 perform various kinds of data exchange with the server 10 through the network 30. The user terminals 20 transmit user data obtained from the users to the server 10. For example, the user data may include operation data representing the content of operation about the virtual space by the users.

[0034] The user terminals 20 acquire operation about the virtual space by the users. For example, the user terminals 20 acquire operation on user avatars arranged in the virtual space. For example, the operation on user avatars is operation for controlling the positions, facial expressions, gestures (emotional expressions by hands or the whole bodies), and action states (sitting, standing up, jumping, running, etc.) of the user avatars. For example, the operation on user avatars may be sensed by controllers operated by the users. In addition, operation information regarding viewer / listener avatars may be input from various types of operation input sections such as keyboards, mouses, or touch pads.

[0035] As operation by the users, the user terminals 20 may acquire motions that represent the positions and postures of body parts of the users, and the like with coordinates, rotation angles, and the like. User motion data is an example of the operation data. The motions of the users may be sensed by motion capturing apparatuses or may be sensed by cameras. The sensed motions of the users are reflected in user avatars in real time. Note that the operation on user avatars mentioned above can also be reflected in the avatars in real time. Hereinbelow, in a case where there is the phrase "user motion data," such data is not limited to action information regarding users obtained by sensing, and may include operation information regarding movements of avatars obtained by controllers.

[0036] The user terminals 20 output the acquired user motion data to the server 10. For example, the user terminals 20 output, to the server 10 in real time, motion data to be reflected as motions of user avatars.

[0037] In addition, the user data may include user sound data. The user terminals 20 may acquire user sounds by using sound input apparatuses such as microphones. The user terminals 20 output the user sound data to the server 10 in real time.

[0038] In particular, user terminals 20 used by performer users may acquire performer sounds of the performer users who are singing while sound recordings are being reproduced. In this case, the user terminals 20 may output, to the server 10, sound data including the performer sounds and the sound recordings. Note that the user terminals 20 may upload, to the server 10, sound data of sound recordings in advance at a preparatory stage of the event.

[0039] In addition, the user terminals 20 receive content data of the virtual space from the server 10. The user terminals 20 combine virtual space videos and sound effects according to the content data received from the server 10, and output them to the users. As a result, the users can enjoy the event held in the virtual space.

[0040] Specifically, the user terminals 20 render virtual space videos from user viewpoints by arranging virtual objects on the basis of map information regarding the virtual space. The user viewpoints may be the viewpoints of user avatars or may be bird's eye viewpoints including user avatars in their angles of view. Note that the server 10 may generate videos from certain viewpoints in the virtual space, and transmit the videos to the user terminals 20. For example, the certain viewpoints may be user viewpoints or may be bird's eye viewpoints from which the whole of the event venue can be seen.

[0041] In addition, the user terminals 20 superimpose user sounds of each user, and reproduce virtual space sound effects. For example, the user terminals 20 may reproduce virtual space sound effects according to user avatar coordinates. For example, the user terminals 20 may reproduce virtual space sound effects such that, as the distance between the user avatar coordinates and user avatar coordinates of audience users who use the user terminal 20 decreases, the volumes of audience sounds of the audience users to be superimposed increase.

[0042] Note that it is expected that there is a need for increasing the volumes of sounds superimposed at a user terminal 20 used by a user (hereinbelow, also referred to as a "user A") such as sounds corresponding to a user avatar of a user (hereinbelow, also referred to as a "user B") who is a friend of the user A, or sounds of a favorite artist (i.e., a performer user) who is a generally-called a "favorite" of the user A. Accordingly, these sounds may be superimposed with larger volumes even in a case where the distance between a user avatar of the user A and the user avatar of the user B or the artist described above is long. For example, in such a case, the sounds corresponding to the user avatar of the user B or the artist described above may be superimposed with such a constant volume that the user A can listen to the sounds easily independently of the distance described above, or the volumes of the superimposed sounds may be adjusted to volumes larger than the volumes of sounds of other users who are at the same distance. Whether or not the user A and the user B are friends may be identified on the basis of user attribute information. For example, the user attribute information is user information set through an application or the like by the users before the event (e.g., information regarding ages, genders, heights, information regarding grouping with particular users other than them, tipping information, or user preferences, etc.). The favorite artist of the user A may be identified on the basis of the user attribute information mentioned before (e.g., tipping information, etc.).

[0043] In addition, the user terminals 20 may superimpose performer sounds such that the users can listen to the performer sounds easily independently of the user avatar coordinates of performer users. Note that the server 10 may generate virtual space sound effects (e.g., sound effects according to user avatar coordinates), and transmit sound data of the sound effects to the user terminals 20.

[0044] (Clearing up Problems) Here, as mentioned above, problems related to the processing load of the server 10 can occur in a case where the server 10 transmits and receives motion data and user sound data of many users to and from user terminals 20 in real time. Accordingly, it is possible that the amount of data that the server 10 transmits and receives to and from user terminals 20 is reduced, and the processing load of the server 10 is reduced by restricting the number of users to participate in the event to a small number.

[0045] However, in a case where the number of users to participate in the event is restricted to a small number, it becomes difficult to produce, in the virtual space, an excitement that is produced by motions and cheers of a large audience, and the like, undesirably. Typically, for example, the sense of unity of a large audience is created in the event by the audience moving, giving cheers or the like, and so on all at once along to singing of a performer. Then, an excitement is created by the audience, who has felt such a sense of unity, more wholeheartedly moving, giving cheers or the like, and so on.

[0046] FIG. 2 is a figure for explaining an example to be compared with the present embodiment regarding virtual space videos and sound effects to be generated in a case where the number of users to participate in an event is restricted to a small number. In the example depicted here, the number of audience users to participate in the event is restricted to a small number. A virtual space video 210 is a video in which performer avatars P (P1, P2, and P3) and audience avatars O (O1, O2, etc.) are arranged in a virtual space concert venue V. In the example explained using FIG. 2, the virtual space video 210 is combined with a virtual space sound effect 211 generated by superimposing a sound generated by a user corresponding to each of the performer avatars P and the audience avatars O one on another.

[0047] In the example depicted in FIG. 2, the number of the audience avatars O is small relative to the size of the virtual space concert venue V, and accordingly it is hard for audience users to feel the sense of unity with audience avatars O of other audience users who are around their own audience avatars O. In addition, it is also possible that the virtual space is generated according to the number of the audience avatars O such that the size of the virtual space concert venue V matches the number of the audience avatars O, but it is also possible that the small number of the audience avatars O prevents reproduction of the force of cheering to such an extent that users can feel the sense of unity.

[0048] In view of this, in addition to audience avatars O operated by audience users, the server 10 according to the present embodiment generates audience avatars that are not operated by audience users, and arranges them in the virtual space. Hereinbelow, in order to make a distinction between the audience avatars O operated by the audience users and the audience avatars not operated by audience users, the audience avatars O operated by the audience users are also referred to as "operated audience avatars O." An operated audience avatar O is a first avatar according to the present embodiment. In addition, hereinbelow, an audience avatar not operated by an audience user and generated by the server 10 is also referred to as a "generation audience avatar." A generation audience avatar is a second avatar according to the present embodiment.

[0049] From here, virtual space videos and sound effects generated by the server 10 according to the present embodiment are explained with reference to FIG. 3 and FIG. 4.

[0050] FIG. 3 is a figure for explaining an example of virtual space videos and sound effects generated by the server 10 according to the present embodiment. FIG. 3 depicts a virtual space video 212. The virtual space video 212 is a virtual space video rendered by a user terminal 20 on the basis of content data of the virtual space generated by the server 10.

[0051] Similarly to the virtual space video 210 explained using FIG. 2, the virtual space video 212 is a video of the virtual space concert venue V in which the performer avatars P (P1, P2, and P3) and the operated audience avatars O (O1, O2, etc.) are arranged. Then, generation audience avatars G (G1, G2, etc.) are further arranged in the virtual space of the virtual space video 212.

[0052] FIG. 4 is a figure for explaining an example of generation of the virtual space by the server 10 according to the present embodiment. As depicted in FIG. 4, the server 10 generates a reflection motion 2010 on the basis of user data 2001. As has been explained thus far, the user data 2001 includes motion data of each user, sound data such as user sounds (audience sounds and performer sounds) or sound recordings of music used at the event, and user avatar coordinates in the virtual space.

[0053] The reflection motion 2010 is a motion to be reflected in a generation audience avatar G. The server 10 may generate a reflection motion 2010 of each of multiple generation audience avatars G that are at different positions in the virtual space. The server 10 generates content data by including, in data for rendering a virtual space video, data related to generation audience avatars G such as appearance data or reflection motions 2010 of the generation audience avatars G.

[0054] As a result, a user terminal 20 can reproduce the virtual space video 212 of the virtual space in which user avatars (performer avatars P and operated audience avatars O), each of which is generated on the basis of the user data 2001, and reflects motions of each user, and generation audience avatars G reflecting reflection motions 2010 are merged. The virtual space video 212 may be generated by the server 10, transmitted to the user terminal 20 as the content data, and reproduced at the HMD 21. Note that details of generation of reflection motions 2010 are explained later.

[0055] As depicted in FIG. 3, the server 10 according to the present embodiment can realize generation of the virtual space in which a large number of audience members participate as compared with the example explained using FIG. 2, while reducing the processing load of the server 10. As a result, it can be expected that an excitement of the audience is produced in the virtual space, and accordingly enhancement of the entertainment of the event held in the virtual space can be realized.

[0056] In addition, as depicted in FIG. 4, the server 10 according to the present embodiment generates an area sound effect 2020 on the basis of the user data 2001.

[0057] Area sound effects 2020 are sound effects to be reflected in audience areas where there are generation audience avatars G in the virtual space. Audience areas are areas set by dividing in advance an area where there can be operated audience avatars O and generation audience avatars G in the virtual space. The audience areas are an example of subspaces which are partial spaces of the virtual space. For example, in the example depicted in FIG. 3, audience areas A1 to A4 are set in the virtual concert venue V. Hereinbelow, in a case where a particular distinction is not made among the audience areas A1 to A4, the audience areas A1 to A4 are referred to as "audience areas A."

[0058] Operated audience avatars O may move in audience areas A and between multiple audience areas A on the basis of user operation. In addition, generation audience avatars G may move in audience areas A and between multiple audience areas A under the control of the server 10. It should be noted that seats may be provided in advance in the virtual space, and the positions of operated audience avatars O and generation audience avatars G may be fixed.

[0059] As explained later as details of area sound effects 2020, area sound effects 2020 can include sounds such as cheers. The server 10 generates content data by including sound data of area sound effects 2020 in data for reproducing virtual space sound effects. When reproducing the virtual space sound effects, user terminals 20 further reproduce the area sound effects 2020 superimposing them on user sounds of each user.

[0060] A user terminal 20 may superimpose, on user sounds, only an area sound effect 2020 to be reflected in an audience area A where there is an operated audience avatar O of a user who uses the user terminal 20. In addition, a user terminal 20 may superimpose multiple area sound effects 2020 on user sounds on the basis of the relation between the user avatar coordinates of a user who uses the user terminal 20 and the location of an audience area A. For example, the user terminal 20 may superimpose each area sound effect 2020 on user sounds such that, as the distance between the user avatar coordinates and the location of the audience area A decreases, the volume of the reproduced area sound effect 2020 increases.

[0061] In the example depicted in FIG. 3 and FIG. 4, the virtual space video 212 is combined with a virtual space sound effect 213 obtained by merging an area sound effect 2020 with user sounds. As a result, the types and volumes of cheers or the like in the virtual space sound effect increase, and accordingly further enhancement of the entertainment of the event held in the virtual space can be realized.

[0062] The virtual space sound effect 213 may be generated by the server 10, transmitted to a user terminal 20 as the content data, combined with the virtual space video 212 at the HMD 21, and reproduced. In addition, the virtual space video 212 and the sound effect 213 may be combined by the server 10, transmitted as content data to a user terminal 20, and reproduced at the HMD 21. Note that the user terminal 20 may continuously acquire content data obtained by combining the virtual space video 212 and the sound effect 213 from the server 10 in a streaming scheme, and reproduce the content data at the HMD 21.

[0063] Next, the specific configuration of each apparatus included in the information processing system 1 according to the present embodiment is explained with reference to figures.

[0064] <2. Configuration Example of Server 10> First, a configuration example of the server 10 is explained using FIG. 5. FIG. 5 is a block diagram depicting an example of the configuration of the server 10 according to the present embodiment.

[0065] As depicted in FIG. 5, the server 10 includes a communicating section 110, a control section 120, and a storage section 130.

[0066] (Communicating Section 110) The communicating section 110 performs data transmission / reception by wired or wireless communication connection with the user terminal 20. For example, the communicating section 110 can perform communication using a wired / wireless LAN, Wi-Fi (registered trademark), Bluetooth (registered trademark), infrared communication, a mobile communication network (4G (fourth generation mobile communication scheme), 5G (fifth generation mobile communication scheme)), or the like.

[0067] For example, the communicating section 110 receives user data 2001 of each user from a user terminal 20. In addition, the communicating section 110 transmits content data to each user terminal 20.

[0068] (Control Section 120) The control section 120 functions as a computation processing apparatus and a control apparatus, and controls operation in general in the server 10 according to various types of programs. For example, the control section 120 is realized by an electronic circuit such as a CPU (Central Processing Unit) or a microprocessor. In addition, the control section 120 may include a ROM (Read Only Memory) that stores programs, computation parameters, or the like to be used, and a RAM (Random Access Memory) that temporarily stores parameters or the like that change as appropriate.

[0069] On the basis of data received from an external apparatus, the control section 120 performs processes as appropriate, and performs control of storage on the storage section 130, control of data transmission to an external apparatus, and the like.

[0070] In addition, the control section 120 functions also as a motion generating section 121, a sound effect generating section 122, and a virtual space generating section 123.

[0071] (Motion Generating Section 121) On the basis of each piece of user data 2001 received by the communicating section 110, the motion generating section 121 generates reflection motions 2010 to be reflected in generation audience avatars G. The motion generating section 121 generates the reflection motions 2010 such that the generation audience avatars G represent motions that are appropriate for the situation of the event.

[0072] For example, the motion generating section 121 may generate reflection motions 2010 to be reflected in a predetermined number of frames. The predetermined number is not limited particularly to any number, and, for example, may be set such that reflection motions 2010 to be reproduced in one second to several seconds are generated. The motion generating section 121 repetitively executes the process of generating reflection motions 2010. For example, by repetition of the generation process while the event is being held, reflection motions 2010 may be reflected in generation audience avatars G consecutively while the event is being held.

[0073] The motion generating section 121 may generate reflection motions 2010 with use of machine learning models generated using a machine learning technology such as DNN (Deep Neural Network). FIG. 6 is a figure for explaining an example of generation of reflection motions 2010 with use of machine learning models by the motion generating section 121.

[0074] As depicted in FIG. 6, a reflection motion 2010 is generated using a pre-seed motion generation model M1 and a reflection motion generation model M2. The pre-seed motion generation model M1 is a first machine learning model according to the present embodiment. The reflection motion generation model M2 is a second machine learning model according to the present embodiment.

[0075] The pre-seed motion generation model M1 is a machine learning model that outputs a pre-seed motion 2011 using sound data of a sound recording B as input data. The reflection motion generation model M2 is a machine learning model that generates the reflection motion 2010 using, as input data, the pre-seed motion 2011, user data 2001, and generation avatar coordinates 1001. The user data 2001 includes performer data 2001A, which is user data 2001 of a performer user, and audience data 2001B, which is user data 2001 of an audience user.

[0076] The pre-seed motion 2011 is a motion that serves as the base (seed) in a case where the reflection motion generation model M2 outputs the reflection motion 2010. The motion generating section 121 inputs, to the reflection motion generation model M2, the pre-seed motion 2011 output from the pre-seed motion generation model M1. Then, using the pre-seed motion 2011 as the base, the reflection motion generation model M2 changes the pre-seed motion 2011, and generates the reflection motion 2010 on the basis of the user data 2001 and the generation avatar coordinates 1001.

[0077] The motion generating section 121 may input, to the pre-seed motion generation model M1, the sound data of the sound recording B which is acquired in real time from a user terminal 20 used by a performer user. As mentioned above, the sound recording B may include or may not include a singing part. That is, the motion generating section 121 may input, as the sound recording B, sound data in which sounds of singing by the performer user are superimposed on a backing track. In addition, the motion generating section 121 may input, to the pre-seed motion generation model M1, the sound recording B which is acquired in advance from the user terminal 20 at a preparatory stage of the event. In this case, the sound recording B may include a pre-recorded singing part. In a case where the sound recording B includes a singing part, the pre-seed motion generation model M1 may generate the pre-seed motion 2011 according to the content of lyrics. As a result, it is possible to cause a generation audience avatar G to represent a motion generated taking into consideration the meaning of music sung by the performer user.

[0078] The pre-seed motion generation model M1 may be a machine learning model that generates the pre-seed motion 2011 according to the melody of the sound recording B.

[0079] As an example, the pre-seed motion generation model M1 may generate the pre-seed motion 2011 such that the pre-seed motion 2011 represents a motion to move along to the rhythm of the sound recording B.

[0080] In addition, as another example, the pre-seed motion generation model M1 may generate the pre-seed motion 2011 that moves along to the melody of each part of the sound recording B. For example, the pre-seed motion 2011 may be generated such that the pre-seed motion 2011 represents a less intense motion at the intro part of the sound recording B as compared with other parts. In addition, for example, the pre-seed motion generation model M1 may generate the pre-seed motion 2011 such that the pre-seed motion 2011 represents motions that become more intense toward a chorus part of the sound recording B.

[0081] When singing is represented at the event, the audience typically moves along to music sung. Accordingly, by generating the pre-seed motion 2011, which serves as the base of the reflection motion 2010, on the basis of the sound data of the sound recording B, it is possible to generate the reflection motion 2010 which is appropriate for the situation of the event.

[0082] Note that the pre-seed motion generation model M1 may be trained in any manner on the basis of sound recordings B as long as the pre-seed motion generation model M1 can output the pre-seed motion 2011 appropriate as a motion for singing using the sound recording B at the event.

[0083] By using the pre-seed motion generation model M1, the motion generating section 121 can change the generated pre-seed motion 2011, and generate the reflection motion 2010 that moves along to the melody of the sound recording B. As a result, it is possible to cause the generation audience avatar G to represent a motion appropriate for music sung by the performer user.

[0084] The motion generating section 121 inputs, to the reflection motion generation model M2, the pre-seed motion 2011 output from the pre-seed motion generation model M1. In addition, the motion generating section 121 inputs, to the reflection motion generation model M2, the user data 2001 and the generation avatar coordinates 1001 along with the pre-seed motion 2011. The generation avatar coordinates 1001 are coordinates representing the position of the generation audience avatar G in the virtual space.

[0085] The motion generating section 121 inputs, to the reflection motion generation model M2, the user data 2001 that has been obtained in predetermined time until the current time. For example, the predetermined time may be one second to several seconds.

[0086] The motion generating section 121 may input, to the reflection motion generation model M2, user data 2001 of all users participating in the event.

[0087] In addition, the motion generating section 121 may sample partial user data 2001 out of user data 2001 of all users participating in the event, and input the sampled partial user data 2001 to the reflection motion generation model M2. For example, the motion generating section 121 may sample partial audience data 2001B out of audience data 2001B of all audience users participating in the event, and input the sampled partial audience data 2001B to the reflection motion generation model M2.

[0088] Whether to or not to sample audience data 2001B may be decided on the basis of the processing load of the server 10. For example, the processing load of the server 10 may be represented by the amount of used memory or power usage of a GPU (Graphics Processing Unit) or the CPU (Central Processing Unit) of the server 10, or the like. For example, the motion generating section 121 may generate reflection motions 2010 by sampling audience data 2001B in a case where the processing load of the server 10 is greater than a threshold.

[0089] For example, the sampling of audience data 2001B may be performed by randomly selecting a predetermined percentage of operated audience avatars O for each audience area A.

[0090] The generation avatar coordinates 1001 are the coordinates of the generation audience avatar G in the virtual space. The generation avatar coordinates 1001 are an example of the position of the generation audience avatar G in the virtual space.

[0091] For example, the reflection motion generation model M2 may generate the reflection motion 2010 such that, as user avatar coordinates get closer to the generation avatar coordinates 1001, the influence, on the generation of the reflection motion 2010, of user data 2001 corresponding to the user avatar coordinates increases. The reflection motion generation model M2 may generate the reflection motion 2010 using only user data 2001 of users corresponding to user avatar coordinates at distances to the generation avatar coordinates 1001 which are equal to or shorter than a predetermined distance. The reflection motion generation model M2 may generate the reflection motion 2010 by weighting each piece of user data 2001 according to the distance between the generation avatar coordinates 1001 and corresponding user avatar coordinates.

[0092] It is expected that audience members represent motions being more influenced by performers and audience members who are closer to them than by performers and audience members who are away from them. Accordingly, by generating the reflection motion 2010 on the basis of the generation avatar coordinates 1001 and user avatar coordinates, it is possible to cause the generation audience avatar G to represent a motion appropriate for the situation of the event.

[0093] Note that, in a case where the generation avatar coordinates 1001 change dynamically, the motion generating section 121 acquires the generation avatar coordinates 1001, and generates the reflection motion 2010 every time the generation process is performed. Here, the generation avatar coordinates 1001 may change according to the reflection motion 2010. On the other hand, in a case where the generation avatar coordinates 1001 are fixed throughout the event, the motion generating section 121 may use the same coordinates for the process of generating the reflection motion 2010 executed consecutively during the event.

[0094] Note that, whereas the reflection motion 2010 is generated weighting each piece of user data 2001 according to the distance between the generation avatar coordinates 1001 and corresponding user avatar coordinates in the example explained in the description above, this example of weighting of a piece of user data 2001 is not the sole example. For example, the reflection motion generation model M2 may generate the reflection motion 2010 weighting each piece of user data 2001 according to preset feature information regarding the generation audience avatar G. For example, the feature information regarding the generation audience avatar G may be information that is set as features of the generation audience avatar G, and is about the age, the gender, the height, information regarding grouping with particular users, or favorite artists.

[0095] For example, the information regarding grouping with particular users may be information representing that the generation avatar and the particular users are friends. In this case, the generation avatar and the users may be set as friends on the basis of operation of user terminals 20 by users.

[0096] As a more specific example, the reflection motion generation model M2 may weight each piece of user data 2001 according to user attribute information regarding a user who is a friend of the generation audience avatar G. For example, the reflection motion generation model M2 may weight performer data 2001A such that, as the degree of match between a performer user and favorite artists of a user who is a friend of the generation audience avatar G increases, the influence, on the generation of the reflection motion 2010, of performer data 2001A of the performer user increases. In addition, the reflection motion generation model M2 may weight user data 2001 such that the influences, on the generation of the reflection motion 2010, of user data 2001 of a user who is a friend of the generation audience avatar G and of user data 2001 of users who are near the user are increased.

[0097] As another example, the reflection motion generation model M2 may calculate the degree of match between the favorite artists of the generation audience avatar G and each performer user, and weight performer data 2001A according to the degrees of match. Then, the reflection motion generation model M2 may generate the reflection motion 2010 such that, as the degree of match of a performer user increases, the influence, on the generation of the reflection motion 2010, of performer data 2001A of the performer user increases.

[0098] Next, generation of the reflection motion 2010 based on various types of data included in user data 2001 by the reflection motion generation model M2 is explained. The reflection motion generation model M2 generates the reflection motion 2010 on the basis of performer data 2001A included in user data 2001. For example, the reflection motion generation model M2 generates the reflection motion 2010 on the basis of performer motion data 2002A in the performer data 2001A. The reflection motion generation model M2 may generate the reflection motion 2010 such that the intensity of movement of the generation audience avatar G increases as the intensity of a motion represented by the performer motion data 2002A increases.

[0099] In addition, the reflection motion generation model M2 may generate the reflection motion 2010 such that the motion represented by the performer motion data 2002A and the reflection motion 2010 resemble each other. In particular, in a case where the motion represented by the performer motion data 2002A is a particular motion (a handclapping motion, a motion to wave a hand at a constant rhythm, etc.), the reflection motion generation model M2 may generate the reflection motion 2010 such that the generation audience avatar G represents a particular motion.

[0100] It is expected that, at the event, audience members move being influenced by motions of performers. Accordingly, by generating the reflection motion 2010 on the basis of the performer motion data 2002A, it is possible to cause the generation audience avatar G to represent a motion appropriate for the situation of the event.

[0101] In addition, the reflection motion generation model M2 may generate the reflection motion 2010 on the basis of performer sound data 2003A. For example, the reflection motion generation model M2 may generate the reflection motion 2010 such that the intensity of movement of the generation audience avatar G increases as the volume of a performer sound represented by the performer sound data 2003A increases.

[0102] In addition, the reflection motion generation model M2 may generate the reflection motion 2010 such that the generation audience avatar G represents a particular motion in a case where the performer sound data 2003A includes a sound of the performer user uttering a particular phrase. For example, the reflection motion generation model M2 may generate the reflection motion 2010 such that the generation audience avatar G claps its hands in a case where the performer sound data 2003A includes a sound of the performer user uttering the phrase "clap hands."

[0103] It is expected that, at the event, audience members move being influenced by sounds generated by performers. Accordingly, by generating the reflection motion 2010 on the basis of the performer sound data 2003A, it is possible to cause the generation audience avatar G to represent a motion appropriate for the situation of the event.

[0104] In addition, the reflection motion generation model M2 generates the reflection motion 2010 on the basis of the audience data 2001B in the user data 2001.

[0105] For example, the reflection motion generation model M2 generates the reflection motion 2010 on the basis of audience motion data 2002B included in the audience data 2001B. The reflection motion generation model M2 may generate the reflection motion 2010 such that the intensity of movement of the generation audience avatar G increases as the intensity of a motion represented by the audience motion data 2002B increases. In addition, the reflection motion generation model M2 may generate the reflection motion 2010 such that the motion represented by the audience motion data 2002B and the reflection motion 2010 resemble each other. In particular, the reflection motion generation model M2 may generate the reflection motion 2010 such that the generation audience avatar G represents a particular motion in a case where the motion represented by the audience motion data 2002B is a particular motion (a handclapping motion, a motion to wave a hand at a constant rhythm, etc.).

[0106] It is expected that, at the event, audience members represent motions being influenced by motions of other audience members. Accordingly, by generating the reflection motion 2010 on the basis of the audience motion data 2002B, it is possible to cause the generation audience avatar G to represent a motion appropriate for the situation of the event.

[0107] In addition, the reflection motion generation model M2 may generate the reflection motion 2010 on the basis of audience sound data 2003B. For example, the reflection motion generation model M2 may generate the reflection motion 2010 such that the intensity of movement of the generation audience avatar G increases as the volume of an audience sound represented by the audience sound data 2003B increases.

[0108] It is expected that, at the event, audience members represent motions being influenced by sounds generated by other audience members. Accordingly, by generating the reflection motion 2010 on the basis of the audience sound data 2003B, it is possible to cause the generation audience avatar G to represent a motion appropriate for the situation of the event.

[0109] Note that, whereas the user data 2001 is input to the reflection motion generation model M2 in the example explained thus far, virtual space information including virtual space videos and sound effects generated on the basis of the user data 2001 may be input to the reflection motion generation model M2 along with the user data 2001, or instead of the user data 2001.

[0110] Virtual space videos and sound effects to be reproduced at a user terminal 20 are generated on the basis of content data managed by the virtual space generating section 123 mentioned later. Accordingly, the motion generating section 121 may recreate virtual space videos and sound effects to be generated at each user terminal 20, on the basis of content data managed by the virtual space generating section 123. Then, the motion generating section 121 may generate the reflection motion 2010 on the basis of the recreated virtual space videos and sound effects. It should be noted that the virtual space videos and sound effects generated at each user terminal 20 may be acquired later on by the server 10.

[0111] It is expected that audience members represent motions on the basis of motions of audience members and performers that are visible to them, and sounds that can be heard through their ears. Accordingly, by generating the reflection motion 2010 on the basis of virtual space videos and sound effects reproduced at a user terminal 20 of a user participating in the event, in particular an audience user, it is possible to cause the generation audience avatar G to represent a motion appropriate for the situation of the event.

[0112] In addition, the reflection motion generation model M2 may generate the reflection motion 2010 further on the basis of user attribute information 131 stored on the storage section 130 mentioned later. User attribute information 131 is an example of user data 2001. User attribute information 131 is information representing an attribute of each user. For example, user attribute information 131 may be information regarding the age, gender, or performance-related preference of each user. For example, the information regarding performance-related preference may be information regarding a favorite song of a user, a favorite music video of a user, or the like. User attribute information 131 is acquired by the communicating section 110 in advance on the basis of operation of a user terminal 20 by a user.

[0113] Motions of audience members represent tendencies which are different according to ages, genders, performance-related preferences, or the like, in some cases. Accordingly, by generating a reflection motion 2010 on the basis of user attribute information 131 about each user, it is possible to cause a generation audience avatar G to represent a motion matching a user participating in the event.

[0114] In addition, the reflection motion generation model M2 may generate the reflection motion 2010 according to preset feature information regarding the generation audience avatar G. For example, the reflection motion generation model M2 may generate the reflection motion 2010 such that the reflection motion 2010 is influenced more by a motion of a user who is a friend of the generation audience avatar G than by a user who is not a friend of the generation audience avatar G.

[0115] In addition, as another example, the reflection motion generation model M2 may generate the reflection motion 2010 such that the reflection motion 2010 represents a more intense movement in a case where the distance between the generation avatar coordinates 1001 and performer avatar coordinates of a performer user matching a favorite artist of the generation audience avatar G is shorter than a predetermined distance.

[0116] In such a manner, by the reflection motion generation model M2 generating the reflection motion 2010 according to the feature information regarding the generation audience avatar G, it is possible to give the generation audience avatar G individuality, and accordingly it is possible to recreate a more realistic audience member.

[0117] Note that, whereas it has been explained thus far that the reflection motion generation model M2 uses various types of data included in the user data 2001 in order to generate the reflection motion 2010, types of data to be used for generation of the reflection motion 2010 are not limited to any particular type.

[0118] For example, the reflection motion generation model M2 may generate the reflection motion 2010 using only partial data in the performer data 2001A and the audience data 2001B. In particular, use of the performer data 2001A may be prioritized over use of the audience data 2001B for generation of the reflection motion 2010. This is because it is expected that audience members participating in the event are most interested in motions and sounds of performers.

[0119] In addition, use of the performer motion data 2002A and the audience motion data 2002B may be prioritized over use of the performer sound data 2003A and the audience sound data 2003B for generation of the reflection motion 2010. This is because it is expected that motions of audience members are most influenced by motions of performers and other audience members.

[0120] Note that, in a case where the reflection motion 2010 is generated using multiple pieces of data in data included in the user data 2001, the reflection motion 2010 may be generated weighting each piece of data. For example, each piece of data may be weighted such that the performer motion data 2002A is most reflected in the reflection motion 2010. For example, a degree of priority may be set for each of multiple pieces of data included in the user data 2001. For example, degrees of priority may be set to integer values from 0 to 10. 0 may mean the lowest degree of priority, and 10 may mean the highest degree of priority. The reflection motion 2010 may be generated such that the influence of data on the reflection motion 2010 increases as the degree of priority of the data increases. The degrees of priority may be set automatically on the system side in advance, or may be set or changed as appropriate on the user side. Note that data to be used for generation of the reflection motion 2010 in data included in the user data 2001 may be changed dynamically according to the network condition, the resources or remaining battery level of each piece of processing equipment (the server 10, the user terminals 20, etc.), and the like. For example, in a case where the network condition is good, the reflection motion 2010 may be generated on the basis of data with degrees of priority of 5 to 10 in data included in the user data 2001, and in a case where the network condition is not good, the reflection motion 2010 may be generated only on the basis of data with degrees of priority of 9 to 10.

[0121] Generation of the reflection motion 2010 based on various types of data included in the user data 2001 by the reflection motion generation model M2 has been explained thus far. It should be noted that, as long as the reflection motion generation model M2 can output the reflection motion 2010 that is influenced by performer avatars P and operated audience avatars O at the event, and is appropriate for the situation of the event, the reflection motion generation model M2 is not limited to examples explained thus far, and may be trained in any manner. Note that reflection motions 2010 generated by the motion generating section 121 may be used for training of machine learning models.

[0122] As explained above, by with use of machine learning models for generation of reflection motions 2010, it is possible to generate motions more appropriate for the situation of the event, without complicated designing by the provider of the virtual space. Note that machine learning models may not be used for generation of reflection motions 2010. For example, the motion generating section 121 may generate reflection motions 2010 on the basis of predetermined rules about user data 2001 and reflection motions 2010.

[0123] The motion generating section 121 may generate each reflection motion 2010 to be reflected in one of multiple generation audience avatars G by inputting a pre-seed motion 2011 to the reflection motion generation model M2. In this case, the reflection motion generation model M2 may output multiple reflection motions 2010 each corresponding to one of multiple sets of generation avatar coordinates 1001 using the generation avatar coordinates 1001 as input data. In addition, the reflection motion generation model M2 may output a reflection motion 2010 corresponding to each of multiple sets of generation avatar coordinates 1001, using the set of generation avatar coordinates 1001 as input data.

[0124] In addition, whereas reflection motions 2010 are generated using the two machine learning models in the example explained using FIG. 6 here, only one machine learning model may be used for generation of reflection motions 2010. More specifically, only the reflection motion generation model M2 may be used for generation of reflection motions 2010. Here, the reflection motion generation model M2 may generate the reflection motions 2010 using, as input data, sound recordings B instead of pre-seed motions 2011.

[0125] (Sound Effect Generating Section 122) The sound effect generating section 122 generates area sound effects 2020 on the basis of each piece of user data 2001 received by the communicating section 110. The sound effect generating section 122 generates the area sound effects 2020 such that the area sound effects 2020 create sounds appropriate for the situation of the event.

[0126] The sound effect generating section 122 generates the area sound effects 2020 as sound effects to be reflected in audience areas A where there are generation audience avatars G in the virtual space. More specifically, the sound effect generating section 122 generates the area sound effects 2020 as sounds to be generated by the generation audience avatars G that are in the audience areas A. For example, the sound effect generating section 122 may generate sounds that audience members can generate during the event. For example, the sounds that audience members can generate during the event include cheers; sounds of handclapping; sounds of applauding; sounds generated by sound-generating tools such as musical instruments held by audience members; and the like.

[0127] For example, the sound effect generating section 122 may generate area sound effects 2020 to be reflected in a predetermined number of frames. The predetermined number is not limited particularly to any number, and, for example, may be set such that area sound effects 2020 to be reproduced in one second to several seconds are generated. The sound effect generating section 122 repetitively executes the process of generating area sound effects 2020. For example, by repetition of the generation process while the event is being held, area sound effects 2020 may be reflected in audience areas A consecutively while the event is being held.

[0128] The sound effect generating section 122 may generate area sound effects 2020 with use of machine learning models generated using a machine learning technology such as DNN. FIG. 7 is a figure for explaining an example of generation of area sound effects 2020 with use of machine learning models by the sound effect generating section 122.

[0129] As depicted in FIG. 7, an area sound effect 2020 is generated using a sound effect generation model M3. The sound effect generation model M3 is a machine learning model that outputs the area sound effect 2020 using, as input data, user data 2001 and data representing the location of a generation avatar area Ar. The sound effect generation model M3 outputs the area sound effect 2020 using the user data 2001 of multiple users.

[0130] The sound effect generating section 122 inputs, to the sound effect generation model M3, the user data 2001 that has been obtained in predetermined time until the current time. For example, the predetermined time may be one second to several seconds.

[0131] The sound effect generating section 122 may input, to the sound effect generation model M3, user data 2001 of all users participating in the event.

[0132] In addition, the sound effect generating section 122 may sample partial user data 2001 out of user data 2001 of all users participating in the event, and input the sampled partial user data 2001 to the sound effect generation model M3. The method of sampling the user data 2001 is similar to the method explained regarding the motion generating section 121.

[0133] The generation avatar area Ar is an audience area A where there are generation avatars, that is, where the generated area sound effect 2020 is to be reflected, in the audience areas A in the virtual space. The sound effect generation model M3 generates the area sound effect 2020 on the basis of the relation between the user avatar coordinates of each user and the location of the generation avatar area Ar.

[0134] For example, the sound effect generation model M3 may generate the area sound effect 2020 such that, as user avatar coordinates and the location of the generation avatar area Ar get closer to each other, the influence, on the generation of the area sound effect 2020, of user data 2001 corresponding to the user avatar coordinates increases. More specifically, the sound effect generation model M3 may generate the area sound effect 2020 such that the influence, on the generation of the area sound effect 2020, of audience data 2001B representing audience avatar coordinates that are included in the generation avatar area Ar becomes greater than the influences of other audience data 2001B.

[0135] In addition, the sound effect generation model M3 may generate the area sound effect 2020 without using audience data 2001B about operated audience avatars O that are in the audience areas A other than the generation avatar area Ar. In addition, the sound effect generation model M3 may generate the area sound effect 2020 weighting audience data 2001B on the basis of the distance between the location of the generation avatar area Ar and the audience avatar coordinates.

[0136] It is expected that audience members generate sounds being more influenced by performers and audience members who are closer to them than by performers and audience members who are away from them. Accordingly, by generating the area sound effect 2020 on the basis of the location of the generation avatar area Ar, it is possible to superimpose a sound effect appropriate for the situation of the event on a virtual space sound effect.

[0137] Note that, whereas the area sound effect 2020 is generated weighting audience data 2001B on the basis of the distance between the location of the generation avatar area Ar and the audience avatar coordinates in the example explained in the description above, this is not the sole example of weighting. For example, the sound effect generation model M3 may generate the area sound effect 2020 weighting each piece of user data 2001 according to preset feature information regarding a generation audience avatar G that is in the generation avatar area Ar.

[0138] As a more specific example, the sound effect generation model M3 may weight each piece of user data 2001 according to user attribute information regarding a user who is a friend of the generation audience avatar G that is in the generation avatar area Ar. For example, the sound effect generation model M3 may weight performer data 2001A such that, as the degree of match between a performer user and favorite artists of a user who is a friend of the generation audience avatar G increases, the influence, on the generation of the area sound effect 2020, of performer data 2001A of the performer user increases. In addition, the sound effect generation model M3 may weight user data 2001 such that the influences, on the generation of the area sound effect 2020, of user data 2001 of a user who is a friend of the generation audience avatar G and of user data 2001 of users who are near the user are increased.

[0139] As another example, the sound effect generation model M3 may calculate the degree of match between the favorite artists of the generation audience avatar G and each performer user, and weight performer data 2001A according to the degrees of match. Then, the sound effect generation model M3 may generate the area sound effect 2020 such that, as the degree of match of a performer user increases, the influence, on the generation of the area sound effect 2020, of performer data 2001A of the performer user increases.

[0140] Next, generation of the area sound effect 2020 based on various types of data included in user data 2001 by the sound effect generation model M3 is explained. The sound effect generation model M3 generates the area sound effect 2020 on the basis of the performer data 2001A included in the user data 2001.

[0141] For example, the sound effect generation model M3 generates the area sound effect 2020 on the basis of the performer sound data 2003A in the performer data 2001A. The sound effect generation model M3 may generate the area sound effect 2020 such that, as the volume of a performer sound represented by the performer sound data 2003A increases, the volume of the area sound effect 2020 increases. In addition, the sound effect generation model M3 may generate the area sound effect 2020 such that the area sound effect 2020 includes a particular sound in a case where the performer sound data 2003A includes a sound of the performer user uttering a particular phrase. For example, the sound effect generation model M3 may generate the area sound effect 2020 such that the area sound effect 2020 includes a sound of handclapping in a case where the performer sound data 2003A includes a sound of the performer user uttering the phrase "clap hands."

[0142] It is expected that, at the event, audience members generate sounds being influenced by sounds generated by performers. Accordingly, by generating the area sound effect 2020 on the basis of the performer sound data 2003A, it is possible to superimpose a sound effect appropriate for the situation of the event on a virtual space sound effect.

[0143] In addition, the sound effect generation model M3 generates the area sound effect 2020 on the basis of the performer motion data 2002A. For example, the sound effect generation model M3 may generate the area sound effect 2020 such that the volume of the area sound effect 2020 increases as the intensity of a motion represented by the performer motion data 2002A increases. Here, the volume of the area sound effect 2020 may be changed linearly or non-linearly according to the intensity of the motion represented by the performer motion data 2002A. In addition, the sound effect generation model M3 may generate the area sound effect 2020 such that the area sound effect 2020 includes a particular sound in a case where the motion represented by the performer motion data 2002A is a particular motion (a handclapping motion, a motion to wave a hand at a constant rhythm, etc.). For example, the sound effect generation model M3 may generate the area sound effect 2020 such that the area sound effect 2020 includes a sound of handclapping in a case where the motion represented by the performer motion data 2002A is a handclapping motion. In addition, "emotes" for causing avatars to represent preset movements are set for a game or the like that uses a virtual space such as a metaverse in some cases. "Emotes" can be included in the user data 2001. The sound effect generation model M3 may generate the area sound effect 2020 using an "emote" that is included in the performer data 2001A and preset by the performer user. The sound effect generation model M3 may generate the area sound effect 2020 on the basis of a motion represented by an emote by a method similar to that in a case where the performer motion data 2002A is used.

[0144] It is expected that, at the event, audience members generate sounds being influenced by motions of performers. Accordingly, by generating the area sound effect 2020 on the basis of the performer motion data 2002A, it is possible to superimpose a sound effect appropriate for the situation of the event on a virtual space sound effect.

[0145] In addition, the sound effect generation model M3 generates the area sound effect 2020 on the basis of the audience data 2001B in the user data 2001.

[0146] For example, the sound effect generation model M3 generates the area sound effect 2020 on the basis of the audience sound data 2003B in the audience data 2001B. For example, the sound effect generation model M3 may generate the area sound effect 2020 such that, as the volume of an audience sound represented by the audience sound data 2003B increases, the volume of the area sound effect 2020 increases. In addition, the sound effect generation model M3 may generate the area sound effect 2020 such that the audience sound represented by the audience sound data 2003B and the area sound effect 2020 resemble each other.

[0147] It is expected that, at the event, audience members generate sounds being influenced by sounds generated by other audience members. Accordingly, by generating the area sound effect 2020 on the basis of the audience motion data 2002B, it is possible to superimpose a sound effect appropriate for the situation of the event on a virtual space sound effect.

[0148] In addition, the sound effect generation model M3 generates the area sound effect 2020 on the basis of the audience motion data 2002B in the audience data 2001B. For example, the sound effect generation model M3 may generate the area sound effect 2020 such that the volume of the area sound effect 2020 increases as the intensity of a motion represented by the audience motion data 2002B increases. Here, the volume of the area sound effect 2020 may be changed linearly or non-linearly according to the intensity of the motion represented by the audience motion data 2002B. In addition, the sound effect generation model M3 may generate the area sound effect 2020 such that the area sound effect 2020 includes a particular sound in a case where the motion represented by the audience motion data 2002B is a particular motion (a handclapping motion, a motion to wave a hand at a constant rhythm, etc.). For example, the sound effect generation model M3 may generate the area sound effect 2020 such that the area sound effect 2020 includes a sound of handclapping in a case where the motion represented by the audience motion data 2002B is a handclapping motion. In addition, "emotes" by which avatars represent preset movements are set for a game or the like that uses a virtual space such as a metaverse in some cases. The sound effect generation model M3 may generate the area sound effect 2020 using an "emote" that is included in the audience data 2001B and preset by an audience user. The sound effect generation model M3 may generate the area sound effect 2020 on the basis of a motion represented by an emote by a method similar to that in a case where the audience motion data 2002B is used.

[0149] It is expected that, at the event, audience members generate sounds being influenced by motions of other audience members. Accordingly, by generating the area sound effect 2020 on the basis of the audience motion data 2002B, it is possible to superimpose a sound effect appropriate for the situation of the event on a virtual space sound effect.

[0150] Note that, whereas the user data 2001 is input to the sound effect generation model M3 in the example explained thus far, virtual space information including virtual space videos and sound effects generated on the basis of the user data 2001 may be input to the sound effect generation model M3 along with the user data 2001, or instead of the user data 2001.

[0151] It is expected that audience members generate sounds on the basis of motions of audience members and performers that are visible to them, and sounds that can be heard through their ears. Accordingly, by generating the area sound effect 2020 on the basis of virtual space videos and sound effects to be generated at a user terminal 20 of a user, in particular of an audience user, it is possible to superimpose a sound effect appropriate for the situation of the event on a virtual space sound effect.

[0152] In addition, the sound effect generation model M3 may generate the area sound effect 2020 further on the basis of user attribute information 131. Sounds generated by audience members represent tendencies which are different according to ages, genders, performance-related preferences, or the like, in some cases. Accordingly, by generating an area sound effect 2020 on the basis of user attribute information 131 about each user, it is possible to superimpose, on a virtual space sound effect, a sound matching an operated audience avatar O participating in the event.

[0153] In addition, the sound effect generation model M3 may generate the reflection motion 2010 according to feature information regarding the generation audience avatar G. For example, the generation audience avatar G may generate the area sound effect 2020 such that the area sound effect 2020 is influenced more by a motion of a user who is a friend of the generation audience avatar G than by a user who is not a friend of the generation audience avatar G.

[0154] In addition, as another example, the sound effect generation model M3 may generate the area sound effect 2020 such that the volume of the area sound effect 2020 increases in a case where the distance between the generation avatar coordinates 1001 and performer avatar coordinates of a performer user matching a favorite artist of the generation audience avatar G is shorter than a predetermined distance.

[0155] In such a manner, by the sound effect generation model M3 generating the area sound effect 2020 according to the feature information regarding the generation audience avatar G, it is possible to give the generation audience avatar G individuality, and accordingly it is possible to recreate a more realistic audience sound effect.

[0156] Note that, whereas it has been explained thus far that the sound effect generation model M3 can use various types of data included in the user data 2001 in order to generate the area sound effect 2020, types of data to be used for generation of the area sound effect 2020 are not limited particularly to any type.

[0157] For example, the sound effect generation model M3 may generate the area sound effect 2020 using only partial data in the performer data 2001A and the audience data 2001B. In particular, use of the performer data 2001A may be prioritized over use of the audience data 2001B for generation of the area sound effect 2020. This is because it is expected that audience members participating in the event are most interested in motions and sounds of performers.

[0158] In addition, use of the performer sound data 2003A and the audience sound data 2003B may be prioritized over use of the performer motion data 2002A and the audience motion data 2002B for generation of the area sound effect 2020. This is because it is expected that sounds generated by audience members are most influenced by sounds generated by performers and other audience members.

[0159] Note that, in a case where the area sound effect 2020 is generated using multiple pieces of data in data included in the user data 2001, the area sound effect 2020 may be generated weighting each piece of data. For example, each piece of data may be weighted such that the performer sound data 2003A is most reflected in the area sound effect 2020. For example, a degree of priority may be set for each of multiple pieces of data included in the user data 2001. For example, degrees of priority may be set to integer values from 0 to 10. 0 may mean the lowest degree of priority, and 10 may mean the highest degree of priority. The area sound effect 2020 may be generated and reflect data such that the influence of the data on the area sound effect 2020 increases as the degree of priority of the data increases. The degrees of priority may be set automatically on the system side in advance, or may be set or changed as appropriate on the user side. Note that data to be used for generation of the area sound effect 2020 in data included in the user data 2001 may be changed dynamically according to the network condition, the resources or remaining battery level of each piece of processing equipment (the server 10, the user terminals 20, etc.), and the like. For example, in a case where the network condition is good, the area sound effect 2020 may be generated on the basis of data with degrees of priority of 5 to 10 in data included in the user data 2001, and in a case where the network condition is not good, the area sound effect 2020 may be generated only on the basis of data with degrees of priority of 9 to 10.

[0160] Generation of the area sound effect 2020 based on various types of data included in the user data 2001 and the location of the generation avatar area Ar by the sound effect generation model M3 has been explained thus far. The sound effect generation model M3 may generate the area sound effect 2020 further using, as input data, motion data of a reflection motion 2010 generated by the motion generating section 121.

[0161] The sound effect generation model M3 may generate the area sound effect 2020 such that the area sound effect 2020 includes sounds appropriate as sounds generated by the generation audience avatar G reflecting a reflection motion 2010. For example, the sound effect generation model M3 may generate the area sound effect 2020 such that the area sound effect 2020 includes a sound of handclapping in a case where the reflection motion 2010 is a handclapping motion.

[0162] The area sound effect 2020 generated by the sound effect generation model M3 using the reflection motion 2010 is reflected in the virtual space simultaneously when the reflection motion 2010 is reflected in the generation audience avatar G. As a result, the area sound effect 2020 appropriate as a sound generated by the generation audience avatar G can be generated, and accordingly the virtual space where a large audience is excited can be produced more realistically.

[0163] Whether to or not to use reflection motions 2010 for generation of the area sound effect 2020 can be set as appropriate.

[0164] Here, the difference in time required from collection of user data 2001 to reflection in the virtual space in a case where the sound effect generation model M3 uses reflection motions 2010 for generation of area sound effects 2020 and in a case where the sound effect generation model M3 does not use reflection motions 2010 for generation of area sound effects 2020 is explained using FIG. 8 and FIG. 9.

[0165] FIG. 8 is a figure for explaining time required from collection of user data 2001 to reflection in the virtual space in a case where reflection motions 2010 are used for generation of area sound effects 2020. FIG. 9 is a figure for explaining time required from collection of user data 2001 to reflection in the virtual space in a case where reflection motions 2010 are not used for generation of area sound effects 2020.

[0166] FIG. 8 and FIG. 9 depict, in a time series, processes from collection of the user data 2001 from a user terminal 20 in predetermined time to the current time to reflection of a reflection motion 2010 and an area sound effect 2020 generated on the basis of the user data 2001 in the virtual space.

[0167] In a case where reflection motions 2010 are used for generation of area sound effects 2020, as depicted in FIG. 8, first, the user data 2001 is collected until time T0 (user data collection 1a). Next, a reflection motion 2010 is generated on the basis of the user data 2001 by time T1 (reflection motion generation 1a).

[0168] Next, an area sound effect 2020 is generated on the basis of the user data 2001 and the reflection motion 2010 by time T2 (area sound effect generation 1a). Then, the generated reflection motion 2010 and area sound effect 2020 are reflected in the virtual space by time T3 (reflection 1a).

[0169] Note that, in parallel with the process described above, the user data 2001 is collected newly until time T1 (user data collection 2a), a reflection motion 2010 is generated using the user data 2001 (reflection motion generation 2a), an area sound effect 2020 is generated on the basis of the user data 2001 and the reflection motion 2010 (area sound effect generation 2a), and the reflection motion 2010 and the area sound effect 2020 are reflected in the virtual space (reflection 2a). In such a manner, the process from collection of the user data 2001 to reflection of a reflection motion 2010 and an area sound effect 2020 in the virtual space is performed consecutively.

[0170] On the other hand, in a case where reflection motions 2010 are not used for generation of area sound effects 2020, as depicted in FIG. 9, first, the user data 2001 is collected until time T0 as in the example depicted in FIG. 8 (user data collection 1b).

[0171] Next, a reflection motion 2010 is generated on the basis of the user data 2001 by time T1 (reflection motion generation 1b). In addition, in parallel with the process of the reflection motion generation 1b, an area sound effect 2020 is generated on the basis of the user data 2001 by time T2 (area sound effect generation 1b). That is, generation of a reflection motion 2010 and generation of an area sound effect 2020 are performed simultaneously in the example depicted in FIG. 9. Hereinbelow, simultaneous generation of a reflection motion 2010 and an area sound effect 2020 is also referred to as "simultaneous generation." Note that "simultaneous" in the present specification also include substantially simultaneous. In addition, "simultaneous generation" may mean that timings of generation of a reflection motion 2010 and generation of an area sound effect 2020 overlap at least partially.

[0172] Then, the generated reflection motion 2010 and area sound effect 2020 are reflected in the virtual space by time T2 (reflection 1b). In parallel with the process described above, the user data 2001 is collected newly until time T1, and a reflection motion 2010 and an area sound effect 2020 generated using the user data 2001 are reflected in the virtual space (user data collection 2b to reflection 2b). In such a manner, reflection motions 2010 and area sound effects 2020 are reflected in the virtual space consecutively.

[0173] As has been explained thus far, in the example depicted in FIG. 9, the reflection motion 2010 and the area sound effect 2020 generated on the basis of the user data 2001 collected until time T0 are reflected by time T2. On the other hand, in the example depicted in FIG. 8, the reflection motion 2010 and the area sound effect 2020 generated on the basis of the user data 2001 collected until the same time T0 are reflected by time T3. In such a manner, in a case where an area sound effect 2020 is generated without using a reflection motion 2010, time from collection of the user data 2001 to reflection in the virtual space is shortened. In addition, in a case where an area sound effect 2020 is generated without using a reflection motion 2010, the amount of calculation by the server 10 also decreases.

[0174] Whether to or not to use reflection motions 2010 for generation of area sound effects 2020 may be set by the provider of the virtual space, or may be decided on the basis of the processing load of the server 10. For example, the processing load of the server 10 may be represented by the amount of used memory or power usage of a GPU or a CPU of the server 10, or the like. For example, the sound effect generating section 122 may generate area sound effects 2020 without using reflection motions 2010 in a case where the processing load of the server 10 is greater than a threshold.

[0175] Generation of area sound effects 2020 using sound effect generation models M3 by the sound effect generating section 122 has been explained thus far. It should be noted that, as long as the sound effect generation model M3 can output area sound effects 2020 that are influenced by performer avatars P and operated audience avatars O at the event, and are appropriate for the situation of the event, the sound effect generation model M3 is not limited to examples explained thus far, and may be trained in any manner. Note that area sound effects 2020 generated by the sound effect generating section 122 may be used for training of a machine learning model.

[0176] Note that the sound effect generating section 122 may generate an area sound effect 2020 to be reflected in each of multiple generation avatar areas Ar. In this case, the sound effect generation model M3 may output multiple area sound effects 2020 to be reflected in the multiple generation avatar areas Ar using, as input data, data representing the locations of the generation avatar areas Ar. In addition, the sound effect generation model M3 may output an area sound effect 2020 corresponding to each piece of data representing the locations of the multiple generation avatar areas Ar using the piece of the data as input data.

[0177] In addition, whereas area sound effects 2020 are generated using the machine learning model in the example explained using FIG. 7 here, a machine learning model may not be used for generation of an area sound effect 2020. For example, the sound effect generating section 122 may generate area sound effects 2020 on the basis of predetermined rules about user data 2001 and area sound effects 2020. It should be noted that, by with use of machine learning models for generation of area sound effects 2020, it is possible to generate sound effects more appropriate for the situation of the event, without complicated designing by the provider of the virtual space.

[0178] (Virtual Space Generating Section 123) The virtual space generating section 123 generates the virtual space. Specifically, as generation of the virtual space, the virtual space generating section 123 performs generation, acquisition, addition, updating, and the like of various types of data included in content data. For example, the various types of data are data for rendering virtual space videos, data for reproducing virtual space sound effects, and the like.

[0179] The data for rendering virtual space videos includes data related to a user avatar of each user to be arranged in the virtual space, and data related to generation audience avatars G. For example, the data related to user avatars is user motion data to be reflected in the user avatars, appearance data of the user avatars, user avatar coordinates, and the like.

[0180] For example, the data related to generation audience avatars G is data of reflection motions 2010 to be reflected in the generation audience avatars G, appearance data of the generation audience avatars G, generation avatar coordinates 1001, and the like.

[0181] The appearance data of generation audience avatars G may be generated by the provider of the virtual space presetting face parts, cloths to be worn by the bodies, and the like in advance, or may be generated so as to fit in with user avatars of users participating in the event on the basis of the user data 2001.

[0182] At the start of the event, the generation avatar coordinates 1001 may be set by the provider of the virtual space in advance. Then, the generation avatar coordinates 1001 may move according to reflection motions 2010 that are generated consecutively while the event is being held. In addition, throughout the event, the generation avatar coordinates 1001 may be fixed to coordinates set by the provider of the virtual space in advance. In addition, the virtual space generating section 123 may extract locations where there are few operated audience avatars O, and set the locations as the generation avatar coordinates 1001.

[0183] Note that the virtual space generating section 123 may set the number of generation audience avatars G to be arranged in the virtual space. For example, the number of generation audience avatars G to be arranged in the virtual space may be adjusted as appropriate while the event is being held, according to the processing load of the server 10.

[0184] For example, the data for reproducing virtual space sound effects includes performer sound data 2003A, audience sound data 2003B, and sound data of area sound effects 2020.

[0185] Note that the virtual space generating section 123 may perform control to generate a virtual space video and a virtual space sound effect from each user viewpoint, transmit them as content data to a user terminal 20, and cause the user terminal 20 to reproduce them.

[0186] The virtual space generating section 123 controls the communicating section 110 to transmit content data to user terminals 20 as appropriate. For example, the virtual space generating section 123 may control the communicating section 110 to transmit, to user terminals 20, content data of the virtual space of each one frame as update information regarding the virtual space.

[0187] (Storage Section 130) The storage section 130 is realized using a ROM that stores programs, computation parameters, and the like to be used for processes performed by the control section 120, and a RAM that temporarily stores parameters and the like that change as appropriate. For example, the storage section 130 stores user attribute information 131 about each user obtained from a user terminal 20.

[0188] <3. Operation Processing Example> Next, an example of the procedure of an operation process performed by the server 10 according to the present embodiment is explained using FIG. 10. FIG. 10 is a flowchart depicting an example of the procedure of an operation process performed by the server 10 according to the present embodiment.

[0189] First, the communicating section 110 acquires user data 2001 of each user from a user terminal 20 (S101). Next, the control section 120 determines whether or not a setting for simultaneous generation of reflection motions 2010 and area sound effects 2020 is turned on (S102).

[0190] In a case where the setting for simultaneous generation is turned on (S102 / YES), the process proceeds to S108. In a case where the setting for simultaneous generation is turned off (S102 / NO), the control section 120 determines whether or not the processing load of the server 10 is equal to or greater than a first threshold (S103). The first threshold is a threshold set for deciding whether to or not to execute the simultaneous generation of reflection motions 2010 and area sound effects 2020.

[0191] In a case where the processing load of the server 10 is smaller than the first threshold (S103 / NO), the control section 120 does not execute the simultaneous generation, and generates reflection motions 2010 and area sound effects 2020 sequentially, that is, first causes the motion generating section 121 to generate a reflection motion 2010 to be reflected in each generation audience avatar G (S104). Next, the sound effect generating section 122 generates area sound effects 2020 using the reflection motions 2010 generated by the motion generating section 121 (S105).

[0192] On the other hand, in a case where the processing load of the server 10 is equal to or greater than the first threshold (S103 / YES), the control section 120 further determines whether or not the processing load of the server 10 is equal to or greater than a second threshold (S106). The second threshold is a threshold set for deciding whether to or not to sample audience data 2001B in the user data 2001 of the users participating in the event, when reflection motions 2010 and area sound effects 2020 are to be generated. The second threshold is a threshold different from the first threshold, and set to a value greater than the first threshold.

[0193] In a case where the processing load of the server 10 is equal to or greater than the second threshold (S106 / YES), the control section 120 samples partial audience data 2001B to be used for generation of reflection motions 2010 and area sound effects 2020 out of multiple pieces of the audience data 2001B (S107). Next, on the basis of the sampled audience data 2001B, the motion generating section 121 and the sound effect generating section 122 simultaneously generate reflection motions 2010 and area sound effects 2020 (i.e., executes a process of generating each piece of data in parallel with each other) (S108). By executing the process of generating each piece of data using the sampled audience data 2001B here, the processing load of the server 10 is reduced.

[0194] On the other hand, in a case where the processing load of the server 10 is smaller than the second threshold (S106 / NO), each of reflection motions 2010 and area sound effects 2020 are generated simultaneously by the motion generating section 121 and the sound effect generating section 122 (S108). At this time, the motion generating section 121 and the sound effect generating section 122 may generate the reflection motions 2010 and the area sound effects 2020, respectively, by inputting all the pieces of the audience data 2001B to machine learning models.

[0195] The virtual space generating section 123 generates the virtual space by generating content data including data of the generated reflection motions 2010 and area sound effects 2020. The communicating section 110 outputs the generated content data to user terminals 20 (S109).

[0196] Then, the virtual space generating section 123 determines whether or not to end the generation of the virtual space (S110). The generation of the virtual space may be ended on the basis of end operation by the provider of the virtual space. In a case where the virtual space generating section 123 does not end the generation of the virtual space (S110 / NO), the process returns to S101, and the virtual space is kept being generated consecutively by repeating the processes of S101 to S109. On the other hand, in a case where the generation of the virtual space is to be ended (S110 / YES), the process ends.

[0197] <4. Modification Examples> Next, modification examples according to the present embodiment are explained.

[0198] (First Modification Example) First, in a first modification example to be explained, each generation audience avatar G is rendered such that the rendering processing load for some generation audience avatars G is reduced.

[0199] In the first modification example, the virtual space generating section 123 generates the virtual space such that the rendering processing load for some generation audience avatars G is reduced. Here, an example of the virtual space generated by the virtual space generating section 123 in the first modification example is explained with reference to FIG. 11. FIG. 11 is a figure for explaining an example of the virtual space generated by the virtual space generating section 123 in the first modification example. FIG. 11 depicts a virtual space video 215. The virtual space video 215 is a virtual space video rendered by the server 10 or a user terminal 20 on the basis of content data of the virtual space generated by the virtual space generating section 123.

[0200] As represented by the virtual space video 215, the virtual concert venue V has locations with different brightness, in some cases. For example, the audience area A1 is bright as compared with the other audience areas A2 to A4. Generation audience avatars G rendered in the audience areas A with lower brightness are difficult to be visually recognized as compared with generation audience avatars G rendered in audience areas A with higher brightness.

[0201] In view of this, the virtual space generating section 123 determines the brightness of each audience area A in the virtual space. Then, the virtual space generating section 123 generates the virtual space such that the rendering processing load for generation audience avatars G that are in audience areas A with brightness lower than a threshold is reduced.

[0202] For example, the virtual space generating section 123 may generate content data such that the appearances of generation audience avatars G that are in audience areas A with brightness lower than the threshold are rendered instead with simple appearances. For example, the virtual space generating section 123 may generate content data such that the appearances of the generation audience avatars G that are in the audience areas A are rendered instead with silhouettes, with the entire generation audience avatars G being painted completely with a single color such as black or gray.

[0203] In addition, the virtual space generating section 123 may generate content data such that reflection motions 2010 generated for generation audience avatars G that are in audience areas A with brightness lower than the threshold are not reflected in the generation audience avatar G. For example, the virtual space generating section 123 may generate content data such that the generation audience avatars G that are in the audience areas A stand stock-still.

[0204] It is assumed that, in the example depicted in FIG. 11, the virtual space generating section 123 determines that the brightness of the audience area A1 is equal to or greater than the threshold, and determines that the brightness of the audience areas A2 to A4 is lower than the threshold. In this case, for example, generation audience avatars G1 and G2 that are in the audience area A1 are caused to reflect appearances and reflection motions 2010. Hereinbelow, generating generation audience avatars G such that the generation audience avatars G are caused to reflect appearances and reflection motions 2010 is also called "normal generation."

[0205] In contrast, a generation audience avatar G3 and the like that are in the audience area A2 are generated such that changes are made to the generation audience avatar G3 and the like by giving them silhouette appearances instead, and causing them not to reflect reflection motions 2010. Hereinbelow, generating generation audience avatars G in a changed manner such that the rendering processing load is reduced is also called "changed generation." By performing changed generation in such a manner, it is possible to reduce the rendering processing load for the virtual space without making users viewing and listening to virtual space videos feel a sense of discomfort.

[0206] Note that whether to or not to perform changed generation according to the brightness of the virtual space may be decided on the basis of whether or not a setting for performing changed generation is turned on by the provider of the virtual space in advance.

[0207] The brightness of each audience area A can change due to a change of the illumination environment of the virtual concert venue V. Accordingly, the virtual space generating section 123 may continuously (e.g., for each frame) perform determination of the brightness of each audience area A in the virtual space.

[0208] (Second Modification Example) In a second modification example explained from here, each area sound effect 2020 is output with volume adjustment.

[0209] The virtual space generating section 123 according to the second modification example decides the volumes of area sound effects 2020 for each audience area A according to the degree of excitement of the event. For example, the degree of excitement of the event may be sensed on the basis of the intensities of motions represented by audience motion data 2002B, or the volumes of sounds represented by audience sound data 2003B. For example, the degree of excitement of the event may be sensed as either of two levels, "low" and "high." In addition, the degree of excitement of the event may be represented by numerical values.

[0210] For example, the virtual space generating section 123 may generate content data such that, as the degree of excitement increases, the volumes of area sound effects 2020 in each audience area A increase.

[0211] In addition, the virtual space generating section 123 may generate content data such that area sound effects 2020 are not reflected in some audience areas A in a case where the degree of excitement is low.

[0212] FIG. 12 is a figure for explaining examples of reflection of area sound effects 2020 in each of multiple audience areas A in a case where the degree of excitement is "low" and in a case where the degree of excitement is "high." As depicted on the left side in FIG. 12, in a case where the degree of excitement is "low," content data may be generated such that area sound effects 2020 are reflected only in audience areas A5, A8, and A9 in multiple audience areas A5 to A10. Note that, whereas it is explained as an example here that area sound effects 2020 are reflected only in the audience areas A5, A8, and A9, it is sufficient if audience areas in which area sound effects 2020 are reflected are some audience areas A, and are not limited to the example depicted on the left side in FIG. 12. For example, as such some audience areas A, the virtual space generating section 123 may choose audience areas A which are not adjacent to each other, may choose only the middle audience area A, may choose only audience areas A close to the stage, or may choose only far audience areas A. In addition, audience areas A designated by a user may be selected.

[0213] In addition, as depicted on the right side in FIG. 12, in a case where the degree of excitement is "high," content data may be generated such that area sound effects 2020 are reflected in all of the multiple audience areas A5 to A10. Whereas it is assumed here as an example that area sound effects 2020 are reflected in all the audience areas A, this is not the sole example. Area sound effects 2020 may be reflected in more than half of the audience areas A according to the degree of excitement, or may be reflected in 80% of the audience areas A.

[0214] Note that, in a case where the degree of excitement is "high," area sound effects 2020 with volumes larger than those in a case where the degree of excitement is "low" may be reflected. For example, in the examples depicted in FIG. 12, in a case where the degree of excitement is "high," area sound effects 2020 with volumes larger than those in a case where the degree of excitement is "low" may be reflected in the audience areas A5, A8, and A9, which are audience areas A where area sound effects 2020 are reflected both in the case where the degree of excitement is "low" and in the case where the degree of excitement is "high."

[0215] Adjustment of the volume of each audience area A results in superimposition of area sound effects 2020 with volumes which increase as the degree of excitement increases, and accordingly it can be expected that a further excitement is produced in a case where the degree of excitement in the event is increasing. On the other hand, it is expected that a case where the degree of excitement is low is a case where audience users want to listen to a performance by performer users quietly, or the like. Since area sound effects 2020 with small volumes are superimposed in a case where the degree of excitement is low according to the configuration described above, it is possible to produce the virtual space according to such a need.

[0216] Note that whether to or not to adjust the volume of each area sound effect 2020 may be decided on the basis of whether or not a setting for performing adjustment of volumes is turned on by the provider of the virtual space in advance.

[0217] (Operation Processing Example of First Modification Example and Second Modification Example) Next, an example of the procedure of an operation process performed by the server 10 in a case where the first modification example and the second modification example are applied is explained using FIG. 13. FIG. 13 is a flowchart depicting the example of the procedure of an operation process performed by the server 10 in a case where the first modification example and the second modification example are applied.

[0218] The process depicted in FIG. 13 is executed subsequent to the process of S105 or S108 depicted in FIG. 10.

[0219] First, the virtual space generating section 123 determines whether or not an adjustment setting that is set by the provider of the virtual space and represents that changed generation of generation audience avatars G is to be performed is turned on (S201).

[0220] In a case where the setting for adjustment of generation audience avatars G is turned off (S201 / NO), the virtual space generating section 123 decides to normally generate all generation audience avatars G (S202). That is, the virtual space generating section 123 generates content data such that appearances and reflection motions 2010 are reflected in all the generation audience avatars G.

[0221] In a case where the setting for adjustment of generation audience avatars G is turned on (S201 / YES), the virtual space generating section 123 determines the brightness of each audience area A in the virtual space (S203). Then, the virtual space generating section 123 decides generation audience avatars G to be generated by changed generation (S204). Specifically, the virtual space generating section 123 may decide that a generation audience avatar G that is in an audience area A with brightness which is lower than a threshold as a generation audience avatar G to be generated by changed generation.

[0222] Next, the virtual space generating section 123 determines whether or not an adjustment setting that is set by the provider of the virtual space and represents that the volumes of area sound effects 2020 are to be adjusted is turned on (S205).

[0223] In a case where the setting for volume adjustment of area sound effects 2020 is turned off (S205 / NO), the virtual space generating section 123 decides to superimpose area sound effects 2020 of all audience areas A without volume adjustment (S206).

[0224] In a case where the setting for volume adjustment of area sound effects 2020 is turned on (S205 / YES), the virtual space generating section 123 determines the degree of excitement of the virtual space (S207). Then, the virtual space generating section 123 decides the volumes of area sound effects 2020 for each audience area A (S208). For example, the virtual space generating section 123 may decide the volume of each area sound effect 2020 such that the volume increases as the degree of excitement of the virtual space increases.

[0225] Next, the process proceeds to S109 depicted in FIG. 10. At S109, the virtual space may be generated such that the appearances of the generation audience avatars G decided as generation audience avatars G to be generated by changed generation at S204 are replaced with silhouettes, and the generation audience avatars G do not reflect reflection motions 2010. In addition, at S109, the virtual space is generated such that the volume of each audience area A decided at S209 is reflected.

[0226] <5. Hardware Configuration> An embodiment and modification examples according to the present disclosure have been explained thus far. Next, a hardware configuration example of the server 10 and the user terminals 20 according to an embodiment of the present disclosure is explained with reference to FIG. 14.

[0227] Processes performed by the server 10 and the user terminals 20 mentioned above can be realized by one or more information processing apparatuses. FIG. 14 is a block diagram depicting a hardware configuration example of an information processing apparatus 900 that realizes the server 10 and the user terminals 20 according to an embodiment of the present disclosure. Note that the information processing apparatus 900 need not necessarily have the entire hardware configuration depicted in FIG. 14. In addition, part of the hardware configuration depicted in FIG. 14 may not be present in the server 10 or the user terminals 20.

[0228] As depicted in FIG. 14, the information processing apparatus 900 includes a CPU 901, a ROM (Read Only Memory) 903, and a RAM 905. In addition, the information processing apparatus 900 may include a host bus 907, a bridge 909, an external bus 911, an interface 913, an input apparatus 915, an output apparatus 917, a storage apparatus 919, a drive 921, a connection port 923, and a communication apparatus 925. The information processing apparatus 900 may have a processing circuit like one called a GPU (Graphics Processing Unit), a DSP (Digital Signal Processor), or an ASIC (Application Specific Integrated Circuit), instead of or along with the CPU 901.

[0229] The CPU 901 functions as a computation processing apparatus and a control apparatus, and controls the whole or part of operation in the information processing apparatus 900 according to various types of programs recorded on the ROM 903, the RAM 905, the storage apparatus 919, or a removable recording medium 927. The ROM 903 stores programs, computation parameters, and the like used by the CPU 901. The RAM 905 temporarily stores programs to be used in execution by the CPU 901, and parameters or the like that change as appropriate during the execution. The CPU 901, the ROM 903, and the RAM 905 are interconnected by the host bus 907 including an internal bus such as a CPU bus. Furthermore, the host bus 907 is connected to the external bus 911 such as a PCI (Peripheral Component Interconnect / Interface) bus via the bridge 909.

[0230] For example, the input apparatus 915 is an apparatus such as a button to be operated by a user. The input apparatus 915 may include a mouse, a keyboard, a touch panel, a switch, a lever, and the like. In addition, the input apparatus 915 may include a microphone that senses sounds of a user. For example, the input apparatus 915 may be a remote control apparatus using infrared rays or other radio waves or may be externally connected equipment 929 such as a mobile phone that supports operation of the information processing apparatus 900. The input apparatus 915 includes an input control circuit that generates an input signal on the basis of information input by a user, and outputs the input signal to the CPU 901. The user inputs various types of data or gives an instruction about process operations to the information processing apparatus 900 by operating the input apparatus 915.

[0231] In addition, the input apparatus 915 may include an image-capturing apparatus and sensors. For example, the image-capturing apparatus is an apparatus that captures images of a real space using image-capturing elements such as a CCD (Charge Coupled Device) or a CMOS (Complementary Metal Oxide Semiconductor) and various types of members such as lenses for controlling image-formation of image-capturing-subject images onto the image-capturing elements, and generates captured images. The image-capturing apparatus may be one that captures still images or may be one that captures videos.

[0232] For example, the sensors are various types of sensors such as a distance measurement sensor, an acceleration sensor, a gyro sensor, a geomagnetic sensor, a vibration sensor, an optical sensor, or a sound sensor. For example, the sensors acquire information regarding the state of the information processing apparatus 900 itself such as the posture of the housing of the information processing apparatus 900, and information regarding the surrounding environment of the information processing apparatus 900 such as the brightness and noises around the information processing apparatus 900. In addition, the sensors may include a GPS sensor that receives GPS (Global Positioning System) signals, and measures the latitude, longitude, and altitude of the apparatus.

[0233] The output apparatus 917 includes apparatuses that can visually or auditorily notify acquired information to a user. For example, the output apparatus 917 can be a display apparatus such as an LCD (Liquid Crystal Display) or an organic EL (Electro-Luminescence) display, a sound output apparatus such as a speaker or headphones, and the like. In addition, the output apparatus 917 may include a PDP (Plasma Display Panel), a projector, a hologram, a printer apparatus, and the like. The output apparatus 917 outputs results obtained by processes performed by the information processing apparatus 900 as videos of text, images, or the like or as auditory information such as sounds or sound effects, and so on. In addition, the output apparatus 917 may include an illuminating apparatus or the like that makes the surrounding space bright.

[0234] The storage apparatus 919 is an apparatus for data storage configured as an example of a storage section of the information processing apparatus 900. For example, the storage apparatus 919 includes a magnetic storage device, a semiconductor storage device, an optical storage device, a magneto-optical storage device, or the like such as an HDD (Hard Disk Drive). This storage apparatus 919 stores programs to be executed by the CPU 901, various types of data, various types of data acquired from the outside, and the like.

[0235] The drive 921 is a reader / writer for the removable recording medium 927 such as a magnetic disc, an optical disc, a magneto-optical disc, or a semiconductor memory, and is built in or externally attached to the information processing apparatus 900. The drive 921 reads out information recorded on the attached removable recording medium 927, and outputs the information to the RAM 905. In addition, the drive 921 writes records on the attached removable recording medium 927.

[0236] The connection port 923 is a port for directly connecting equipment to the information processing apparatus 900. For example, the connection port 923 can be a USB (Universal Serial Bus) port, an IEEE 1394 port, an SCSI (Small Computer System Interface) port, or the like. In addition, the connection port 923 may be an RS-232C port, an optical audio terminal, an HDMI (registered trademark) (High-Definition Multimedia Interface) port, or the like. By connecting the externally connected equipment 929 to the connection port 923, various types of data can be exchanged between the information processing apparatus 900 and the externally connected equipment 929.

[0237] For example, the communication apparatus 925 is a communication interface including a communication device or the like for connection to the network 30. For example, the communication apparatus 925 can be a communication card or the like for a wired or wireless LAN (Local Area Network), Bluetooth (registered trademark), Wi-Fi (registered trademark), or WUSB (Wireless USB). In addition, the communication apparatus 925 may be a router for optical communication, a router for ADSL (Asymmetric Digital Subscriber Line), a modem for various types of communication, or the like. For example, the communication apparatus 925 transmits and receives signals and the like to and from the Internet and other communication equipment using a predetermined protocol such as TCP / IP. In addition, the network 30 connected to the communication apparatus 925 is a network connected through a cable or wirelessly, and is the Internet, a home LAN, infrared communication, radio wave communication, satellite communication, or the like, for example.

[0238] <6. Supplementary Notes> Whereas a suitable embodiment of the present disclosure has been explained in detail thus far with reference to the attached figures, the technical scope of the present disclosure is not limited to the example. For those with ordinary knowledge in the technical field of the present disclosure, it is obvious that various types of modified examples or corrected examples can be conceived of within the scope of the technical idea described in claims, and it is understood that those various types of modified examples or corrected examples certainly belong to the technical scope of the present disclosure.

[0239] For example, the user data 2001 explained in an embodiment described above may further include data different from the data explained in the description above. For example, the user data 2001 may include bio-information such as the heart rate or body temperature of each user.

[0240] In addition, a computer program for causing hardware such as the CPU, the ROM, or the RAM built in the server 10 or the user terminals 20 mentioned above to exhibit the functions of the server 10 or the user terminals 20 can also be created. In addition, a computer-readable storage medium on which the computer program is stored is also provided.

[0241] In addition, the advantages described in the present specification are presented merely for explanation or illustration, and not for limitation. That is, the technology according to the present disclosure can exhibit other advantages that are obvious for those skilled in the art from the description of the present specification, along with the advantages described above or instead of the advantages described above.

[0242] Note that configurations like the ones below also belong to the technical scope of the present disclosure. (1) An information processing apparatus including: circuitry configured to generate a virtual space, and initiate display of the virtual space, in which the virtual space includes a first avatar operated by a user and a second avatar different from the first avatar, and in which the circuitry is further configured to generate a reflection motion to be reflected in the second avatar, based on user data obtained from the user. (2) The information processing apparatus according to (1) above, in which the circuitry is further configured to generate a sound effect to be reflected in a subspace where the second avatar is located among one or more subspaces in the virtual space, by using at least one of the user data or the reflection motion. (3) The information processing apparatus according to (1) or (2) above, in which the user includes a performer in the virtual space, the user data includes sound data of music used at the event, and the circuitry is configured to generate the reflection motion, based on the sound data. (4) The information processing apparatus according to any one of (1) to (3) above, in which the user data includes at least one of motion data of the performer user or sound data of a sound generated by the performer user. (5) The information processing apparatus according to any one of (1) to (4) above, in which the user includes an audience user who operates a first avatar and is an audience member of the event, and the user data includes at least any one of motion data of the audience user and sound data of a sound generated by the audience user. (6) The information processing apparatus according to any one of (1) to (5) above, in which the user data includes a position of the first avatar in the virtual space, and the circuitry is configured to generate the reflection motion, further based on a position of the second avatar in the virtual space. (7) The information processing apparatus according to any one of (1) to (6) above, in which the circuitry is configured to generate the reflection motion based on multiple pieces of the user data obtained from a plurality of users corresponding to a plurality of first avatars in the virtual space, and generate the reflection motion based on a position of each respective first avatar in the virtual space with respect to the position of the second avatar, such that an influence of respective user data of each respective user corresponding to each respective first avatar increases as the position of the respective avatar gets closer to the position of the second avatar. (8) The information processing apparatus according to any one of (1) to (7) above, in which the user data includes a position of the first avatar in the virtual space, and the circuitry is configured to generate the sound effect, based on a relation between the position of the first avatar and a subspace where the second avatar is located in the virtual space. (9) The information processing apparatus according to any one of (1) to (8) above, in which the circuitry is further configured to acquire the reflection motion by inputting, to a second machine learning model, at least one pre-seed motion output by a first machine learning model using, as input data, sound data of music used in the virtual space, and in which the circuitry is further configured to acquire the reflection motion by further inputting to the second machine learning model at least one of the user data, virtual space information generated based on the user data, or a position of the second avatar in the virtual space. (10) The information processing apparatus according to any one of (1) to (9) above, in which, in a case where a processing load of the information processing apparatus is greater than a threshold, the circuitry is configured to generate the sound effect without using the reflection motion. (11) The information processing apparatus according to any one of (1) to (10) above, in which, in a case where the processing load of the information processing apparatus is greater than a threshold, the circuitry is further configured to reduce the number of pieces of user data to be used when generating the reflection motion, as compared with a case where the processing load is smaller than the threshold. (12) The information processing apparatus according to any one of (1) to (11) above, in which the circuitry is configured to determine brightness of each subspace of a plurality of subspaces in the virtual space, and generate the virtual space such that the reflection motion is generated for the second avatar according to a brightness of a subspace where the second avatar is located. (13) The information processing apparatus according to any one of (1) to (12) above, in which the circuitry is configured to generate a respective motion of each respective second avatar of a plurality of second avatars that are located at different positions in the virtual space. (14) The information processing apparatus according to any one of (1) to (13) above, in which the circuitry is configured to generate a sound effect to be reflected in each subspace of a plurality of subspaces where a plurality of second avatars are located at different positions in the virtual space. (15) The information processing apparatus according to any one of (1) to (14) above, in which the circuitry is configured to acquire a degree of excitement in the virtual space, and determine a volume of the sound effect to be reflected in each subspace according to the degree of excitement. (16) The information processing apparatus according to any one of (1) to (15) above, in which the user data includes information regarding at least one of an age, a gender, or a performance-related preference of the user. (17) An information processing apparatus including: circuitry configured to generate a virtual space, and initiate display of the generated virtual space, in which the displayed virtual space includes a first avatar operated by a user and a second avatar different from the first avatar, and in which the circuitry is further configured to generate a sound effect to be reflected in a subspace where the second avatar is located among one or more subspaces in the virtual space, based on user data obtained from the user. (18) The information processing apparatus according to (17) above, in which the user data includes a position of the first avatar in the virtual space, and the circuitry is configured to generate the sound effect based on multiple pieces of the user data obtained from a plurality of users corresponding to a plurality of first avatars in the virtual space, and generate the sound effect based on a distance between each respective first avatar and a location of the subspace of the second avatar, such that an influence of respective user data corresponding to each respective first avatar increases as the distance between the respective first avatar and the location of the subspace of the second avatar decreases. (19) A method executed by a processor, the method including: generating a virtual space, the virtual space including a first avatar operated by a user and a second avatar different from the first avatar; displaying the generated virtual space; and generating a reflection motion to be reflected in the second avatar, based on user data obtained from the user. (20) A non-transitory computer-readable storage medium having embodied thereon a program which when executed by a computer causes the computer to execute a method, the method including: generating a virtual space, the virtual space including a first avatar operated by a user and a second avatar different from the first avatar; displaying the generated virtual space; and generating a reflection motion to be reflected in the second avatar, based on user data obtained from the user.

[0243] 10: Server 110: Communicating section 120: Control section 121: Motion generating section 122: Sound effect generating section 123: Virtual space generating section 130: Storage section 131: User attribute information 20: User terminal 21: HMD 30: Network

Claims

1. An information processing apparatus comprising: circuitry configured to generate a virtual space, and initiate display of the virtual space, wherein the virtual space includes a first avatar operated by a user and a second avatar different from the first avatar, and wherein the circuitry is further configured to generate a reflection motion to be reflected in the second avatar, based on user data obtained from the user.

2. The information processing apparatus according to claim 1, wherein the circuitry is further configured to generate a sound effect to be reflected in a subspace where the second avatar is located among one or more subspaces in the virtual space, by using at least one of the user data or the reflection motion.

3. The information processing apparatus according to claim 1, wherein the user includes a performer in the virtual space, wherein the user data includes sound data of music used at the event, and wherein the circuitry is configured to generate the reflection motion, based on the sound data.

4. The information processing apparatus according to claim 3, wherein the user data includes at least one of motion data of the performer user or sound data of a sound generated by the performer user.

5. The information processing apparatus according to claim 3, wherein the user includes an audience user who operates a first avatar and is an audience member of the event, and wherein the user data includes at least any one of motion data of the audience user and sound data of a sound generated by the audience user.

6. The information processing apparatus according to claim 1, wherein the user data includes a position of the first avatar in the virtual space, and wherein the circuitry is configured to generate the reflection motion, further based on a position of the second avatar in the virtual space.

7. The information processing apparatus according to claim 6, wherein the circuitry is configured to generate the reflection motion based on multiple pieces of the user data obtained from a plurality of users corresponding to a plurality of first avatars in the virtual space, and generate the reflection motion based on a position of each respective first avatar in the virtual space with respect to the position of the second avatar, such that an influence of respective user data of each respective user corresponding to each respective first avatar increases as the position of the respective avatar gets closer to the position of the second avatar.

8. The information processing apparatus according to claim 2, wherein the user data includes a position of the first avatar in the virtual space, and wherein the circuitry is configured to generate the sound effect, based on a relation between the position of the first avatar and a subspace where the second avatar is located in the virtual space.

9. The information processing apparatus according to claim 1, wherein the circuitry is further configured to acquire the reflection motion by inputting, to a second machine learning model, at least one pre-seed motion output by a first machine learning model using, as input data, sound data of music used in the virtual space, and wherein the circuitry is further configured to acquire the reflection motion by further inputting to the second machine learning model at least one of the user data, virtual space information generated based on the user data, or a position of the second avatar in the virtual space.

10. The information processing apparatus according to claim 2, wherein, in a case where a processing load of the information processing apparatus is greater than a threshold, the circuitry is configured to generate the sound effect without using the reflection motion.

11. The information processing apparatus according to claim 1, wherein, in a case where the processing load of the information processing apparatus is greater than a threshold, the circuitry is further configured to reduce a number of pieces of user data to be used when generating the reflection motion, as compared with a case where the processing load is smaller than the threshold.

12. The information processing apparatus according to claim 1, wherein the circuitry is configured to determine brightness of each subspace of a plurality of subspaces in the virtual space, and generate the virtual space such that the reflection motion is generated for the second avatar according to a brightness of a subspace where the second avatar is located.

13. The information processing apparatus according to claim 1, wherein the circuitry is configured to generate a respective motion of each respective second avatar of a plurality of second avatars that are located at different positions in the virtual space.

14. The information processing apparatus according to claim 2, wherein the circuitry is configured to generate a sound effect to be reflected in each subspace of a plurality of subspaces where a plurality of second avatars are located at different positions in the virtual space.

15. The information processing apparatus according to claim 14, wherein the circuitry is configured to acquire a degree of excitement in the virtual space, and determine a volume of the sound effect to be reflected in each subspace according to the degree of excitement.

16. The information processing apparatus according to claim 1, wherein the user data includes information regarding at least one of an age, a gender, or a performance-related preference of the user.

17. An information processing apparatus comprising: circuitry configured to generate a virtual space, and initiate display of the generated virtual space, wherein the displayed virtual space includes a first avatar operated by a user and a second avatar different from the first avatar, and wherein the circuitry is further configured to generate a sound effect to be reflected in a subspace where the second avatar is located among one or more subspaces in the virtual space, based on user data obtained from the user.

18. The information processing apparatus according to claim 17, wherein the user data includes a position of the first avatar in the virtual space, and wherein the circuitry is configured to generate the sound effect based on multiple pieces of the user data obtained from a plurality of users corresponding to a plurality of first avatars in the virtual space, and generate the sound effect based on a distance between each respective first avatar and a location of the subspace of the second avatar, such that an influence of respective user data corresponding to each respective first avatar increases as the distance between the respective first avatar and the location of the subspace of the second avatar decreases.

19. A method executed by a processor, the method comprising: generating a virtual space, the virtual space including a first avatar operated by a user and a second avatar different from the first avatar; displaying the generated virtual space; and generating a reflection motion to be reflected in the second avatar, based on user data obtained from the user.

20. A non-transitory computer-readable storage medium having embodied thereon a program which when executed by a computer causes the computer to execute a method, the method comprising: generating a virtual space, the virtual space including a first avatar operated by a user and a second avatar different from the first avatar; displaying the generated virtual space; and generating a reflection motion to be reflected in the second avatar, based on user data obtained from the user.

Citation Information

Patent Citations

  • Program and electronic apparatus

    JP2018007828A

  • Game system and program

    JP2018075260A

  • Program and information processing apparatus

    JP2023092332A