Device and method for providing metaverse service for exercising with avatar
Patent Information
- Application Number
- PCT/KR2024/000225
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-04-24
- Filing Date
- 2024-01-04
- Publication Date
- 2025-05-22
AI Technical Summary
Conventional fitness exercises lack realism and engagement, leading to user boredom and disinterest, as they cannot effectively replicate the sense of presence and social interaction found in natural environments, limiting their ability to maintain user interest and motivation.
A metaverse service provision device and method that reflects user motion data in virtual space, utilizing a pre-trained neural network model to predict and control avatar movements, allowing users and avatars to exercise together, thereby enhancing the sense of presence, interest, and fun.
The solution provides a more immersive and engaging exercise experience by synchronizing user and avatar movements, increasing user engagement and motivation through enhanced realism and social interaction in virtual fitness environments.
Smart Images

Figure KR2024000225_22052025_PF_FP_ABST
Abstract
Description
Device and method for providing a metaverse service that allows exercise with an avatar
[0001] The present disclosure relates to a metaverse service providing device, and more specifically, to a metaverse service providing device and method capable of providing a metaverse service in which a user's movement in a real space is reflected to an avatar in a virtual space, thereby enabling a user and an avatar to move together.
[0002] As technology advances, modern people have become more comfortable and affluent, but at the same time, diseases caused by lack of exercise have become a growing problem both personally and socially.
[0003] Accordingly, interest in exercise for health management is increasing throughout society, and facilities such as fitness centers that allow for more intensive exercise in indoor spaces that are easily accessible and easily used in busy daily lives are being expanded and broadened.
[0004] However, while these fitness exercises have the advantage of allowing people to get the amount of exercise they want at a relatively low cost compared to other exercises, they have the disadvantage of not being able to feel the real thing, so many people easily lose interest.
[0005] Recently, fitness devices that provide an environment similar to the natural environment through various video screens have been developed to stimulate people's interest. However, there are limitations in realistically implementing the realism of an actual exercise course, and there is a problem that users may feel bored because they still exercise alone.
[0006] Therefore, in the future, there is a need for the development of a metaverse service providing device that can provide users with a sense of presence, interest, and fun by reflecting the user's movement in real space to an avatar in virtual space, allowing the user and avatar to move together.
[0007] One purpose of the present disclosure to solve the above-described problems is to provide a metaverse service providing device and method that can provide a sense of presence, interest, and fun to the user by reflecting user motion data on an avatar uploaded to a virtual space of an exercise session and controlling the motion of the avatar by predicting the next movement through a pre-learned neural network model, thereby enabling the user and the avatar to exercise together.
[0008] The problems to be solved by the present disclosure are not limited to the problems mentioned above, and other problems not mentioned will be clearly understood by those skilled in the art from the description below.
[0009] The metaverse service providing device according to the present disclosure for achieving the above-described technical task comprises: a communication unit communicating with a user terminal; a database storing motion data of an avatar;And a control unit that controls the motion of an avatar uploaded to a virtual space of a metaverse, wherein the control unit requests a user-captured video to the user terminal when a request for participation in an exercise session is received from the user terminal, extracts user motion data from the user-captured video when the user-captured video is received from the user terminal, reflects the user motion data to the avatar uploaded to the virtual space of the exercise session, inputs user motion data corresponding to the current motion of the user into a pre-learned neural network model to predict the next motion of the user, and controls the motion of the avatar based on the prediction result, and when the control unit extracts the user motion data, extracts 3D key points for joint positions from the user's body included in the user-captured video using the pre-learned neural network model, and extracts user motion data based on a rotation parameter corresponding to each joint key point, and when the control unit extracts the 3D key points, if the number of the extracted key points is less than a reference ratio or the visibility of the predicted key points is less than a reference ratio, determines that the user motion is invalid, and stops updating the user motion data or updates the user motion data. The data is predicted in advance and then replaced with other user motion data corresponding to the motion, and the control unit determines that the user motion is invalid if the motion change rate of the user motion becomes greater than the reference motion change rate while the user-captured video is received at a constant number of frames per second, and the control unit is characterized in that if the area occupied by joints corresponding to the user motion is predicted to collide with each other based on the rotation parameter, the rotation parameter of the joint is corrected.;
[0010] In addition, a method for providing a metaverse service according to the present disclosure for achieving the above-described technical task comprises the steps of: receiving a request for participation in an exercise session from the user terminal; requesting a user-captured video to the user terminal; receiving the user-captured video from the user terminal; extracting user motion data from the user-captured video; reflecting the user motion data to an avatar uploaded to a virtual space of the exercise session; inputting user motion data corresponding to the user's current motion into a pre-learned neural network model to predict the user's next motion; And a step of controlling the motion of the avatar based on the prediction result; wherein, when extracting the user motion data, the control unit extracts 3D key points for joint positions from the user's body included in the user-captured video using a pre-learned neural network model, and extracts user motion data based on a rotation parameter corresponding to each joint key point, and when extracting the 3D key points, if the number of the extracted key points is less than a reference ratio or the visibility of the predicted key points is less than a reference ratio, the control unit determines that the user motion is invalid, and stops updating the user motion data or predicts the user motion data in advance and then replaces it with other user motion data corresponding to the motion, and the control unit determines that the user motion is invalid if the motion change rate of the user motion becomes greater than the reference motion change rate while the user-captured video is received at a constant number of frames per second, and the control unit corrects the rotation parameter of the joint if it is predicted that areas occupied by joints corresponding to the user motion collide with each other based on the rotation parameter.
[0011] In addition, a computer program stored in a computer-readable recording medium for executing a method for implementing the present disclosure may be further provided.
[0012] In addition, a computer-readable recording medium recording a computer program for executing a method for implementing the present disclosure may be further provided.
[0013] As described above, according to the present disclosure, by reflecting user motion data on an avatar uploaded to a virtual space of an exercise session and controlling the motion of the avatar by predicting the next movement through a pre-learned neural network model, it is possible to provide a sense of realism, interest, and fun to the user by having the user and the avatar exercise together.
[0014] The effects of the present disclosure are not limited to the effects mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art from the description below.
[0015] Figure 1 is a schematic diagram for explaining a metaverse service providing device according to the present disclosure.
[0016] Figures 2 to 6 are flowcharts for explaining a method for providing a metaverse service according to the present disclosure.
[0017] FIGS. 7 to 19 are drawings showing screen implementations of a metaverse service providing device according to the present disclosure.
[0018] The advantages and features of the present invention, and methods for achieving them, will become clearer with reference to the embodiments described in detail below together with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below and may be implemented in various different forms. These embodiments are provided solely to ensure that the present disclosure is complete and to fully inform those skilled in the art of the scope of the present disclosure, and the present disclosure is defined solely by the scope of the claims.
[0019] The terminology used herein is for the purpose of describing embodiments only and is not intended to limit the present disclosure. In this specification, the singular also includes the plural unless specifically stated otherwise. As used herein, the terms "comprises" and / or "comprising" do not exclude the presence or addition of one or more other components in addition to the mentioned components. Like reference numerals refer to like components throughout the specification, and "and / or" includes each and any combination of one or more of the mentioned components. Although "first", "second", etc. are used to describe various components, these components are not limited by these terms. These terms are only used to distinguish one component from another. Therefore, it should be understood that a first component mentioned below may also be a second component within the technical spirit of the present disclosure.
[0020] Unless otherwise defined, all terms (including technical and scientific terms) used herein may be used in their common sense to those of ordinary skill in the art to which this disclosure pertains. Furthermore, terms defined in commonly used dictionaries are not to be interpreted ideally or excessively unless explicitly and specifically defined otherwise.
[0021] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the attached drawings.
[0022] Before proceeding, the meanings of terms used in this specification will be briefly explained. However, it should be noted that the explanation of terms is intended to aid understanding of this specification and, unless explicitly stated to limit the disclosure, they are not intended to limit the technical concepts of this disclosure.
[0023] Additionally, throughout this specification, the terms "neural network," "neural network," and "network function" may be used interchangeably. A neural network may be comprised of a set of interconnected computational units, which may generally be referred to as "nodes." These "nodes" may also be referred to as "neurons." A neural network comprises at least two or more nodes. The nodes (or neurons) constituting the neural networks may be interconnected by one or more "links."
[0024] Figure 1 is a schematic diagram for explaining a metaverse service providing device according to the present disclosure.
[0025] As illustrated in FIG. 1, the metaverse service providing device according to the present disclosure may include a server (200) of a metaverse service platform that is connected to a user terminal (100) via a network.
[0026] Here, the user terminal (100) may include both a stationary device (standing device) such as a personal computer (PC), a network TV, a hybrid broadcast broadband TV (HBBTV), a smart TV, an internet protocol TV (IPTV), and the like, and a mobile device (mobile device or handheld device) such as a smart phone, a tablet PC, a notebook, a personal digital assistant (PDA), and the like. In some cases, the user terminal (100) may further include a device that utilizes a computer chip and image recognition, and the image recognition device may be a webcam, but this is only one example and is not limited thereto.
[0027] And, the network that provides communication connection between the user terminal (100) and the server (200) includes both wired and wireless networks, and is a general term for a communication network that supports various communication standards or protocols for pairing or / and data transmission and reception between the user terminal (100) and the server (200).
[0028] These wired / wireless networks include all current or future communication networks supported by the standard, and can support one or more communication protocols for them.
[0029] These wired / wireless networks can be formed by networks for wired connections such as USB (Universal Serial Bus), CVBS (Composite Video Banking Sync), Component, S-Video (analog), DVI (Digital Visual Interface), HDMI (High Definition Multimedia Interface), RGB, and D-SUB, and communication standards or protocols therefor, and networks for wireless connections such as Bluetooth, RFID (Radio Frequency Identification), IrDA (infrared Data Association), UWB (Ultra Wideband), ZigBee, DLNA (Digital Living Network Alliance), WLAN (Wireless LAN) (Wi-Fi), Wibro (Wireless broadband), Wimax (World Interoperability for Microwave Access), HSDPA (High Speed Downlink Packet Access), LTE / LTE-A (Long Term Evolution / LTE-Advanced), and Wi-Fi Direct, and communication standards or protocols therefor.
[0030] In addition, the server (200) may include a communication unit that communicates with the user terminal (100), a database that stores motion data of the avatar, and a control unit that controls the motion of the avatar uploaded to the virtual space of the metaverse.
[0031] When a request for participation in an exercise session is received from a user terminal (100), the control unit requests a user-captured video from the user terminal (100), extracts user motion data from the user-captured video when the user-captured video is received from the user terminal (100), reflects the user motion data to an avatar uploaded to a virtual space of the exercise session, inputs motion sequence data corresponding to the user's current motion sequence into a pre-learned neural network model to predict the user's next motion, and controls the motion of the avatar based on the prediction result.
[0032] Here, when extracting motion data, the control unit can extract motion data through a depth sensor, etc., in addition to the user-captured video, or through motion tracking equipment / suit, etc.
[0033] Additionally, the control unit can control avatar motion without performing an avatar motion prediction process using a neural network model. In the present disclosure, the use of a neural network model is merely one embodiment, and avatar motion can be controlled without a neural network model.
[0034] Here, before receiving a request to participate in an exercise session, the control unit can check whether the user terminal (100) is registered as a member when an app execution request is received from the user terminal (100), and if the user terminal (100) is registered as a member, provide exercise session information to the user terminal (100).
[0035] In addition, when providing exercise session information, the control unit can extract exercise session information based on an exercise session start schedule from an exercise session list pre-stored in a database, and provide the extracted exercise session information to a user terminal (100).
[0036] For example, exercise session information may include at least one of an exercise schedule, exercise start time, exercise duration, exercise title, calories burned, exercise composition movements, URL (Uniform Resource Locator) information of the corresponding content, and information for composing an exercise result video, but this is only an example and is not limited thereto.
[0037] And, when providing exercise session information, the control unit can extract exercise session information having an exercise schedule for a predetermined period of time and sequentially provide the extracted exercise session information according to the order of the exercise schedule.
[0038] Here, when sequentially providing exercise session information, the control unit may first provide exercise session information having the earliest schedule among the extracted exercise session information, and lastly provide exercise session information having the latest schedule.
[0039] In some cases, when providing exercise session information, the control unit may extract personalized recommended exercise session information based on user information of the user terminal (100) and sequentially provide the extracted exercise session information according to the order of the exercise schedule.
[0040] For example, the control unit may extract personalized recommended exercise session information based on at least one physical condition among age, gender, medical history, and weight among the user information of the user terminal (100), but this is only one embodiment and is not limited thereto.
[0041] In another case, when providing exercise session information, the control unit may extract exercise session information based on exercise session participation history information of the user terminal (100) and sequentially provide the extracted exercise session information according to the order of the exercise schedule.
[0042] Here, the control unit can extract exercise session information based on at least one activity condition among the most recently participated exercise session information, the most frequently participated exercise session information, and the exercise session information on which the most time was spent among the exercise session participation history information of the user terminal (100).
[0043] Additionally, the control unit may provide a home screen including a schedule card with exercise session information and an avatar selected by the user when providing exercise session information.
[0044] Here, the control unit can provide exercise session information including at least one of the current or nearest next exercise session date and time, exercise session status, exercise session title, and main action button within the schedule card.
[0045] For example, the status of a workout session may include a waiting state before the workout session starts or an in-progress state after the workout session starts.
[0046] As another example, the main action button may have different text, color, and actions depending on at least one of the following: reservation status, waiting status, waiting and full status, in progress and under capacity status, and in progress and full status.
[0047] Additionally, the control unit can provide menu items, a background image of the virtual space, and a welcome message within the home screen.
[0048] Here, the welcome message may have different text content depending on at least one of the following situations: reservation status, waiting status, waiting and full capacity status, in progress and under capacity status, and in progress and full capacity status.
[0049] In addition, when confirming whether a user terminal (100) has registered as a member, if the user terminal (100) has not registered as a member, an onboarding screen is provided. If an app launch request is received through the onboarding screen, an app usage guide window is provided. If user confirmation is received through the app usage guide window, an avatar setting screen is provided. The avatar and user information set by the user through the avatar setting screen can be stored in a database. In some cases, if membership registration is confirmed, the existing avatar setting process and nickname setting process stored in the database can be omitted.
[0050] Here, the control unit can provide a home screen including a schedule card with exercise session information and an avatar selected by the user when the avatar is set through the avatar setting screen.
[0051] For example, the control unit may provide an onboarding screen only once after the app is installed.
[0052] As another example, when providing an onboarding screen, the control unit may provide an onboarding screen that includes a menu button, a background image, and a start button, and receive an app start request through the start button.
[0053] Additionally, the control unit may provide the app usage guide window only once, in the form of a pop-up window that provides instructions on how to use the app. In some cases, the control unit may provide the app usage guide window again after the onboarding process or according to user input selected from a menu.
[0054] For example, when providing an app usage guide window, the control unit may provide an app usage guide window that includes an app usage description image, app usage description text, a previous button showing a previous app usage description page, a next button showing a next app usage description page, and a current page number.
[0055] Here, the control unit does not provide a previous button if the app usage description page is the first page, changes the next button to a confirmation button if the app usage description page is the last page, and when the current app usage description page changes to the next app usage description page, the app usage description image may also be changed.
[0056] As another example, when providing an avatar setting screen, the control unit may provide an avatar setting screen that includes a previous page button for returning to the previous page, an avatar image, a previous avatar button for showing a previous avatar, a next avatar button for showing a next avatar, and an avatar confirmation button for selecting and saving an avatar.
[0057] In some cases, when providing an avatar setting screen, the control unit may provide an avatar setting screen including a personal avatar creation button, and when the personal avatar creation button is selected, provide a personal avatar creation window through which the user can directly create a personal avatar.
[0058] In another case, when providing an avatar setting screen, the control unit may provide an avatar setting screen including a button for creating a user-like avatar, and when the button for creating a user-like avatar is selected, obtain a user image, input the user image into a pre-learned neural network model to extract user features, and generate and provide an avatar that resembles the user based on the user features.
[0059] Here, when extracting user feature points, the control unit extracts only user face feature points from the user face region if the user image contains only the user body, extracts only user body feature points from the user body region if the user image contains only the user body, and extracts both user face feature points and user body feature points from the user face and body regions if the user image contains both the user face and body.
[0060] In another case, when providing an avatar setting screen, the control unit may provide an avatar setting screen including a button for creating a user-like avatar, and when the button for creating a user-like avatar is selected, obtain a user video, input the user video into a pre-trained neural network model to extract user motion features, and generate and provide an avatar having user motion characteristics resembling the user based on the user motion features.
[0061] Here, when extracting user motion feature points, the control unit extracts only user face motion feature points from the user face region if the user video contains only the user face, extracts only user body motion feature points from the user body region if the user video contains only the user body, and extracts both user face motion feature points and user body motion feature points from the user face and body regions if the user video contains both the user face and body.
[0062] Next, the control unit provides an avatar setting screen, and when the avatar setting is completed, an avatar nickname setting screen is provided, and when the avatar's nickname is set through the nickname setting screen, the set avatar's nickname can be stored in a database.
[0063] Here, the control unit may provide an avatar nickname setting screen that includes a previous page button for returning to a previous page, an avatar nickname text input field, a nickname text entered in the avatar nickname text input field, and an avatar nickname confirmation button for setting and saving the avatar nickname when providing an avatar nickname setting screen.
[0064] Additionally, the control unit can generate a random nickname text automatically selected at random within the avatar nickname text input field when the avatar nickname setting screen is first provided, and modify the random nickname text according to the user's modification request.
[0065] Next, when a request for creating an exercise session is received from a user terminal (100), the control unit provides an exercise session creation window through which an individual user can directly create an exercise session, and can store personal exercise session information created through the exercise session creation window in a database.
[0066] And, when requesting a user-captured video, the control unit configures the main screen of the exercise session so that an avatar selected by the user is uploaded when a request for participation in an exercise session is received from the user terminal (100), requests the user-captured video from the user terminal (100), and when the user-captured video is received, provides the main screen of the exercise including the user-captured video.
[0067] Here, when configuring the main exercise screen of the exercise session, if there are multiple users who have requested to participate in the exercise session, the control unit can configure the main exercise screen of the exercise session so that the avatars of the users who have requested to participate and the participant avatars of other users are arranged at a certain interval in the virtual space of the exercise session.
[0068] For example, when arranging avatars in the virtual space of an exercise session, the control unit may arrange the avatars in the virtual space of the exercise session in one of the following methods: a first arrangement method that sequentially arranges the avatars according to the order of participation in the exercise session, a second arrangement method that arranges the avatars centered on the user's own avatar, and a third arrangement method that arranges the avatars in preset fixed positions.
[0069] In addition, the control unit may arrange the avatars in the virtual space of the exercise session in any one of a fourth arrangement method for arranging the avatars based on at least one of the motion size and voice size of the avatars participating in the exercise session, a fifth arrangement method for arranging the avatars based on a score calculated based on the motion counter and motion scoring of the avatars participating in the exercise session, and a sixth arrangement method for arranging only the avatars designated by the user among the avatars participating in the exercise session.
[0070] In some cases, the control unit may arrange only a random number of avatars among the avatars participating in the exercise session when the number of avatars participating in the exercise session is too large to display all of them on the screen or the visibility of each avatar is reduced.
[0071] Here, the control unit, when arranging avatars in the virtual space of the exercise session, can place name tags including nickname text on the upper side of the avatars participating in the exercise session, and can add an identifier different from the name tags of other participant avatars to the name tag of my avatar among the avatars.
[0072] For example, another identifier may include at least one of the color of the name tag, the shape of the name tag, or the addition of a specific icon.
[0073] In some cases, when arranging avatars in the virtual space of the exercise session, the control unit may place a name tag including nickname text on at least one of the upper and lower sides of the avatars participating in the exercise session, and add a specific identifier to the name tag of the room manager avatar who created the exercise session among the avatars.
[0074] Here, a specific identifier may include an icon of a specific shape.
[0075] In another case, when arranging avatars in the virtual space of an exercise session, the control unit may place name tags including nickname text on the upper side of the avatars participating in the exercise session and add an identifier to the name tag of an avatar that is speaking among the avatars.
[0076] For example, the identifier may include at least one of a border highlight and a speech icon on the name tag of the speaking avatar.
[0077] In addition, the control unit can configure the exercise main screen of the exercise session so that, when configuring the exercise main screen of the exercise session, a room manager screen area including input text of the room manager who created the exercise session is formed on one side of the area where avatars are arranged before and after playing the exercise session.
[0078] Here, the control unit can provide a specific image or a specific video to the room screen area.
[0079] Next, the control unit can configure the main exercise screen of the exercise session so that a user screen area including a user-captured video and exercise play time is formed on one side of an area where avatars are arranged during play of the exercise session.
[0080] Here, when a user-captured image is received, the control unit can check whether the user's full body is included in the user-captured image, and if the user's full body included in the user-captured image is less than a standard ratio, a field of view check pop-up is generated and provided, and if the user's full body included in the user-captured image is greater than a standard ratio, the generated field of view check pop-up can be removed.
[0081] For example, the control unit may generate a view angle check pop-up that includes at least one of a user-captured video, guidance text for the view angle, and an operation button for operating the screen.
[0082] In some cases, the control unit may display the user's currently detected key points and a skeleton connecting the key points over the captured video.
[0083] Additionally, the control unit can generate and provide a view angle check pop-up when the exercise main screen is first provided.
[0084] Additionally, the control unit can check whether the user's full body is included in the user-captured image for each frame while the exercise main screen is provided, and can generate and provide a field of view check pop-up if the user's full body included in the user-captured image is below a standard ratio.
[0085] Here, the control unit generating the angle check pop-up is only one example and is not limited thereto, and various other forms of UI can be provided in addition to the angle check pop-up.
[0086] In addition, when checking whether the user's full body is included in the user-captured image, the control unit can input the user's captured image into a pre-learned neural network model to infer skeleton points corresponding to the user's full body, and analyze the skeleton points corresponding to the user's full body to check whether the user's full body is included in the user-captured image.
[0087] In addition, when providing a view angle check pop-up, the control unit may generate and provide a view angle check pop-up if a user-captured image with a user's whole body ratio below a standard ratio is continuously received for a standard period of time or longer.
[0088] In some cases, when the control unit provides a field of view check pop-up, if a user-captured image with a user's full body ratio greater than or equal to a standard ratio is continuously received for a standard period of time, the generated field of view check pop-up can be removed.
[0089] Here, the control unit generates a view angle check pop-up, but this is only one example and is not limited thereto, and can be provided in various other forms of UI in addition to the view angle check pop-up.
[0090] Next, the control unit can predict the distance between the camera position of the user terminal and the user based on the user motion data, and adjust the position of the avatar based on the predicted distance.
[0091] When extracting user motion data, a pre-learned neural network model can be used to extract 3D keypoints for joint locations from the user's body included in the user's captured video, and user motion data can be extracted based on the rotation parameters corresponding to each joint keypoint.
[0092] Here, when extracting 3D keypoints, if the number of extracted keypoints is less than a reference ratio or the visibility of predicted keypoints is less than a reference ratio, the control unit determines that the user motion is invalid, and stops updating the user motion data, or predicts the user motion data in advance and then replaces it with other user motion data corresponding to the motion.
[0093] In some cases, when the control unit determines the user motion based on the keypoint, if the user motion deviates from the preset reference motion by more than a reference range, the control unit may determine that the user motion is invalid.
[0094] Additionally, when the control unit determines the user motion based on the key point, if the rate of change in the user motion becomes greater than the reference rate of change in the motion while the user-captured video is received at a constant number of frames per second (fps), the control unit can determine that the user motion is invalid.
[0095] Additionally, when extracting user motion data, the control unit can correct the three-dimensional rotation parameters of the joints if it is predicted that areas occupied by joints corresponding to the user motion collide with each other based on the three-dimensional rotation parameters of the joints.
[0096] Next, when controlling the motion of the avatar, if there are multiple users who have requested to participate in the exercise session, the control unit can determine whether the sync for the same motion of the avatar of the user who requested to participate matches the sync for the same motion of the participant avatars of other users, and if the sync for the same motion does not match, control the avatar motion so as to match the sync for the same motion of the avatars participating in the exercise session.
[0097] And, when controlling the motion of the avatar, the control unit can adjust the size of the speaking avatar to identify the speaking avatar among the avatars participating in the exercise session.
[0098] Here, the control unit can adjust the exercise main screen so that, if an avatar participating in the exercise session is speaking outside the exercise main screen, the speaking avatar is located within the exercise main screen.
[0099] Additionally, the control unit can highlight and identify the participant avatar calling the nickname of the user's own avatar among the avatars participating in the exercise session, if there is a participant avatar of another user calling the nickname of the user's own avatar.
[0100] Additionally, the control unit can animate the shape of the avatar's mouth based on the user's voice.
[0101] Here, the control unit can input the user's voice into a pre-trained neural network model to infer the shape of the avatar's mouth, and animate the shape of the avatar's mouth to be synchronized with the user's voice based on the inferred shape of the mouth.
[0102] Additionally, the control unit may provide a menu pop-up including a first menu button for moving to an app usage guide screen, a second menu button for moving to an avatar setting screen, a third menu button for moving to a nickname setting screen, a fourth menu button for moving to a user inquiry channel, a fifth menu button for moving to an open chat room, and a fifth menu button for moving to a metaverse service providing website when a user request for selecting a menu item within the home screen is received.
[0103] Additionally, the control unit may provide a volume control pop-up including an exercise video volume control controller, a user-to-user conversation volume control controller, a confirmation button for saving a volume control state, and a volume status display unit for showing a current volume control state when a user request for selecting a volume control item within the exercise main screen is received.
[0104] Additionally, the control unit may provide an exit popup that includes an exit button that navigates to the home screen and a cancel button that cancels the exit and returns to the exercise main screen when a user request to select an exit item within the exercise main screen is received.
[0105] Additionally, the control unit may provide an additional pop-up when a user request to select another user's avatar within the exercise main screen is received.
[0106] An example of a button included in a pop-up may include at least one of a button for individually adjusting the voice volume of the current user who clicked, a button for muting, a button for adding friends, a button for reporting, a button for adjusting the position / size of the avatar, and a button for kicking, depending on whether the current user has room moderator rights or not, but this is only an example and is not limited thereto.
[0107] Additionally, the control unit may provide a workout authentication sharing pop-up that includes a user avatar workout recording video, a workout video save button, a workout video share button, and a pop-up close button when the workout session ends.
[0108] Here, the user avatar exercise recording video may include at least one of the user's exercise result information, the user's avatar exercise motion, the number of times the user exercised, and a specific logo, but this is only an example and is not limited thereto.
[0109] For example, the user's exercise result information may include at least one of the exercise type, exercise time, and calories consumed.
[0110] In addition, when a request for recording an exercise video for a user avatar is received from a user terminal while the avatar participating in the exercise session is exercising, the control unit can record the exercise appearance of the user avatar from the time the exercise video recording request is received.
[0111] In some cases, the control unit may record all avatars participating in the exercise session simultaneously while the exercise session is in progress.
[0112] In another case, the control unit may compare the exercise movements of an avatar participating in the exercise session with a preset reference movement when the exercise session is in progress, calculate a score based on the degree of agreement, and record the avatar exercising at the time when the exercise movement with the highest score of agreement is performed.
[0113] Additionally, if the recording point for each exercise video is stored in the exercise session information, the control unit can perform recording at that point.
[0114] For example, the control unit can record the first 10 seconds after the exercise begins or from the start of the full squat movement (7 minutes 10 seconds) to the end (7 minutes 50 seconds).
[0115] In another case, the control unit may analyze the exercise movements of an avatar participating in an exercise session when the exercise session is in progress, and record the avatar exercising at the point in time when the exercise movement with the largest motion size is performed.
[0116] In another case, the control unit may record the avatar exercising at a point in time when the exercise session is in progress and the avatar is performing the exercise movements in the latter part of the exercise session.
[0117] In another case, the control unit can check the average volume of voice chat of avatars participating in the exercise session when the exercise session is in progress, and record the avatars exercising at the time when the average volume of voice chat is the loudest.
[0118] In addition, when saving a user avatar movement recording video, the control unit can extract only the motion data of the user avatar from the user avatar movement recording video and save it in a database.
[0119] Here, when playback of a user avatar movement recording video is requested, the control unit obtains motion data corresponding to the user avatar from a database, reflects the motion data to the user avatar, generates a user avatar movement recording video, and can play the generated user avatar movement recording video.
[0120] In addition, when consent to recording a user-captured video is received from a user terminal (100), the control unit can simultaneously record a video of a user avatar participating in an exercise session exercising and a user-captured video received from the user terminal (100) to store a video including a user avatar exercising in a virtual space and a user exercising in a real space.
[0121] Next, the control unit provides an avatar photo shooting pop-up for taking an exercise authentication photo when the exercise session ends, and when a request for taking a group authentication photo is received through the avatar photo shooting pop-up, it counts down a set time for the avatars participating in the exercise session to take individual poses or group poses, and when the set time arrives, an authentication photo of the avatars participating in the exercise session can be taken.
[0122] For example, the avatar photo shoot pop-up may include at least one of the following: exercise date, exercise time, exercise calories burned, exercise type, a button to take a group photo, and a button to do exercise one more time.
[0123] Meanwhile, the neural network model described above may be a deep neural network. Throughout this specification, the terms neural network, network function, and neural network may be used interchangeably. A deep neural network (DNN) may refer to a neural network that includes multiple hidden layers in addition to an input layer and an output layer. A deep neural network can be used to identify latent structures in data. That is, it can identify latent structures of photos, text, videos, voices, and music (e.g., what objects are in a photo, what the content and emotion of a text is, what the content and emotion of a voice is, etc.). A deep neural network may include a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a Q network, a U network, a Siamese network, etc.
[0124] According to one embodiment of the present disclosure, the control unit may be configured with one or more cores and may include a processor for data analysis and deep learning, such as a central processing unit (CPU), a general purpose graphics processing unit (GPGPU), and a tensor processing unit (TPU) of a computing device. The control unit may read a computer program stored in a database and perform data processing for machine learning according to one embodiment of the present disclosure. According to one embodiment of the present disclosure, the control unit may perform operations for learning a neural network. The control unit may perform calculations for learning a neural network, such as processing input data for learning in deep learning (DL), extracting features from input data, calculating errors, and updating weights of a neural network using backpropagation. At least one of the CPU, GPGPU, and TPU of the control unit may process learning of a network function. For example, the CPU and GPGPU may together process learning of a network function and classification of data using a network function. Additionally, in one embodiment of the present disclosure, the processors of multiple computing devices can be used together to process network function learning and data classification using network functions. Furthermore, a computer program executed on a computing device according to one embodiment of the present disclosure may be a CPU, GPGPU, or TPU executable program.
[0125] Additionally, the communication unit may include a wireless Internet module, a short-range communication module, a location information module, etc.
[0126] A wireless Internet module is a module for wireless Internet access, and is configured to transmit and receive wireless signals in a communication network according to wireless Internet technologies.
[0127] Wireless Internet technologies include, for example, WLAN (Wireless LAN), Wi-Fi (Wireless-Fidelity), Wi-Fi (Wireless Fidelity) Direct, DLNA (Digital Living Network Alliance), WiBro (Wireless Broadband), WiMAX (World Interoperability for Microwave Access), HSDPA (High Speed Downlink Packet Access), HSUPA (High Speed Uplink Packet Access), LTE (Long Term Evolution), and LTE-A (Long Term Evolution-Advanced), and the wireless Internet module transmits and receives data according to at least one wireless Internet technology, including Internet technologies not listed above.
[0128] From the perspective that wireless Internet access through WiBro, HSDPA, HSUPA, GSM, CDMA, WCDMA, LTE, LTE-A, etc. is achieved through a mobile communication network, a wireless Internet module that performs wireless Internet access through a mobile communication network can be understood as a type of the above mobile communication module.
[0129] The short-range communication module is for short-range communication and can support short-range communication using at least one of Bluetooth™, RFID (Radio Frequency Identification), Infrared Data Association (IrDA), UWB (Ultra Wideband), ZigBee, NFC (Near Field Communication), Wi-Fi (Wireless-Fidelity), Wi-Fi Direct, and Wireless USB (Wireless Universal Serial Bus) technologies.
[0130] A location information module is a module for obtaining the location (or current location) of a server, and representative examples include a GPS (Global Positioning System) module or a WiFi (Wireless Fidelity) module.
[0131] And, the database may include at least one type of storage medium among flash memory type, hard disk type, multimedia card micro type, card type memory (e.g., SD or XD memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, and optical disk.
[0132] The server (200) of the present disclosure may operate in connection with web storage that performs the storage function of a database (120) on the Internet. The description of the database described above is merely an example and is not limited thereto.
[0133] Additionally, the server (200) of the present disclosure may further include an output unit and an input unit.
[0134] Here, the output unit can display a user interface (UI) for providing a virtual space of the metaverse. The output unit can output any form of information generated or determined by the control unit and any form of information received by the communication unit.
[0135] For example, the output unit may include at least one of a liquid crystal display (LCD), a thin film transistor-liquid crystal display (TFT LCD), an organic light-emitting diode (OLED), a flexible display, and a three-dimensional display (3D display). Some of these display modules may be configured as transparent or light-transmitting so that the outside can be viewed through them. This may be referred to as a transparent display module, and a representative example of a transparent display module is TOLED (Transparent OLED).
[0136] Additionally, the input unit can receive user input. The input unit may include keys and / or buttons on a user interface for receiving user input, or physical keys and / or buttons. A computer program according to embodiments of the present disclosure may be executed based on user input through the input unit.
[0137] Additionally, the input unit may receive a signal by detecting the user's button operation or touch input, or may receive the user's voice or motion through a camera or microphone and convert it into an input signal. For this purpose, speech recognition technology or motion recognition technology may be used.
[0138] Additionally, the input unit may be implemented as an external input device connected to the server (200). For example, the input device may be at least one of a touch pad, a touch pen, a keyboard, or a mouse for receiving user input, or may include at least one of a VR device and a motion tracking suit / device, but this is merely an example and is not limited thereto.
[0139] For example, the input unit can recognize a user touch input. In some cases, the input unit may have the same configuration as the output unit. The input unit may be configured as a touch screen implemented to receive a user's selection input. The touch screen may use any one of a contact-type electrostatic capacitance method, an infrared light sensing method, a surface acoustic wave (SAW) method, a piezoelectric method, and a resistive film method. The detailed description of the touch screen described above is merely an example according to one embodiment of the present disclosure, and various touch screen panels may be employed in the platform server (100). The input unit configured as a touch screen may include a touch sensor. The touch sensor may be configured to convert a change in pressure applied to a specific portion of the input unit or an electrostatic capacitance generated at a specific portion of the input unit into an electrical input signal. The touch sensor may be configured to detect not only the position and area of a touch, but also the pressure at the time of the touch. When there is a touch input to the touch sensor, a corresponding signal(s) is sent to the touch controller. The touch controller can process the signal(s) and then transmit corresponding data to the processor. This allows the processor to recognize which area of the input unit has been touched, etc.
[0140] The server (200) of the present disclosure may include other components for implementing a server environment. The server (200) may include any type of device. The server (200) may be a digital device equipped with a processor, memory, and computing power, such as a laptop computer, notebook computer, desktop computer, web pad, or mobile phone.
[0141] The server (200) of the present disclosure may be a computing system capable of generating a user interface according to embodiments of the present disclosure and providing information to a user terminal via a network. The server (200) may transmit a virtual space of the metaverse space to a user terminal (100). In this case, the user terminal (100) may be any type of computing device capable of accessing the server (200). The control unit of the server (200) may transmit the user interface to the user terminal (100) via a communication unit. The server (200) of the present disclosure may be a cloud server. The server (200) may be a web server that processes services. The types of servers (200) described above are merely examples and are not limited thereto.
[0142] In this way, the present disclosure can provide a sense of realism, interest, and fun to the user by allowing the user and the avatar to exercise together by reflecting the user's motion data on the avatar uploaded to the virtual space of the exercise session and controlling the motion of the avatar by predicting the next movement through a pre-learned neural network model.
[0143] Figures 2 to 6 are flowcharts for explaining a method for providing a metaverse service according to the present disclosure.
[0144] FIG. 2 is a flowchart for explaining the process of providing metaverse services before a user participates in an exercise session, FIG. 3 is a flowchart for explaining the process of providing metaverse services after a user participates in an exercise session, FIG. 4 is a flowchart for explaining the process of synchronizing user motions in FIG. 3, FIG. 5 is a flowchart for explaining the process of real-time voice chatting in FIG. 3, and FIG. 6 is a flowchart for explaining the process of generating an exercise result image in FIG. 3.
[0145] First, as illustrated in FIG. 2, the present disclosure checks whether the user has registered as a member (S102) when the user terminal runs the metaverse service providing app (S101), and if the user has registered as a member, the user terminal can be connected to the play server (S103).
[0146] Here, the present disclosure may provide an onboarding screen to a user terminal if the user is not a member (S104), and may provide an app usage guide window, such as a tutorial, if the user terminal requests to start the app through the onboarding screen (S105).
[0147] Next, the present disclosure provides an avatar setting screen when user confirmation is received through the app usage guide window, and allows the user to set an avatar through the avatar setting screen (S106), and stores the set avatar and user information in a database (S107).
[0148] In addition, the present disclosure can provide a home screen to a user terminal that includes a list of scheduled exercise sessions from a content server and information on currently ongoing exercise sessions from a multiplayer server when the user terminal is connected to a play server (S108).
[0149] Here, the present disclosure can extract exercise session information based on an exercise session start schedule from an exercise session list pre-stored in a database when providing exercise session information, and provide the extracted exercise session information to a user terminal.
[0150] For example, exercise session information may include at least one of an exercise schedule, exercise start time, exercise duration, exercise title, calories burned, exercise composition movements, URL (Uniform Resource Locator) information of the corresponding content, and information for composing an exercise result video, but this is only an example and is not limited thereto.
[0151] In addition, the present disclosure can extract exercise session information having an exercise schedule for a predetermined period of time when providing exercise session information, and sequentially provide the extracted exercise session information according to the order of the exercise schedule.
[0152] Here, when providing exercise session information sequentially, the present disclosure can first provide exercise session information having the earliest schedule among the extracted exercise session information, and lastly provide exercise session information having the latest schedule.
[0153] In some cases, the present disclosure may extract personalized recommended exercise session information based on user information of a user terminal when providing exercise session information, and sequentially provide the extracted exercise session information according to the order of the exercise schedule.
[0154] For example, the present disclosure can extract personalized recommended exercise session information based on at least one physical condition among age, gender, medical history, and weight among user information of a user terminal, but this is only one embodiment and is not limited thereto.
[0155] In another case, the present disclosure may, when providing exercise session information, extract exercise session information based on exercise session participation history information of a user terminal, and sequentially provide the extracted exercise session information according to the order of the exercise schedule.
[0156] Here, the present disclosure can extract exercise session information based on at least one activity condition among the most recently participated exercise session information, the most frequently participated exercise session information, and the most time-spent exercise session information among the exercise session participation history information of the user terminal.
[0157] Next, the present disclosure allows a user of a user terminal to check whether there is a currently ongoing exercise session through exercise session information (S109), and if there is an currently ongoing exercise session, to participate in it (S110), or if there is no currently ongoing exercise session, to set an exercise notification for an exercise session scheduled to be in progress (S111).
[0158] Here, the present disclosure can provide an exercise screen when a user of a user terminal requests to participate in an exercise session currently in progress (S112).
[0159] As illustrated in FIG. 3, the present disclosure provides an exercise session preparation screen (S114) when a user terminal enters an exercise screen (S113), and can check the play server connection status (S115).
[0160] Here, the present disclosure can synchronize user motions to an avatar if the user terminal is in a good connection state with the play server (S116), start real-time voice chat (S117), and generate a result video (S118).
[0161] However, in the present disclosure, if the user terminal is not in a good connection state with the play server, it can return to the home screen (S119).
[0162] Next, the present disclosure performs a user motion synchronization process, a real-time voice chat process, and a result video generation process, and then checks the initialization status (S120), and if the initialization status is good, starts the exercise session when a start event is transmitted from the multiplayer server (S121), and if the initialization status is bad, returns to the home screen (S119).
[0163] Next, the present disclosure provides exercise content video from a content server to enable streaming playback of exercise content (S122).
[0164] In addition, the present disclosure can continuously synchronize the playback time of all users (S123), continuously check whether content is complete (S124), and continuously check the play server connection status (S125).
[0165] Next, the present disclosure provides an exercise session completion screen when an exercise session is completed (S126), and an exercise result image can be shared and saved (S127).
[0166] Next, the present disclosure can return to the home screen (S129) when the host of the exercise session ends the exercise session or the user ends it directly (S128).
[0167] The present disclosure can perform a user motion synchronization process (S116) as shown in FIG. 4.
[0168] As illustrated in FIG. 4, the present disclosure activates the front camera of a user terminal (S201), starts an artificial intelligence pose model (S202), filters pose results from a user-captured video captured through the front camera (S203), and performs a forward kinematic algorithm and an inverse kinematic algorithm (S204, S205) to generate a final avatar pose (S206). Here, the filtering may be performed before or after the forward kinematic or after the inverse kinematic.
[0169] As an example, the present disclosure can extract 3D key points for joint locations from a user's body included in a user-captured video using a pre-learned neural network model, and extract user motion data based on rotation parameters corresponding to each joint key point.
[0170] In addition, the present disclosure can confirm the validity of the result from the user-captured video (S208) and initialize the avatar pose if it is invalid.
[0171] For example, the present disclosure determines that the user motion is invalid if the number of key points corresponding to the user body joints from the user-captured video is less than a reference ratio or the visibility of the predicted key points is less than a reference ratio, and stops updating the user motion data or predicts the user motion data in advance and then replaces it with other user motion data corresponding to the motion.
[0172] In some cases, the present disclosure may determine that a user motion is invalid if the user motion deviates from a preset reference motion by more than a reference range.
[0173] In another case, the present disclosure may determine that a user motion is invalid if the rate of change in the motion of the user motion is greater than a reference rate of change in the motion while the user-captured video is received at a constant number of frames per second (fps).
[0174] Next, the present disclosure can check the user position within the camera angle from the user captured video (S210) and provide an angle check pop-up if it is invalid.
[0175] For example, the present disclosure can check whether a user's full body is included in a user-captured image, and if the user's full body included in the user-captured image is below a standard ratio, a field of view check pop-up can be generated and provided, and if the user's full body included in the user-captured image is above a standard ratio, the generated field of view check pop-up can be removed.
[0176] In the present disclosure, when providing a view angle check pop-up, if a user-captured image having a user's full body ratio below a standard ratio is continuously received for a standard period of time or longer, the view angle check pop-up can be generated and provided.
[0177] In some cases, the present disclosure may remove the generated angle check pop-up when a user-captured image having a user body ratio greater than or equal to a standard ratio is continuously received for a standard period of time when providing the angle check pop-up.
[0178] Here, the present disclosure, generating a view angle check pop-up is only one embodiment and is not limited thereto, and various other forms of UI can also be provided in addition to the view angle check pop-up.
[0179] Next, the present disclosure can generate a final avatar pose (S206) by adjusting the avatar position from the user-captured video (S213).
[0180] Here, the present disclosure can predict the distance between the camera position of the user terminal and the user based on user motion data, and adjust the position of the avatar based on the predicted distance.
[0181] In this way, the present disclosure can update the user's own avatar posture uploaded to the exercise screen (S207) and continuously synchronize the avatar posture of another user uploaded to the exercise screen based on the final avatar posture generated from another user terminal (S214).
[0182] In addition, the present disclosure can perform a real-time voice chat process (S117) as shown in FIG. 5.
[0183] As illustrated in FIG. 5, the present disclosure can connect real-time voice chat (S301) and transmit and play voice through a voice chat server (S302).
[0184] And, the present disclosure can perform icon notation for avatar identification corresponding to the confirmed speaker after speaker identification (S303).
[0185] Next, the present disclosure can be performed repeatedly until the service ends (S304).
[0186] As an example, the present disclosure may add an identifier to the name tag of an avatar that is speaking among the avatars.
[0187] Here, the identifier may include at least one of a border highlight and a utterance icon on the name tag of the avatar that is speaking.
[0188] In addition, the present disclosure can adjust the size of an avatar that is speaking in order to identify the avatar that is speaking among the avatars participating in the exercise session.
[0189] Here, the present disclosure can identify by highlighting the participant avatar calling the nickname of another user's avatar among the avatars participating in the exercise session.
[0190] Additionally, the present disclosure can animate the mouth shape of an avatar based on the user's voice.
[0191] Here, the present disclosure can input a user's voice into a pre-trained neural network model to infer the mouth shape of an avatar, and animate the mouth shape of the avatar to be synchronized with the user's voice based on the inferred mouth shape.
[0192] In addition, the present disclosure can perform the exercise result image generation process (S118) as shown in FIG. 6.
[0193] As illustrated in FIG. 6, the present disclosure can configure a recording screen corresponding to exercise session information and member information from the exercise session database and member information database of the content server (S401).
[0194] And, the present disclosure records an exercise result video (S402), saves the recorded video as a temporary file (S403), and then shares the recorded video through a social network service (S404) or saves it in a gallery (S405).
[0195] As an example, the present disclosure can extract only the motion data of a user avatar from a recorded video of the user avatar's movements and store the extracted data in a database.
[0196] Here, the present disclosure can obtain motion data corresponding to the user avatar from a database when playback of a user avatar motion recording video is requested, reflect the motion data to the user avatar, generate a user avatar motion recording video, and play the generated user avatar motion recording video.
[0197] In addition, the present disclosure can simultaneously record a video of a user avatar participating in an exercise session exercising and a user-captured video received from the user terminal when consent to recording a user-captured video is received from the user terminal, thereby storing a video including the user avatar exercising in a virtual space and the user exercising in a real space.
[0198] FIGS. 7 to 19 are drawings showing screen implementations of a metaverse service providing device according to the present disclosure.
[0199] As illustrated in FIG. 7, the present disclosure may provide an onboarding screen (310) when the membership registration of the user terminal is not confirmed.
[0200] Here, the present disclosure can be provided only once after app installation when providing an onboarding screen (310).
[0201] And, the present disclosure provides an onboarding screen (310) including a menu button (312), a background image (314), and a start button (316), and can receive an app start request through the start button (316).
[0202] As illustrated in FIG. 8, the present disclosure may provide an app usage guide window (320) when an app launch request is received through an onboarding screen.
[0203] Here, in the present disclosure, when providing an app usage guide window (320), the app usage guide window (320) can be provided only once in the form of a pop-up that provides information on how to use the app.
[0204] In some cases, the control unit may re-present the app usage guide window after the onboarding process or may re-present the app usage guide window based on user input selected from a menu.
[0205] For example, when providing an app usage guide window (320), the present disclosure may provide an app usage guide window (320) that includes an app usage description image (322), app usage description text (324), a previous button (326) that shows a previous app usage description page, a next button (328) that shows a next app usage description page, and a current page number (327).
[0206] Here, the present disclosure does not provide a previous button (326) if the app usage description page is the first page, changes the next button (328) to a confirmation button if the app usage description page is the last page, and when the current app usage description page is changed to the next app usage description page, the app usage description image (322) may also be changed.
[0207] As illustrated in FIG. 9, the present disclosure can provide an avatar setting screen (330) when user confirmation is received through an app usage guide window, and can store the avatar and user information set by the user through the avatar setting screen (330) in a database.
[0208] The present disclosure may provide an avatar setting screen (330) that includes a previous page button (332) for returning to a previous page, an avatar image (334), a previous avatar button (335) for showing a previous avatar, a next avatar button (336) for showing a next avatar, and an avatar confirmation button (337) for selecting and saving an avatar.
[0209] In some cases, the present disclosure may provide an avatar setting screen (330) that includes a personal avatar creation button when providing an avatar setting screen (330), and when the personal avatar creation button is selected, provide a personal avatar creation window that allows the user to directly create a personal avatar.
[0210] In another case, the present disclosure provides an avatar setting screen (330) including a button for creating a user-like avatar when providing an avatar setting screen (330), and when the button for creating a user-like avatar is selected, a user image is acquired, the user image is input into a pre-learned neural network model to extract user features, and an avatar that resembles the user is created and provided based on the user features.
[0211] Here, the present disclosure extracts user feature points when extracting user feature points, if the user image contains only the user face, only the user face feature points are extracted from the user face region, if the user image contains only the user body, only the user body feature points are extracted from the user body region, and if the user image contains both the user face and body, both the user face feature points and the user body feature points can be extracted from the user face and body regions.
[0212] In another case, the present disclosure provides an avatar setting screen (330) including a button for creating a user-like avatar when providing an avatar setting screen (330), and when the button for creating a user-like avatar is selected, a user video is acquired, the user video is input into a pre-learned neural network model to extract user motion characteristics, and an avatar having user motion characteristics similar to the user motion characteristics is created and provided based on the user motion characteristics.
[0213] Here, in the present disclosure, when extracting user motion feature points, if only the user face is included in the user video, only the user face motion feature points are extracted from the user face region, if only the user body is included in the user video, only the user body motion feature points are extracted from the user body region, and if both the user face and body are included in the user video, both the user face motion feature points and the user body motion feature points can be extracted from the user face and body regions.
[0214] As illustrated in FIG. 10, the present disclosure provides an avatar nickname setting screen (340) when avatar setting is completed through the avatar setting screen, and when the nickname of the avatar is set through the nickname setting screen (340), the set nickname of the avatar can be stored in a database.
[0215] Here, the present disclosure may provide an avatar nickname setting screen (340) that includes a previous page button (342) for returning to a previous page, an avatar nickname text input field (344), a nickname text (347) entered in the avatar nickname text input field, and an avatar nickname confirmation button (346) for setting and saving the avatar nickname.
[0216] In addition, the present disclosure can generate a random nickname text (347) that is automatically selected randomly within the avatar nickname text input field (344) when the avatar nickname setting screen (340) is first provided, and modify the random nickname text (347) according to a user's modification request.
[0217] As illustrated in FIG. 11, the present disclosure may provide a home screen (350) including a schedule card (358) having exercise session information and an avatar (354) selected by the user when an avatar is set through the avatar setting screen.
[0218] Here, the present disclosure may provide workout session information including at least one of the current or nearest next workout session date and time, workout session status, workout session title, and main action button (359) within a schedule card (358).
[0219] For example, the status of a workout session may include a waiting state before the workout session starts or an in-progress state after the workout session starts.
[0220] As another example, the main action button (359) may have different text, color, and actions depending on at least one of the following: reservation status, waiting status, waiting and full capacity status, progress and under capacity status, and progress and full capacity status.
[0221] Additionally, the present disclosure may provide a menu item (352), a background image (356) of a virtual space, and a welcome message (357) within a home screen (350).
[0222] Here, the welcome message (357) may have different text contents depending on at least one of the following situations: reservation status, waiting status, waiting and over capacity status, progress and under capacity status, and progress and over capacity status.
[0223] In addition, the present disclosure can extract exercise session information based on an exercise session start schedule from an exercise session list pre-stored in a database when providing exercise session information, and provide the extracted exercise session information.
[0224] For example, exercise session information may include at least one of an exercise schedule, exercise start time, exercise duration, exercise title, calories burned, exercise composition movements, URL (Uniform Resource Locator) information of the corresponding content, and information for composing an exercise result video, but this is only an example and is not limited thereto.
[0225] In addition, the present disclosure can extract exercise session information having an exercise schedule for a predetermined period of time when providing exercise session information, and sequentially provide the extracted exercise session information according to the order of the exercise schedule.
[0226] Here, when providing exercise session information sequentially, the present disclosure can first provide exercise session information having the earliest schedule among the extracted exercise session information, and lastly provide exercise session information having the latest schedule.
[0227] In some cases, the present disclosure may extract personalized recommended exercise session information based on user information when providing exercise session information, and sequentially provide the extracted exercise session information according to the order of the exercise schedule.
[0228] For example, the present disclosure may extract personalized recommended exercise session information based on at least one physical condition among age, gender, medical history, and weight among user information, but this is only an example and is not limited thereto.
[0229] In another case, the present disclosure may, when providing exercise session information, extract exercise session information based on the user's exercise session participation history information, and sequentially provide the extracted exercise session information according to the order of the exercise schedule.
[0230] Here, the present disclosure can extract exercise session information based on at least one activity condition among the most recently participated exercise session information, the most frequently participated exercise session information, and the exercise session information on which the user spends the most time, among the exercise session participation history information of the user.
[0231] As illustrated in FIG. 12, the present disclosure configures the exercise main screen (360) of the exercise session so that an avatar selected by the user is uploaded when a request for participation in an exercise session is received from a user terminal, requests a user-captured video from the user terminal, and provides the exercise main screen (360) including the user-captured video when the user-captured video is received.
[0232] Here, in the present disclosure, when configuring the exercise main screen (360) of an exercise session, if there are multiple users who have requested to participate in the exercise session, the exercise main screen (360) of the exercise session can be configured so that the avatars of the users who have requested to participate and the participant avatars of other users are arranged at a certain interval in the virtual space of the exercise session.
[0233] For example, the present disclosure may arrange avatars in the virtual space of an exercise session in any one of a first arrangement method of sequentially arranging avatars according to the order of participation in the exercise session, a second arrangement method of arranging avatars centered on the user's own avatar, and a third arrangement method of arranging avatars at preset fixed positions.
[0234] In addition, the present disclosure can arrange avatars in the virtual space of an exercise session by any one of a fourth arrangement method of arranging avatars by at least one of the motion size and voice size of the avatars participating in the exercise session, a fifth arrangement method of arranging avatars by scores calculated based on the motion counter and motion scoring of the avatars participating in the exercise session, and a sixth arrangement method of arranging only avatars designated by the user among the avatars participating in the exercise session.
[0235] Here, the present disclosure can arrange avatars in a virtual space of an exercise session, place name tags including nickname text on the upper side of avatars participating in the exercise session, and add an identifier different from the name tags of other participant avatars to the name tag of my avatar among the avatars.
[0236] For example, another identifier may include at least one of the color of the name tag, the shape of the name tag, or the addition of a specific icon.
[0237] In some cases, the present disclosure may, when arranging avatars in a virtual space of an exercise session, place name tags including nickname text on the upper side of the avatars participating in the exercise session, and add a specific identifier to the name tag of the room manager avatar who created the exercise session among the avatars.
[0238] Here, a specific identifier may include an icon of a specific shape.
[0239] In another case, the present disclosure may, when arranging avatars in a virtual space of an exercise session, place a name tag including a nickname text on at least one of the upper and lower sides of the avatars participating in the exercise session, and add an identifier to the name tag of an avatar that is speaking among the avatars.
[0240] For example, the identifier may include at least one of a border highlight and a speech icon on the name tag of the speaking avatar.
[0241] In addition, the present disclosure can configure the exercise main screen (360) of the exercise session so that a room manager screen area (354) including input text (356) of the room manager who created the exercise session is formed on one side of an area where avatars are arranged before and after playing the exercise session.
[0242] Here, the present disclosure can provide a specific image or a specific video to the room screen area (354).
[0243] Additionally, in the present disclosure, when not during play, the time may not be displayed in the play time display area (352) included in the room screen area (354).
[0244] As illustrated in FIG. 13, the present disclosure can configure the exercise main screen (370) of an exercise session so that a user screen area including a user-captured image (372) and an exercise play time (374) is formed on one side of an area where avatars are arranged during play of an exercise session.
[0245] The present disclosure places a name tag including a nickname text on at least one of the upper and lower sides of avatars participating in an exercise session, and a specific identifier can be added to the name tag (378) of the room manager avatar who created the exercise session among the avatars.
[0246] Here, a specific identifier may include an icon of a specific shape, such as a crown.
[0247] The present disclosure determines whether a user's full body is included in a user's captured image (372) when a user's captured image (372) is received, and if the user's full body included in the user's captured image (372) is below a standard ratio, a field of view check pop-up is generated and provided, and if the user's full body included in the user's captured image is above a standard ratio, the generated field of view check pop-up can be removed.
[0248] Here, the present disclosure, generating a view angle check pop-up is only one embodiment and is not limited thereto, and can be provided in various other forms of UI in addition to the view angle check pop-up.
[0249] Next, the present disclosure can predict the distance between the camera position of the user terminal and the user based on user motion data, and adjust the position of the avatar based on the predicted distance.
[0250] Next, the present disclosure can extract 3D key points for joint locations from a user's body included in a user-captured video using a pre-learned neural network model, and extract user motion data based on rotation parameters corresponding to each joint key point.
[0251] Here, the present disclosure determines that the user motion is invalid if the number of extracted keypoints is less than a reference ratio or the visibility of the predicted keypoint is less than a reference ratio, and stops updating the user motion data or predicts the user motion data in advance and then replaces it with other user motion data corresponding to the motion.
[0252] In some cases, the present disclosure may determine that a user motion is invalid if the user motion deviates from a preset reference motion by more than a reference range.
[0253] In addition, the present disclosure can determine that the user motion is invalid if the rate of change in the motion of the user motion becomes greater than the reference rate of change in the motion while the user-captured video (372) is received at a constant number of frames per second (fps).
[0254] In addition, the present disclosure can correct the three-dimensional rotation parameters of the joints if it is predicted that the areas occupied by joints corresponding to the user motion collide with each other based on the three-dimensional rotation parameters of the joints.
[0255] Next, the present disclosure determines whether the sync for the same motion of the avatar of the user who requested to participate in the exercise session and the avatars of other users are identical when there are multiple users who requested to participate, and if the sync for the same motion is not identical, the avatar motion can be controlled to match the sync for the same motion of the avatars participating in the exercise session.
[0256] In addition, the present disclosure can adjust the size of an avatar that is speaking in order to identify the avatar that is speaking among the avatars participating in the exercise session.
[0257] Here, the present disclosure can adjust the exercise main screen (370) so that, if an avatar participating in an exercise session is speaking outside the exercise main screen (370), the speaking avatar is located within the exercise main screen (370).
[0258] Additionally, the present disclosure can identify by highlighting the participant avatar calling the nickname of another user's avatar among the avatars participating in the exercise session.
[0259] Additionally, the present disclosure can animate the mouth shape of an avatar based on the user's voice.
[0260] Here, the present disclosure can input a user's voice into a pre-trained neural network model to infer the mouth shape of an avatar, and animate the mouth shape of the avatar to be synchronized with the user's voice based on the inferred mouth shape.
[0261] Additionally, the exercise main screen (370) may further include an exit button (375), a microphone on / off button (376), and a volume control button (377).
[0262] As illustrated in FIG. 14, the present disclosure may provide a menu pop-up (380) that includes a first menu button (382) for moving to an app usage guide screen, a second menu button (384) for moving to an avatar setting screen, a third menu button (386) for moving to a nickname setting screen, a fourth menu button (387) for moving to a user inquiry channel, a fifth menu button (388) for moving to an open chat room, and a fifth menu button (389) for moving to a metaverse service providing website when a user request for selecting a menu item within a home screen is received, and further includes a close button (381).
[0263] As illustrated in FIG. 15, the present disclosure, when a user-captured image (392) is received, checks whether the user's full body is included in the user-captured image (392), and if the user's full body included in the user-captured image (392) is below a standard ratio, generates and provides a field of view check pop-up (390), and if the user's full body included in the user-captured image (392) is above a standard ratio, the generated field of view check pop-up (390) can be removed.
[0264] As an example, the present disclosure may generate a view angle check pop-up (390) including at least one of a user-captured image (392), a guidance text for the view angle (394), and an operation button (396) for screen operation.
[0265] In addition, the present disclosure can generate and provide a view angle check pop-up (390) when the exercise main screen is first provided.
[0266] In addition, the present disclosure can check whether the user's full body is included in the user-captured image for each frame while the exercise main screen is provided, and can generate and provide a field of view check pop-up (390) if the user's full body included in the user-captured image is below a standard ratio.
[0267] In addition, the present disclosure can determine whether the user's full body is included in the user-captured image by inputting the user-captured image into a pre-learned neural network model to infer skeleton points corresponding to the user's full body, and analyzing the skeleton points corresponding to the user's full body to determine whether the user's full body is included in the user-captured image.
[0268] In addition, when the present disclosure provides a view angle check pop-up (390), if a user-captured image having a user body ratio below a standard ratio is continuously received for a standard period of time or longer, the view angle check pop-up (390) can be generated and provided.
[0269] In some cases, when the present disclosure provides a field of view check pop-up (390), if a user-captured image having a user body ratio greater than or equal to a standard ratio is continuously received for a standard period of time, the generated field of view check pop-up (390) can be removed.
[0270] As illustrated in FIG. 16, the present disclosure may provide a volume control pop-up (400) including an exercise video volume control controller (402), a user-to-user conversation volume control controller (404), a confirmation button (408) for saving a volume control status, and a volume status display unit (406) for showing a current volume control status when a user request for selecting a volume control item within an exercise main screen is received.
[0271] As illustrated in FIG. 17, the present disclosure may provide an exit pop-up (410) including an exit button (412) that moves to the home screen when a user request to select an exit item within the exercise main screen is received, and a cancel button (414) that cancels the exit and returns to the exercise main screen.
[0272] As illustrated in FIG. 18, the present disclosure may provide an exercise authentication sharing pop-up (420) including a user avatar exercise recording video (421), an exercise video save button (426), an exercise video share button (427), and a pop-up close button when an exercise session ends.
[0273] Here, the user avatar exercise recording video (421) may include at least one of the user's exercise result information (422), the user's avatar exercise motion (423), the number of times the user exercised (425), and a specific logo (424), but this is only an example and is not limited thereto.
[0274] For example, the user's exercise result information (422) may include at least one of exercise type, exercise time, and calories consumed.
[0275] In addition, the present disclosure can record the user avatar's exercise from the time the exercise video recording request is received when the exercise video recording request is received from the user terminal while the avatar participating in the exercise session is exercising.
[0276] In some cases, the present disclosure may simultaneously record all avatars participating in an exercise session exercising as the exercise session progresses.
[0277] In another case, the present disclosure can compare the exercise movements of an avatar participating in an exercise session with a preset reference movement when the exercise session is in progress, calculate a score based on the degree of agreement, and record the avatar exercising at the time when the exercise movement with the highest score of agreement is performed.
[0278] In another case, the present disclosure can analyze the exercise movements of an avatar participating in an exercise session when the exercise session is in progress, and record the avatar exercising at the time when the exercise movement with the largest motion magnitude is performed.
[0279] In another case, the present disclosure can record an avatar exercising at a point in time when the avatar is performing exercise movements during the latter part of the exercise session as the exercise session progresses.
[0280] In another case, the present disclosure can determine the average volume of voice chat of avatars participating in an exercise session when the exercise session is in progress, and record the avatars exercising at the time when the average volume of voice chat is the loudest.
[0281] In addition, the present disclosure can extract only the motion data of the user avatar from the user avatar movement recording video and store the extracted data in a database when storing the user avatar movement recording video.
[0282] Here, the present disclosure can obtain motion data corresponding to the user avatar from a database when playback of a user avatar motion recording video is requested, reflect the motion data to the user avatar, generate a user avatar motion recording video, and play the generated user avatar motion recording video.
[0283] In addition, the present disclosure can simultaneously record a video of a user avatar participating in an exercise session exercising and a user-captured video received from the user terminal when consent to recording a user-captured video is received from the user terminal, thereby storing a video including the user avatar exercising in a virtual space and the user exercising in a real space.
[0284] As illustrated in FIG. 19, the present disclosure provides an avatar photo shooting pop-up (430) for taking an exercise authentication photo when an exercise session ends, and when a request for taking a group authentication photo is received through the avatar photo shooting pop-up, a set time (436) is counted for the avatars participating in the exercise session to take individual poses or group poses, and when the set time arrives, an authentication photo of the avatars participating in the exercise session can be taken.
[0285] For example, the avatar photo shoot pop-up (430) may include at least one of the exercise date, exercise time, exercise calories burned, exercise type, group photo take button (432), and exercise one more time button (434).
[0286] In this way, the present disclosure can provide a sense of realism, interest, and fun to the user by allowing the user and the avatar to exercise together by reflecting the user's motion data on the avatar uploaded to the virtual space of the exercise session and controlling the motion of the avatar by predicting the next movement through a pre-learned neural network model.
[0287] In the present disclosure, the use of a neural network model is only one embodiment, and avatar motion can be controlled even without a neural network model.
[0288] The method according to one embodiment of the present disclosure described above can be implemented as a program (or application) and stored in a medium to be executed in combination with a hardware server.
[0289] The above-described program may include codes coded in a computer language, such as C, C++, JAVA, or machine language, that can be read by the processor (CPU) of the computer through the device interface of the computer, so that the computer reads the program and executes the methods implemented as a program. Such codes may include functional codes related to functions that define functions necessary for executing the methods, and may include control codes related to execution procedures necessary for the processor of the computer to execute the functions according to a predetermined procedure. In addition, such codes may further include memory reference-related codes regarding which location (address address) of the internal or external memory of the computer should reference additional information or media necessary for the processor of the computer to execute the functions. In addition, if the processor of the computer needs to communicate with any other computer or server located remotely in order to execute the functions, the code may further include communication-related code regarding how to communicate with any other computer or server located remotely using the communication module of the computer, and what information or media to send and receive during communication.
[0290] The above storage medium refers to a medium that stores data semi-permanently and can be read by a device, rather than a medium that stores data for a short period of time, such as a register, cache, or memory. Specifically, examples of the storage medium include, but are not limited to, ROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage device. That is, the program can be stored in various recording media on various servers that the computer can access or in various recording media on the user's computer. In addition, the medium can be distributed across network-connected computer systems, so that computer-readable code can be stored in a distributed manner.
[0291] The steps of a method or algorithm described in connection with the embodiments of the present disclosure may be implemented directly in hardware, implemented as a software module executed by hardware, or implemented by a combination thereof. The software module may reside in a random access memory (RAM), a read only memory (ROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), a flash memory, a hard disk, a removable disk, a CD-ROM, or any other form of computer-readable recording medium well known in the art to which the present disclosure pertains.
[0292] While the embodiments of the present disclosure have been described above with reference to the attached drawings, those skilled in the art will appreciate that the present disclosure can be implemented in other specific forms without altering the technical spirit or essential features thereof. Therefore, the embodiments described above should be understood to be illustrative in all respects and not restrictive.
Claims
1. A communication unit that communicates with the user terminal; A database that stores the avatar's motion data; and Includes a control unit that controls the motion of an avatar uploaded to the virtual space of the metaverse, The above control unit, When a request for participation in an exercise session is received from the user terminal, a user-captured video is requested from the user terminal, and when a user-captured video is received from the user terminal, user motion data is extracted from the user-captured video, the user motion data is reflected in an avatar uploaded to a virtual space of the exercise session, user motion data corresponding to the current motion of the user is input into a pre-learned neural network model to predict the user's next motion, and the motion of the avatar is controlled based on the prediction result. The above control unit, When extracting the above user motion data, a pre-learned neural network model is used to extract 3D key points for joint positions from the user's body included in the user's captured video, and user motion data is extracted based on the rotation parameters corresponding to each joint key point. The above control unit, When extracting the above 3D key points, if the number of extracted key points is less than a reference ratio or the visibility of the predicted key points is less than a reference ratio, the user motion is determined to be invalid, and the update of the user motion data is stopped or the user motion data is predicted in advance and then replaced with other user motion data corresponding to the motion. The above control unit, If the rate of change in the user motion is greater than the standard rate of change in the motion while the above user-captured video is received at a constant number of frames per second, the user motion is judged to be invalid. The above control unit, A metaverse service providing device characterized in that the rotation parameters of the joints are corrected when the areas occupied by joints corresponding to the user motion are predicted to collide with each other based on the rotation parameters.
2. In paragraph 1, The above control unit, A metaverse service providing device characterized in that, when controlling the above avatar motion, real-time voice chat is connected and voice transmission and playback control is performed through a voice chat server.
3. In paragraph 1, The above control unit, A metaverse service providing device characterized in that, when controlling the motion of the avatar, if there are multiple users who have requested to participate in the exercise session, it is determined whether the sync for the same motion of the avatar of the user who requested to participate and the participant avatars of other users are identical, and if the sync for the same motion is not identical, the avatar motion is controlled to match the sync for the same motion of the avatars participating in the exercise session.
4. In paragraph 1, The above control unit, A metaverse service providing device characterized in that, when the above exercise session ends, an exercise authentication sharing pop-up is provided that includes a user avatar exercise recording video, an exercise video save button, an exercise video share button, and a pop-up close button.
5. In paragraph 4, The above control unit, A metaverse service providing device characterized in that, when saving the user avatar movement recording video, only the motion data of the user avatar is extracted from the user avatar movement recording video and stored in the database.
6. In paragraph 5, The above control unit, A metaverse service providing device characterized in that, when playback of the user avatar movement recording video is requested, motion data corresponding to the user avatar is acquired from the database, the motion data is reflected to the user avatar to generate the user avatar movement recording video, and the generated user avatar movement recording video is played.
7. In paragraph 1, The above control unit, A metaverse service providing device characterized in that when the above exercise session ends, an avatar photo shooting pop-up is provided for taking an exercise authentication photo, and when a request for taking a group authentication photo is received through the avatar photo shooting pop-up, a set time is counted for the avatars participating in the exercise session to take individual poses or group poses, and when the set time arrives, an authentication photo of the avatars participating in the exercise session is taken.
8. A method for providing a metaverse service of a device that is connected to a user terminal and communicates with the device, performed by the device's control unit. A step of receiving a request to participate in an exercise session from the user terminal; A step of requesting a user captured video to the user terminal; A step of receiving a user-captured image from the user terminal; A step of extracting user motion data from the user-captured video; A step of reflecting the user motion data to an avatar uploaded to a virtual space of the above exercise session; A step of predicting the user's next action by inputting user motion data corresponding to the user's current action into a pre-learned neural network model; and A step of controlling the motion of the avatar based on the above prediction result; The above control unit, When extracting the above user motion data, a pre-learned neural network model is used to extract 3D key points for joint positions from the user's body included in the user's captured video, and user motion data is extracted based on the rotation parameters corresponding to each joint key point. The above control unit, When extracting the above 3D key points, if the number of extracted key points is less than a reference ratio or the visibility of the predicted key points is less than a reference ratio, the user motion is determined to be invalid, and the update of the user motion data is stopped or the user motion data is predicted in advance and then replaced with other user motion data corresponding to the motion. The above control unit, If the rate of change in the user motion is greater than the standard rate of change in the motion while the above user-captured video is received at a constant number of frames per second, the user motion is judged to be invalid. The above control unit, A method for providing a metaverse service, characterized in that the rotation parameters of the joints are corrected when it is predicted that the areas occupied by joints corresponding to the user motion collide with each other based on the rotation parameters.
9. In paragraph 8, The above control unit, A method for providing a metaverse service, characterized in that when controlling the above avatar motion, real-time voice chat is connected and voice transmission and playback control is performed through a voice chat server.
10. In paragraph 8, The above control unit, A method for providing a metaverse service, characterized in that when controlling the motion of the avatar, if there are multiple users who have requested to participate in the exercise session, the sync for the same motion of the avatar of the user who requested to participate and the avatars of other users is determined to be the same, and if the sync for the same motion is not the same, the avatar motion is controlled to match the sync for the same motion of the avatars participating in the exercise session.
11. In paragraph 8, The above control unit, A method for providing a metaverse service, characterized in that when the above exercise session ends, an exercise authentication sharing pop-up is provided that includes a user avatar exercise recording video, an exercise video save button, an exercise video share button, and a pop-up close button.
12. In paragraph 11, The above control unit, A method for providing a metaverse service, characterized in that when saving the user avatar movement recording video, only the motion data of the user avatar is extracted from the user avatar movement recording video and stored in the database.
13. In paragraph 12, The above control unit, A method for providing a metaverse service, characterized in that when playback of the user avatar movement recording video is requested, motion data corresponding to the user avatar is acquired from the database, the motion data is reflected to the user avatar to generate the user avatar movement recording video, and the generated user avatar movement recording video is played.
14. In paragraph 8, The above control unit, A method for providing a metaverse service, characterized in that when the above exercise session ends, an avatar photo shooting pop-up is provided to take an exercise authentication photo, and when a request for a group authentication photo is received through the avatar photo shooting pop-up, a set time is counted for the avatars participating in the exercise session to take individual poses or group poses, and when the set time arrives, an authentication photo of the avatars participating in the exercise session is taken.
Citation Information
Patent Citations
Apparatus for predicting intention of user using multi modal information and method thereof
KR1020110009614A
Substrate processing apparatus
KR1020230019734A
Manufacturing method of packing paper
KR1020240054609A
System for providing vitual reality service related to sports contents
KR102443949B1
Clothes Dryer
KR102516277B1