Motion-based user authentication device and method therefor

The motion-based user authentication system in the metaverse uses AI models to continuously authenticate users based on avatar movements, addressing the limitations of conventional methods by enhancing security and reducing costs.

WO2026005444A1PCT designated stage Publication Date: 2026-01-02FOUND OF SOONGSIL UNIV IND COOP +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/008811
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-26
Filing Date
2025-06-24
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing user authentication methods in the metaverse, such as facial or iris recognition, are inadequate due to the challenges of user anonymity and the need for expensive equipment, and they fail to provide continuous authentication against potential threats like account theft and identity impersonation.

Method used

A motion-based user authentication system that utilizes a processor to extract posture information from an avatar's movements, normalize this data using AI models, and continuously authenticate users through a transformer encoder and auto-encoder model, eliminating the need for expensive biometric equipment.

Benefits of technology

Enhances security by preventing account theft and identity switching in the metaverse through continuous authentication based on user behavior patterns, reducing costs and improving accuracy without requiring additional hardware.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025008811_02012026_PF_FP_ABST
    Figure KR2025008811_02012026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a technology for authenticating a user on the basis of motions of an avatar. A motion-based user authentication device according to the present invention receives a motion image of an avatar controlled by a user, extracts posture information of the avatar from the received image, normalizes the posture information, summarizes the normalized posture information using a first artificial intelligence model to extract information required for user identification and identify the user, and processes the normalized posture information using a second artificial intelligence model to authenticate the user, thereby being able to continuously identify and authenticate the user in a metaverse environment.
Need to check novelty before this filing date? Find Prior Art

Description

Motion-based user authentication device and method thereof

[0001] The present invention relates to user authentication in an online environment such as a metaverse, and more particularly, to a technology that enables continuous authentication based on the movements of an avatar.

[0002] Services based on the metaverse present various security and privacy concerns. These issues are exacerbated by the unique characteristics of the metaverse. The combination of four core characteristics—socialization, immersive interaction, virtual-real-world construction, and scalability—can potentially lead to significant security and privacy breaches.

[0003] The metaverse, as it strives to create an immersive cyberspace, incorporates more personal elements, which in turn increases the risk of personal information leaks. Users provide information, such as their profiles and behavioral data, to enjoy social services, making them more vulnerable to attackers.

[0004] There is also a risk of becoming a target of crimes such as pretending to be someone else in the metaverse, providing false information, or stealing important information such as financial information.

[0005] Moreover, it has recently become possible to imitate the face or voice of a specific person using generative AI, and therefore, technology to verify that the currently logged-in user is the actual account owner is an essential technology for security in the metaverse environment.

[0006] Conventional technologies have attempted to use biometric authentication methods like facial or iris recognition for user authentication. However, in the metaverse, a user's actual face is not revealed, and iris recognition requires additional, expensive equipment, making it difficult to implement.

[0007] Additionally, if only one-time authentication is performed when accessing the metaverse using facial recognition or iris recognition, it is difficult to respond to methods such as changing users in the middle of the process.

[0008] The inventors of the present invention have dedicated extensive research efforts to addressing the aforementioned problems in prior art. They have developed a system that recognizes user behavior patterns within the metaverse, eliminating the need for conventional biometric authentication methods. This allows for continuous authentication while the user is connected to the metaverse.

[0009] In order to solve the problems of the prior art described above, the present invention aims to provide a device and method that can strengthen the security of the metaverse by continuously providing an authentication method on the metaverse.

[0010] In addition, another object of the present invention is to provide a device and method that can perform accurate authentication without using expensive equipment such as biometric authentication such as facial recognition or iris recognition.

[0011] However, the problems to be solved by the present invention are not limited to the problems mentioned above, and other problems not mentioned can be clearly understood by a person having ordinary skill in the technical field to which the present invention belongs from the description below.

[0012] In order to solve the above-described problem, a motion-based user authentication device according to a preferred embodiment of the present invention is characterized by receiving a motion image of an avatar controlled by a user, extracting posture information of the avatar from the received image, normalizing the posture information, extracting information necessary for user identification by summarizing the normalized posture information using a first artificial intelligence model, identifying the user, and calculating the normalized posture information using a second artificial intelligence model, thereby authenticating the user.

[0013] The above first artificial intelligence model is characterized by using a transformer encoder and a fully connected (FC) layer.

[0014] The above second artificial intelligence model is characterized by using an auto-encoder model.

[0015] The above detailed information is characterized by including the number of frames of the input video, joint coordinates, and confidence score.

[0016] The processor is characterized in that it continuously identifies and authenticates the user while the avatar continues to operate.

[0017] A motion-based user authentication method according to another preferred embodiment of the present invention,

[0018] The method is characterized by comprising: receiving a motion image of an avatar controlled by a user; extracting posture information of the avatar from the received image; normalizing the posture information; extracting information necessary for user identification by summarizing the normalized posture information using a first artificial intelligence model to identify the user; and authenticating the user by calculating the normalized posture information using a second artificial intelligence model.

[0019] The above first artificial intelligence model is characterized by using a transformer encoder and a fully connected layer.

[0020] The above second artificial intelligence model is characterized by using an autoencoder model.

[0021] The above detailed information is characterized by including the number of frames of the input video, joint coordinates, and confidence score.

[0022] The step of receiving the motion image of the avatar or the step of authenticating the user is characterized in that it is continuously repeated while the motion of the avatar continues.

[0023] According to the present invention, there is an effect of reducing security threats such as metaverse account theft by providing a continuous authentication method on the metaverse.

[0024] Additionally, it has the advantage of saving authentication costs by performing authentication using only avatar movement data without expensive additional hardware.

[0025] The effects that can be obtained from the present invention are not limited to the effects mentioned above, and other effects not mentioned can be clearly understood by a person having ordinary skill in the art to which the present invention belongs from the description below.

[0026] FIG. 1 is a schematic structural diagram of a motion-based user authentication device according to a preferred embodiment of the present invention.

[0027] Figure 2 is a schematic flowchart of a feature data normalization method according to a preferred embodiment of the present invention.

[0028] Figure 3 shows the difference in user identification accuracy according to a normalization method according to a preferred embodiment of the present invention.

[0029] Figure 4 is a schematic flowchart of a user identification method according to a preferred embodiment of the present invention.

[0030] Figure 5 is a schematic structural diagram of an autoencoder model according to a preferred embodiment of the present invention.

[0031] FIG. 6 is a schematic flowchart of a motion-based user authentication method according to another preferred embodiment of the present invention.

[0032] The above-described objects, means, and resulting effects of the present invention will become more apparent through the following detailed description, taken in conjunction with the accompanying drawings. Accordingly, those skilled in the art will be able to readily implement the technical concepts of the present invention. Furthermore, in describing the present invention, if a detailed description of known technology related to the present invention is deemed to unnecessarily obscure the gist of the invention, such detailed description will be omitted.

[0033] The terminology used herein is for the purpose of describing embodiments and is not intended to limit the present invention. In this specification, the singular also includes the plural, unless specifically stated otherwise. In this specification, terms such as "include," "provide," "provide," or "have" do not exclude the presence or addition of one or more other components other than the mentioned components.

[0034] In this specification, terms such as “or,” “at least one,” and the like can refer to one of the words listed together, or a combination of two or more. For example, “or B” or “and at least one of B” can include only one of A or B, or can include both A and B.

[0035] In this specification, descriptions using the phrase “for example” or the like should not be construed as limiting the embodiments of the invention in terms of the effects of variations such as tolerances, measurement errors, limitations of measurement accuracy, and other commonly known factors, including the information presented, such as cited characteristics, variables, or values, which may not be exact matches.

[0036] In this specification, when a component is described as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components in between. Conversely, when a component is described as being "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between.

[0037] In this specification, when a component is described as being "on" or "in contact with" another component, it should be understood that it may be directly on or connected to the other component, but there may be another component in between. Conversely, when a component is described as being "directly on" or "in direct contact with" another component, it should be understood that there is no other component in between. Other expressions that describe the relationship between components, such as "between" and "directly between", can be interpreted similarly.

[0038] In this specification, terms such as "first" and "second" may be used to describe various components, but the components should not be limited by these terms. Furthermore, these terms should not be construed to limit the order of each component, but rather may be used to distinguish one component from another. For example, a "first component" may be referred to as a "second component," and similarly, a "second component" may also be referred to as a "first component."

[0039] Unless otherwise defined, all terms used herein may be used in their common sense by those of ordinary skill in the art to which the present invention pertains. Furthermore, terms defined in commonly used dictionaries are not to be interpreted ideally or excessively unless explicitly and specifically defined otherwise.

[0040]

[0041] Hereinafter, a preferred embodiment according to the present invention will be described in detail with reference to the attached drawings.

[0042] FIG. 1 is a schematic structural diagram of a motion-based user authentication device according to a preferred embodiment of the present invention.

[0043] A motion-based user authentication device (100) according to the present invention may include one or more processors (110) and memories (120).

[0044] The memory (120) may store instructions, data structures, and program codes that can be read by the processor (110). In embodiments, at least the operations performed by the processor (110) may be implemented by executing instructions or codes of the program stored in the memory (120).

[0045] The memory (120) may include a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), and may include a non-volatile memory including at least one of a ROM (Read-Only Memory), an EEPROM (Electrically Erasable Programmable Read-Only Memory), a PROM (Programmable Read-Only Memory), a magnetic memory, a magnetic disk, and an optical disk, and a volatile memory such as a RAM (Random Access Memory) or an SRAM (Static Random Access Memory).

[0046] The memory (120) can store one or more instructions or programs that the motion-based user authentication device (100) can use to recognize the motion of an avatar and identify and authenticate a user from it.

[0047] The processor (110) controls the overall operations of the motion-based user authentication device (100). For example, the processor (110) can control the overall operations of the motion-based user authentication device (100) to identify the behavioral pattern of an avatar within the metaverse environment and perform user identification and authentication using an artificial intelligence model by executing one or more commands stored in the memory (120).

[0048] The processor (110) may be configured as at least one of, for example, a central processing unit, a microprocessor, a graphic processing unit, an application specific integrated circuits (ASICs), a digital signal processor (DSPs), a digital signal processing device (DSPDs), a programmable logic device (PLDs), a field programmable gate array (FPGAs), an application processor, a neural processing unit, or an artificial intelligence processor designed with a hardware structure specialized for processing an artificial intelligence model, but is not limited thereto.

[0049] The processor (110) first extracts features from the avatar's movements for user identification and authentication. An avatar is a character representing a user in a virtual space that operates under the user's control, such as through user input. The avatar's movements may reflect the user's intentions. For example, the avatar may be a character in a metaverse environment that reflects the user's movements through a VR device. For example, the avatar's posture can be estimated from the movements of an avatar performing a game in a virtual space such as the metaverse.

[0050] Open source pose estimation models such as OpenPose can be used for pose extraction, but are not limited to these.

[0051] Information extracted from an avatar's motion video may have the form (T, A, B), for example.

[0052] T represents the number of frames in the video, A represents the number of joints, and B represents the number of features.

[0053] For example, an output of the form (T, 25, 3) could be pose information with three features for T frames and 25 joints. B could consist of, for example, the x-coordinate of the joint, the y-coordinate, and a confidence score.

[0054] The confidence score represents the accuracy with which the pose estimation model predicts the location of the joint. It is expressed as a value between 0 and 1, which quantifies the probability that the joint location predicted by the pose estimation model and the actual joint exist at that location.

[0055] If the confidence score is included in the features of the avatar's pose information, the pose estimation model can be trained by taking into account cases where the joint coordinates are measured inaccurately due to the joint being obscured in the image, such as when the hand is behind the back or the knee is obscured by the hand.

[0056] The confidence score can be defined as follows:

[0057]

[0058]

[0059]

[0060] Table 1 below shows the user identification accuracy with and without confidence scores. Accuracy is higher when confidence scores are included in the avatar's posture information features (with confidence).

[0061] window sizew confidencew / o confidence30079.81%77.41%10080.96%76.48%

[0062]

[0063]

[0064] The processor (110) uses the extracted detail information as the final user characteristic through a normalization process.

[0065] The normalization process is designed to minimize the influence of the avatar's initial position on the authentication result, thereby better understanding the essential characteristics of the avatar's movements and improving authentication accuracy.

[0066] The normalization process standardizes the initial positions of each user, allowing the AI ​​model to more accurately distinguish changes in the avatar's movement. Because features extracted through the pose estimation model have different values ​​depending on the user's initial starting position, there's a problem with classifying users based on the avatar's position rather than its behavior, i.e., its movement.

[0067] To solve this problem, features are extracted from the detailed information and then normalized.

[0068] Figure 2 is a schematic flowchart showing a normalization process according to the present invention.

[0069] The processor (110) first loads feature information extracted from the avatar's motion data (S210) and reshapes the feature data (S220).

[0070] The extracted feature information has the form of a (T, 25, 3)-dimensional matrix, for example, since there are 25 joints and 3 features for the T frame.

[0071] To view this as a concept of the same line, 3 features for 25 joints for each frame, i.e. 25*3=75 features, the dimension of the matrix is ​​converted to (T, 75).

[0072] Reshape means changing only the shape of the matrix while keeping the values ​​of the matrix intact.

[0073] A normalization model is selected to normalize the following feature data (S230).

[0074] Normalization models can include temporal normalization, joint normalization, and mixed normalization.

[0075] The time normalization model normalizes based on the first frame of the authentication unit data.

[0076] The joint normalization model normalizes based on the center point of the pose that appears in each frame.

[0077] Mixed normalization performs normalization by concatenating the data derived from time normalization and the data derived from joint normalization.

[0078] Finally, the processor (110) normalizes the feature data using the selected normalization model and stores it (S240).

[0079] Figure 3 shows the accuracy of the normalization results obtained in this manner.

[0080] Compared to time normalization, joint normalization and mixed normalization can be confirmed to maintain performance even when the authentication unit length is shortened.

[0081] This is because time normalization expresses how much movement has occurred since the first frame, so when the authentication unit length is short, it is difficult to capture all of the avatar's movement in one authentication unit.

[0082] Additionally, joint normalization shows the highest performance when the authentication unit length is the same, because joint normalization expresses how far apart it is from the center joint, so information about a person's posture is contained in a single token, allowing for more accurate modeling of behavior.

[0083] Therefore, the highest accuracy can be achieved by using joint normalization as a regularization model.

[0084] After completing normalization, the processor (110) uses the normalized feature data to perform user identification and authentication.

[0085] Figure 4 is a schematic flowchart of a user identification method according to the present invention.

[0086] The first artificial intelligence model can be used for user identification.

[0087] The second artificial intelligence model can use a transformer model and a fully connected layer.

[0088] The processor (110) loads feature data (S410) and extracts key information therefrom (S420).

[0089] Key information extraction can be achieved using, but is not limited to, the encoder structure of the Transformer model, the first AI model. Transformer encoders have the advantage of performing better than decoders on tasks such as classification and regression.

[0090] Therefore, when feature information is input into the encoder of the transformer model, the encoder summarizes the feature information and outputs it as key information necessary for user identification.

[0091] The summarized core information is input into the fully connected layer to perform user identification (S430).

[0092] The parameters used in the encoder are as follows:

[0093] - Encoder layer: 8

[0094] - Attention head: 8

[0095] - Model dimension: 64

[0096] - Learning rate: 1e -3

[0097] By this process, the processor (110) provides an accurate and efficient method for user identification and can accurately analyze the behavior of the avatar.

[0098] Finally, the processor (110) performs user authentication using feature data.

[0099] A pre-trained second artificial intelligence model can be used for user authentication.

[0100] The second artificial intelligence model may be, but is not limited to, an auto-encoder model.

[0101] Figure 5 shows a schematic structure of an autoencoder model.

[0102] The autoencoder model (500) is structured with an input unit (510), an encoder (520), a decoder (530), and an output unit (540).

[0103] Data input to the input unit (510) is dimensionally reduced by the encoder (520) to extract core data, and the core data is dimensionally expanded by the decoder (530) to output restored feature data by the output unit (540).

[0104] That is, since the input data and output data have the same dimension, it can be trained to reduce the error value (Least Square Error, LSE) between the input data and the output data.

[0105] In addition, the autoencoder model (500) is an unsupervised learning model, so it does not require a label, and thus has the characteristic of being able to build an effective authentication model using only the user's normal behavior data.

[0106] In this way, the second AI model is structured to learn specific user behavior patterns and authenticate their identity based on these patterns. This is because the autoencoder model can efficiently learn and extract key features from input data.

[0107] When the processor (110) inputs the user's feature data as input to the second artificial intelligence model, if the LSE value between the input feature information and the restored feature information is less than a threshold value, the processor can determine that the user is the user and perform authentication.

[0108] The processor (110) can continuously perform user identification and authentication as described above while the user continues to perform actions. Therefore, compared to conventional authentication methods that only perform authentication at certain points in time, this provides the advantage of further enhancing security.

[0109]

[0110] FIG. 6 is a schematic flowchart of a motion-based user authentication method according to another preferred embodiment of the present invention.

[0111] The motion-based user authentication method according to the present invention can be performed by a motion-based user authentication device including one or more processors and memory.

[0112] First, a user's motion video is received (S110), and from this, the avatar's posture information can be extracted (S120).

[0113] Extraction of posture information can be done using open source posture estimation models such as OpenPose, but is not limited thereto.

[0114] Information extracted from the avatar's motion video can be in the form of (T, A, B), and the detailed information is as described above.

[0115] The posture information extracted from the avatar motion video is normalized and used as user characteristic information (S130).

[0116] The normalization process standardizes the initial positions of each user, allowing the AI ​​model to more accurately distinguish changes in the avatar's movements. However, features extracted through the pose estimation model have different values ​​depending on the user's initial position, leading to the problem of classifying users based on the avatar's position rather than its behavior (i.e., movement). Therefore, the normalization process is implemented to address this issue.

[0117] Normalization can use temporal normalization, joint normalization, and mixed normalization models, among which joint normalization can achieve the highest accuracy.

[0118] After normalization is complete, user identification is performed using the normalized feature data (S140).

[0119] User identification uses a first artificial intelligence model, which may include, but is not limited to, a transformer model and a fully connected layer.

[0120] The encoder of the transformer model summarizes feature information into core information, and by inputting the summarized information into the fully connected layer, the user can be identified.

[0121] Finally, the processor performs user authentication using feature data (S150).

[0122] User authentication may be performed by a second artificial intelligence model, which may be, but is not limited to, an auto-encoder model.

[0123] The autoencoder model can be trained to reduce the least square error (LSE) between the input and output data because the input and output data have the same dimension.

[0124] When the user's feature data is input to the second artificial intelligence model using the learned autoencoder model, if the LSE value between the input feature information and the restored feature information is less than the threshold, it can be determined that the user is the user and authentication can be performed.

[0125] The processor may repeatedly perform the steps from receiving motion images (S110) to performing user authentication (S150) while the user's actions continue. This enhances the security of user authentication compared to performing user identification and authentication only at specific points in time.

[0126]

[0127] In this way, according to the motion-based user authentication device and method according to the present invention, by continuously extracting characteristic information from the motion of an avatar reflecting the user's intention and performing user authentication based on the extracted characteristic information, problems such as user information hacking or user switching can be prevented, thereby enhancing security in the metaverse environment.

[0128]

[0129] While the detailed description of the present invention has described specific embodiments, it should be understood that various modifications are possible without departing from the scope of the present invention. Therefore, the scope of the present invention is not limited to the described embodiments, but should be defined by the claims and their equivalents.

Claims

1. Memory containing one or more instructions; and A processor that executes one or more instructions stored in the memory; Including, but not limited to, The above processor, Receive motion video of the avatar controlled by the user, Extracting the posture information of the avatar from the received image, Normalize the above detailed information, The above normalized detailed information is summarized using the first artificial intelligence model to extract information necessary for user identification and identify the user. A motion-based user authentication device characterized in that it authenticates a user by calculating the normalized posture information using a second artificial intelligence model.

2. In paragraph 1, A motion-based user authentication device, characterized in that the first artificial intelligence model uses a transformer encoder and a fully connected (FC) layer.

3. In paragraph 1, A motion-based user authentication device, characterized in that the second artificial intelligence model uses an auto-encoder model.

4. In paragraph 1, A motion-based user authentication device, characterized in that the above-mentioned detailed information includes the number of frames of an input video, joint coordinates, and a confidence score.

5. In paragraph 1, A motion-based user authentication device, characterized in that the processor continuously identifies and authenticates the user while the avatar continues to move.

6. A motion-based user authentication method performed by a motion-based user authentication device including one or more processors and memories: A step of receiving a motion image of an avatar controlled by a user; A step of extracting posture information of the avatar from the received image; A step of normalizing the above detailed information; A step of extracting information necessary for user identification by summarizing the above normalized detailed information using a first artificial intelligence model and identifying the user; and A step of authenticating a user by calculating the normalized posture information using a second artificial intelligence model; A motion-based user authentication method, characterized in that it includes:

7. In paragraph 6, A motion-based user authentication method, characterized in that the first artificial intelligence model uses a transformer encoder and a fully connected layer.

8. In paragraph 6, A motion-based user authentication method, characterized in that the second artificial intelligence model uses an autoencoder model.

9. In paragraph 6, A motion-based user authentication method, characterized in that the above-mentioned detailed information includes the number of frames of an input video, joint coordinates, and a confidence score.

10. In paragraph 6, A motion-based user authentication method, characterized in that the step of receiving the motion image of the avatar or the step of authenticating the user are continuously repeated while the motion of the avatar continues.

Citation Information

Patent Citations

  • Method for user vertification using face recognition and head pose estimation and apparatus thereof

    KR101464446B1

  • System for access control using hand gesture cognition, method thereof and computer recordable medium storing the method

    KR1020160101228A

  • Authentication method of mobile phone using implicit authentication

    KR102081266B1

  • Realtime Pose recognition system using artificial intelligence and recognition method

    KR102369152B1

  • Information processing device, information processing method, and computer-readable storage medium

    WO2023032170A1