Interactive 3D human motion generation system using artificial intelligence

The conversational 3D human motion generation system uses AI to infer and verify motion vectors from instruction sentences, addressing the inefficiencies in existing methods by automating the creation and modification of 3D character movements.

WO2026089095A1PCT designated stage Publication Date: 2026-04-30CREADTO INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/016519
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-10-25
Filing Date
2024-10-28
Publication Date
2026-04-30

AI Technical Summary

Technical Problem

Existing methods for generating 3D character animations in the game and entertainment industries require manual parameter-by-parameter modification to create new movements, lacking an efficient system for creating and modifying human model movements based on instruction sentences.

Method used

A conversational 3D human motion generation system using artificial intelligence that infers motion data from instruction sentences, generates motion vectors using generative AI, and verifies these vectors on a 3D human model, allowing for precise adjustment and increased diversity of movements.

Benefits of technology

Enables the creation of diverse and precise 3D human movements by automatically generating and verifying actions based on natural language inputs, reducing the need for manual parameter adjustment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024016519_30042026_PF_FP_ABST
    Figure KR2024016519_30042026_PF_FP_ABST
Patent Text Reader

Abstract

Provided is an interactive 3D human motion generation system using artificial intelligence, the system comprising: a user terminal that inputs a natural language-based instruction text to generate 3D human motion and outputs a 3D human whose motion is generated on the basis of the instruction text; and a generation service-providing server including a receiving unit that receives the instruction text from the user terminal, a determination unit that infers and defines the motion included in the instruction text by dividing the motion by body parts, a generation unit that generates a motion vector on the basis of generative AI to animate the body parts on the basis of the motion defined for each body part, and a verification unit that verifies the motion based on the instruction text by applying the motion vector to a pre-built 3D human model.
Need to check novelty before this filing date? Find Prior Art

Description

Conversational 3D human motion generation system using artificial intelligence

[0001] The present invention relates to an interactive 3D human motion generation system using artificial intelligence, and provides a system that, when an instruction sentence is input, infers the motion within the instruction sentence by body part, generates a motion vector with the inferred motion, and verifies it by testing it on a 3D human model.

[0002] In the game and entertainment industries, which are growing alongside the rapid advancement of digital technology, 3D character animation is providing a new dimension of realism and creativity. Recently, AI platform services have been increasing exponentially; Generative Artificial Intelligence (GAI) is an AI technology that utilizes data to generate results based on user requirements, having emerged through the development of AI—a technology where machines perform intelligent tasks similar to those of humans. Driven by advancements in machine learning and deep learning, this AI technology is being utilized across various service sectors and has achieved significant progress over the past few years. With the continuous development of computer technology and graphics, an increasing number of game, film, and animation production companies have begun adopting technologies that automatically generate 3D character models and animations.

[0003] At this time, methods for controlling movement in a 3D model or generating motion in a 3D model have been researched and developed. In this regard, prior art Korean Published Patent No. 2023-0135449 (published September 25, 2023) and Korean Published Patent No. 2024-0013610 (published January 30, 2024) each disclose a configuration in which, after capturing the motion of a user, each body part of an object corresponding to the user is tracked by a skeleton model and reflected on a virtual character, and a configuration in which an object within the video is extracted, a person character corresponding to the object is created, the motion of the object within the video is recognized based on a 3D mesh to extract motion data, and retargeting is performed to reflect the motion data on the person character.

[0004] However, both the former and the latter only initiate the process of digitizing human movements to convert them into characters, and do not initiate the creation of new movements based on already created movements. Furthermore, even if a new movement for a 3D model is painstakingly created, the operator must go through the process of modifying the 3D data on a parameter-by-parameter basis until the desired movement is actually achieved. Accordingly, research and development of a system capable of creating and modifying the movements of a 3D human model based on instruction sentences is required.

[0005] One embodiment of the present invention provides a conversational 3D human motion generation system using artificial intelligence, wherein when a command sentence is input from a user terminal, the action of the command sentence is divided by body part, motion data for each body part is inferred based on a CLIP (Contrastive Language-Image Pre-training) model, motion vectors are generated using a generative AI to generate actions based on the motion data, and then the motion vectors are tested against a pre-established 3D human model to verify whether there are any abnormalities in the generated actions. The verified motion vectors are then mapped to the command sentence and provided to the user terminal. Furthermore, by utilizing a model that remembers the context of the conversation, various poses or postures can be assumed even within a single action, thereby allowing for precise adjustment of the actions and simultaneously increasing the diversity of the motion data. However, the technical problem that this embodiment aims to solve is not limited to the technical problem described above, and other technical problems may exist.

[0006] As a technical means for achieving the technical task described above, one embodiment of the present invention includes a user terminal that inputs a command sentence based on natural language to generate 3D human motion and outputs a 3D human with generated motion based on the command sentence, a receiving unit that receives the command sentence from the user terminal, a discrimination unit that infers and defines motions included in the command sentence by dividing them by body part, a generation unit that generates motion vectors based on generative AI to move body parts based on motions defined by body part, and a verification unit that verifies motions based on the command sentence by applying motion vectors to a pre-constructed 3D human model.

[0007] According to any one of the means for solving the problem of the present invention described above, when a command sentence is input at a user terminal, the action of the command sentence is divided by body part, the action data for each body part is inferred based on a CLIP (Contrastive Language-Image Pre-training) model, and a motion vector is generated using a generative AI to generate an action based on the action data, and then the action vector is tested on a pre-established 3D human model to verify whether there is any abnormality in the generated action, and the verified motion vector is mapped to the command sentence and provided to the user terminal, and by using a model that remembers the context of the conversation, various poses or postures can be taken even in a single action, thereby allowing for precise adjustment of the action and increasing the diversity of the action data.

[0008] FIG. 1 is a diagram illustrating an interactive 3D human motion generation system using artificial intelligence according to an embodiment of the present invention.

[0009] Figure 2 is a block diagram illustrating a creation service providing server included in the system of Figure 1.

[0010] FIGS. 3 and FIGS. 4 are drawings for illustrating an embodiment in which an interactive generation service according to an embodiment of the present invention is implemented.

[0011] FIG. 5 is an operation flowchart illustrating a method for providing an interactive generation service according to an embodiment of the present invention.

[0012] Embodiments of the present invention are described below in detail with reference to the attached drawings so that those skilled in the art can easily implement the invention. However, the present invention may be embodied in various different forms and is not limited to the embodiments described herein. Furthermore, in order to clearly explain the present invention in the drawings, parts unrelated to the explanation have been omitted, and similar parts throughout the specification are denoted by similar reference numerals.

[0013] Throughout the specification, when a part is described as being "connected" to another part, this includes not only cases where they are "directly connected" but also cases where they are "electrically connected" with other elements interposed between them. Furthermore, when a part is described as "including" a component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components, and it should be understood that this does not preclude the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0014] Terms such as “about,” “substantially,” etc., used throughout the specification, are used to mean at or near the stated value when inherent manufacturing and material tolerances are presented in the stated meaning, and are used to prevent unscrupulous infringers from unfairly exploiting the disclosure in which precise or absolute values ​​are mentioned to aid in understanding the invention. Terms such as “step” or “step of” used throughout the specification of the invention do not mean “step for”.

[0015] In this specification, the term "part" includes a unit realized by hardware, a unit realized by software, and a unit realized using both. Additionally, one unit may be realized using two or more pieces of hardware, and two or more units may be realized by one piece of hardware. Meanwhile, "part" is not limited to software or hardware, and "part" may be configured to reside in an addressable storage medium or configured to run on one or more processors. Accordingly, as an example, "part" includes components such as software components, object-oriented software components, class components, and task components, as well as processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. The functions provided within the components and "parts" may be combined into a smaller number of components and "parts" or further separated into additional components and "parts." In addition, the components and '~parts' may be implemented to play one or more CPUs within the device or secure multimedia card.

[0016] Some of the operations or functions described herein as being performed by a terminal, device, or device may instead be performed by a server connected to said terminal, device, or device. Likewise, some of the operations or functions described as being performed by a server may also be performed by a terminal, device, or device connected to said server.

[0017] In this specification, some of the operations or functions described as mapping or matching with a terminal may be interpreted as meaning mapping or matching the terminal's unique number or personal identification information, which is the terminal's identifying data.

[0018] The present invention will be described in detail below with reference to the attached drawings.

[0019] FIG. 1 is a diagram illustrating an interactive 3D human motion generation system using artificial intelligence according to an embodiment of the present invention. Referring to FIG. 1, the interactive 3D human motion generation system using artificial intelligence (1) may include at least one user terminal (100) and a generation service providing server (300). However, since the interactive 3D human motion generation system using artificial intelligence (1) of FIG. 1 is merely an embodiment of the present invention, the present invention is not to be interpreted as being limited through FIG. 1.

[0020] At this time, each component of FIG. 1 is generally connected through a network (Network, 200). For example, as shown in FIG. 1, at least one user terminal (100) can be connected to a generation service provider server (300) through the network (200).

[0021] Here, a network refers to a connection structure capable of exchanging information among individual nodes, such as multiple terminals and servers. Examples of such networks include Local Area Networks (LANs), Wide Area Networks (WANs), the World Wide Web (WWW), wired and wireless data networks, telephone networks, and wired and wireless television networks. Examples of wireless data communication networks include, but are not limited to, 3G, 4G, 5G, 3GPP (3rd Generation Partnership Project), 5GPP (5th Generation Partnership Project), 5G NR (New Radio), 6G (6th Generation of Cellular Networks), LTE (Long Term Evolution), WIMAX (World Interoperability for Microwave Access), Wi-Fi, Internet, LAN (Local Area Network), Wireless LAN (Wireless Local Area Network), WAN (Wide Area Network), PAN (Personal Area Network), RF (Radio Frequency), Bluetooth network, NFC (Near-Field Communication) network, satellite broadcasting network, analog broadcasting network, DMB (Digital Multimedia Broadcasting) network, etc.

[0022] In the following, the term "at least one" is defined as a term including both singular and plural forms, and it will be obvious that even if the term "at least one" does not exist, each component may exist in a singular or plural form and may mean singular or plural. Furthermore, whether each component is provided in a singular or plural form may be changed according to the embodiment.

[0023] At least one user terminal (100) may be a terminal of a user that inputs a command sentence using a web page, app page, program, or application related to an interactive generation service and outputs a 3D human corresponding to the command sentence.

[0024] Here, at least one user terminal (100) may be implemented as a computer capable of connecting to a remote server or terminal via a network. Here, the computer may include, for example, a navigation system, a laptop equipped with a web browser, a desktop, a laptop, etc. At this time, at least one user terminal (100) may be implemented as a terminal capable of connecting to a remote server or terminal via a network. At least one user terminal (100) may include all kinds of handheld-based wireless communication devices, such as navigation, PCS (Personal Communication System), GSM (Global System for Mobile communications), PDC (Personal Digital Cellular), PHS (Personal Handyphone System), PDA (Personal Digital Assistant), IMT (International Mobile Telecommunication)-2000, CDMA (Code Division Multiple Access)-2000, W-CDMA (W-Code Division Multiple Access), Wibro (Wireless Broadband Internet) terminal, smartphone, smartpad, tablet PC, etc.

[0025] The generation service providing server (300) may be a server that provides an interactive generation service web page, app page, program, or application. Additionally, the generation service providing server (300) may be a server that divides the actions of the instruction sentences entered by the user terminal (100) by body part and infers the action data for each body part based on a CLIP (Contrastive Language-Image Pre-training) model. Additionally, the generation service providing server (300) may be a server that infers the action data of uninstructed body parts based on an action database and then integrates each action data to generate integrated data. Additionally, the generation service providing server (300) may be a server that generates action vectors based on the integrated data and performs verification by testing (simulating) the abnormalities of the generated action vectors on a pre-established 3D human model. Furthermore, the generation service providing server (300) may be a server that maps the verified action vectors and instruction sentences, stores them in an action database, and transmits them to be visualized on the user terminal (100).

[0026] Here, the generation service providing server (300) may be implemented as a computer capable of connecting to a remote server or terminal via a network. Here, the computer may include, for example, a navigation system, a laptop equipped with a web browser, a desktop, a laptop, etc.

[0027] FIG. 2 is a block diagram for explaining a creation service providing server included in the system of FIG. 1, and FIG. 3 and FIG. 4 are drawings for explaining an embodiment in which an interactive creation service according to an embodiment of the present invention is implemented.

[0028] Referring to FIG. 2, the generation service providing server (300) may include a receiving unit (310), a discrimination unit (320), a generation unit (330), a verification unit (340), an operation DB construction unit (350), an interpolation unit (360), a visualization unit (370), and a context memory unit (380).

[0029] When a generation service providing server (300) or another server (not shown) operating in conjunction with a generation service according to one embodiment of the present invention transmits an interactive generation service application, program, app page, web page, etc. to at least one user terminal (100), the at least one user terminal (100) may install or open the interactive generation service application, program, app page, web page, etc. Additionally, a service program may be executed on at least one user terminal (100) using a script executed in a web browser. Here, a web browser refers to a program that enables the use of web (WWW: World Wide Web) services and means a program that receives and displays hypertext described in HTML (Hyper Text Mark-up Language), and includes, for example, Chrome, Microsoft Edge, Safari, Firefox, Whale, UC Browser, etc. Additionally, an application refers to an application program on a terminal, and includes, for example, an app executed on a mobile terminal (smartphone).

[0030] Referring to FIG. 2, the receiver (310) can receive instruction sentences from the user terminal (100). The user terminal (100) can input instruction sentences based on natural language to generate 3D human actions. For example, it inputs a sentence that instructs an action to be performed by the 3D human. For example, to make the 3D human perform the action of sitting and typing, the instruction sentence "[sit and type]" can be input. Or, to make the 3D human perform the action of walking a dog, the instruction sentence "[walk with a dog]" can be input.

[0031] The discriminant unit (320) can infer and define actions included in the instruction sentence by dividing them by body part. If there is an entity name referring to a body part within the instruction sentence, it can be used; otherwise, a natural language processing model such as a CLIP model can be used. That is, when the instruction sentence [walk with a dog] described above is input, in the case of a Human, one can imagine a dog wearing a harness, an owner holding a leash in one hand, a connection between the dog and the owner by a leash, and the owner walking, but in the case of a Machine, this is not the case. Accordingly, the 3D human corresponding to the owner must be described one by one as holding a leash in one hand and walking. At this time, Human is merely a word corresponding to the Machine and does not refer to the 3D human described above. To this end, in one embodiment of the present invention, the instruction sentence is input into a CLIP model to obtain information on which [body part] the action [walk with a dog] is performed. To this end, in addition to the CLIP model described above, a Text-to-Text generative AI may be used to output the body part where a word is performed when a word is input.

[0032] Additionally, the discrimination unit (320) can divide the actions included in the instruction sentence into body parts such as the face, upper body, hands, and lower body, designate the body part where the action must be performed, and infer the action data for each body part of the instruction sentence based on a CLIP (Contrastive Language-Image Pre-training) model. At this time, the CLIP model may be fine-tuned based on an action database according to an embodiment of the present invention. The action database stores [word]-[body part]-[action data] by mapping them, and the [action vector] to be described later, which is created when the user finally completes the action, is also mapped to [word]-[body part]-[action data], thereby finally constructing a database of [word]-[body part]-[action data]-[action vector]. This will be explained in detail below.

[0033] <clip>

[0034] The CLIP model is a neural network architecture developed by OpenAI that enables computers to process human language and images, such as understanding natural language and implementing computer vision. This CLIP model combines the Vision Transformer (ViT) and the Transformer-Based Language Model to process both images and text. The Transformer-Based Language Model is a model that has been trained on text data through pre-training. In summary, CLIP is a deep learning-based model based on the relationship between text and images. OpenAI's DALL·E is a tool created using this CLIP model. In one embodiment of the present invention, it is used text-based rather than multimodally, but it does not completely exclude multimodal use.

[0035] Once it is determined how the instruction sentences move each body part, the movements must be divided into those for each body part based on this information and expressed as motion data; at this stage, the aforementioned CLIP model can be used again. In this way, motion data regarding what movements should be performed for each body part is created. For example, referring to instruction sentence 1 in Fig. 3c, for the instruction sentence [clench fist], the body part is [hand], and the motion data is [the motion of clenching a fist]. Also, for the instruction sentence [stretch], the body parts are [upper body], [hand], and [lower body], and the motion data is [upper body]-[the motion of extending the hand upward], [hand]-[the motion of clenching a fist], and [lower body]-[the motion of stretching the leg].

[0036] [Table 1]

[0037]

[0038] The determination unit (320) can select motion data for each body part that is not designated as an instruction sentence based on the correlation between motion data for each body part designated based on the motion database. According to one embodiment of the present invention, the body parts are divided into four parts: [face]-[upper body]-[hand]-[lower body]. At this time, looking at Table 1, motion data exists for [upper body], [hand], and [lower body], but there is no motion data describing the motion (expression) of [face]. At this time, if a blank face is left as is when stretching, it does not blend with the rest of the body and the face floats awkwardly, so this face must also be generated to match the stretching motion.

[0039] [Table 2]

[0040]

[0041] [Table 3]

[0042]

[0043] Generally, when a person stretches, tension is applied to the face, and a grimace is bound to appear. The purpose is to determine how closely this facial expression relates to motion data of the rest of the body. In this context, correlation is a statistical concept that indicates the association between two variables, representing how one variable changes when the other changes. To continue citing the example mentioned above, the goal is to determine the correlation between the fact that stretching causes the face to grimace. For instance, if stretching results in a grimace with approximately 80% probability and a calm face with 10% probability, then the [grimace] is more correlated with stretching than the [calm] expression; in one embodiment of the present invention, the [grimace] expression is selected rather than the [calm] expression. At this time, the criterion for determining the correlation may be a motion database. This motion database may be a database in which [word-motion data] are mapped and stored.

[0044] The determination unit (330) can generate integrated data by integrating motion data for designated body parts and motion data for undesignated body parts. In this way, motion data for body parts designated by the user's instruction sentence and motion data for undesignated body parts are combined to obtain full body motion data (integrated data) for [face]-[upper body]-[hand]-[lower body].

[0045] The generation unit (330) can generate motion vectors based on generative AI to move body parts based on movements defined for each body part. In this case, a motion vector refers to a vector representing the distance moved when an object moves. For example, motion vectors are used when tracking or performing compensation on the movement of an object, and can represent the difference between position coordinates in a time-series manner along with direction. Such motion vectors are necessary to enable 3D humans to move in animation. To generate such motion vectors, the generative AI must be fine-tuned in advance to generate vector data called motion vectors from text data called motion data. To this end, a dataset of [motion data - motion vector] can be constructed in the motion database. In addition, as described above, as users continue to generate 3D human movements, [instruction sentence-motion vector] which is the learning material accumulates, and as the dataset for the generative AI to learn becomes richer and more diverse through the linkage of the [instruction sentence-word-body part-motion data-motion vector] dataset, the platform of the present invention does not need to spend separate manpower and time to build a database for motion vectors on its own.

[0046] The verification unit (340) can verify actions based on instructions by applying motion vectors to a pre-established 3D human model. At this time, the 3D human model may be a 3D human of the user terminal (100) or a standard 3D model created for actual testing. Both are 3D models, but the two terms are named differently to distinguish between the model of the user terminal (100) and the model of the already established system (1). Of course, the model of the user terminal (100) is a 3D model made from a user's photo, and the standard 3D model of the system (1) may be a 3D model like Fig. 4m. At this time, for a detailed method of creating one's own 3D human in the user terminal (100), refer to the applicant's Korean Registered Patent No. 10-2559717 (published July 26, 2023) and Korean Registered Patent No. 10-2713027 (published October 7, 2024). The present invention can be linked with the applicant's prior registered patent. If a 3D human is created on a user terminal (100) using a user's photo (1-Stage), the 3D human can be made to move using a motion vector generated according to one embodiment of the present invention (2-Stage).

[0047] The verification unit (340) can determine that a motion is abnormal if, after applying a motion vector to a 3D human model, the motion includes an angle that is not possible with the human joint angle, or if the physical collision caused by the motion is a collision that cannot be generated by the physical simulation. In addition, it can also check whether there are collisions between objects when performing 3D modeling, in addition to such physical collisions. For example, when creating a [clapping motion], an animation should not be created where the hand overlaps with the torso and the hand appears to be stuck in the torso. Also, when creating a [walking motion], the gait or arm movement should not be created where the motion looks like a zombie or monster walking, as the gait or arm movement exceeds the range of motion of the joints. In this case, the physical simulation can utilize conventional software such as Unity or Unreal, which provides game and animation tools, and various other known technologies can be applied.

[0048] The action DB construction unit (350) can construct an action database by creating a database of body parts where an action of at least one word is performed and action data for at least one word. Through this, it is possible to determine which body part is being performed even if a word designating a body part is not entered. For example, in the instruction sentence "[do a squat]," there is no entity designating a body part anywhere. In order to distinguish the body parts here, the CLIP model must be able to identify the body parts "[upper body]," "[hand]," and "[lower body]" from the word "[squat]." Accordingly, one embodiment of the present invention constructs an action database including body parts to infer them from the instruction sentence even if the body parts are not designated in the instruction sentence, and by having the CLIP model learn this, it is possible to infer body parts and action data regardless of the instruction sentence. At this time, the action data is created based on words, but phrases, clauses, and sentences are also possible. In addition, as described above, once the user finally completes the action, the action vector is also accumulated and can be stored so that it is mapped to words, phrases, clauses, sentences, etc., and as described above, this can be used as learning material to continuously fine-tune the generative AI.

[0049] The interpolation unit (360) can perform a procedure for interpolating abnormal movements. At this time, interpolation means filling in the gap, and for example, it means generating a new value between two frames. For example, if there is a walking motion and there is a gap in the frame, an intermediate motion between motion A and motion B is inferred using a deep learning model such as a CNN and then inserted to fill the gap. In one embodiment of the present invention, interpolation is defined not only as filling gaps in this way but also as correcting abnormal movements. If the range of motion of a joint is exceeded, it can be corrected so that it is not exceeded, and if a collision cannot occur, it can be changed to a collision that can occur in the actual physical world.

[0050] The visualization unit (370) can transmit the instruction sentence and the action vector of the action verified by the verification unit to the user terminal (100), and can store the instruction sentence and the action vector of the verified action so as to be mapped to the action database. The user terminal (100) can output a 3D human with an action generated based on the instruction sentence.

[0051] The context memory unit (380) may apply a Session-Based Context Management model to remember the context of the instruction sentence entered from the user terminal (100). At this time, although it is described as remembering, a Large Language Model (LM), such as a generative AI, does not actually remember the sentence given by the user. Just as a doctor looks at a medical chart to recall a history when treating a patient, the context is maintained by providing the user's instruction sentence together (new instruction sentence + previous instruction sentence) when providing instruction sentences that continuously generate or modify actions. This is also called Session-Based Context Management because the context is maintained on a session basis. In this case, a session may be the period during which the user connects and disconnects, or the period from when the user makes a request until it is processed. In the latter case, the context may be maintained within the process in which four instruction sentences are entered and an action vector is generated, in which the user continuously changes and modifies actions as shown in FIG. 3d. Since this is subject to definition, it may be changed according to the embodiment.

[0052] Hereinafter, the operation process according to the configuration of the generation service providing server of FIG. 2 described above will be explained in detail with reference to FIG. 3 and FIG. 4. However, it is obvious that the embodiment is merely one of the various embodiments of the present invention and is not limited thereto.

[0053] Referring to FIG. 3a, (a) when a user terminal (100) inputs an instruction sentence, the generation service providing server (300) uses a CLIP model to classify the action of the instruction sentence by body part as shown in (b) and obtain corresponding action data. Then, the generation service providing server (300) infers the action data of unspecified body parts using correlation based on the action database as shown in (c), integrates the action data as shown in (d) to generate an action vector as shown in (a) of FIG. 3b, and verifies it as shown in (b). The action vector, for which [discrimination-generation-verification] is completed, is transmitted to the user terminal (100) as a set of [instruction sentence-action vector], visualized, and stored in the action database. Additionally, as shown in (d) and FIG. 3d, by remembering the context within a single session, it becomes possible to generate actions for a series of movements. FIG. 3c is explained by reference in FIG. 2, so a redundant explanation is omitted.

[0054] According to one embodiment of the present invention, the applicant (Crito Co., Ltd.) provides a service that generates a 3D human using a user's photo or video as shown in FIGS. 4a to 4c, and applies it to a game as shown in FIG. 4d or to the metaverse as shown in FIG. 4e. A 3D human modeling application as shown in FIG. 4f is currently available on Google Play and the App Store as shown in FIGS. 4g to 4k.

[0055] As for the details regarding the interactive generation service provision method of FIGS. 2 to 4 that are not described, they are identical to or can be easily inferred from the details described above regarding the interactive generation service provision method through FIG. 1, so further explanation will be omitted.

[0056] FIG. 5 is a diagram illustrating the process of transmitting and receiving data between each component included in the interactive 3D human motion generation system using artificial intelligence of FIG. 1 according to an embodiment of the present invention. Hereinafter, an example of the process of transmitting and receiving data between each component will be described through FIG. 5, but the present invention is not to be interpreted as being limited to such an embodiment, and it is obvious to those skilled in the art that the process of transmitting and receiving data illustrated in FIG. 5 may be changed according to various embodiments described above.

[0057] Referring to FIG. 5, the generation service providing server receives an instruction sentence from the user terminal (S5100).

[0058] And, the generation service provider server infers and defines the actions included in the instruction sentence by dividing them by body part (S5200).

[0059] Additionally, the generation service providing server generates motion vectors based on generative AI to move body parts based on movements defined for each body part (S5300), and applies the motion vectors to a pre-built 3D human model to verify movements based on instructions (S5400).

[0060] The order of the steps described above (S5100~S5400) is merely an example and is not limited thereto. That is, the order of the steps described above (S5100~S5400) may vary, and some of these steps may be executed simultaneously or deleted.

[0061] As for the details regarding the interactive generation service provision method of Fig. 5 that are not described, they are identical to or can be easily inferred from the details described above regarding the interactive generation service provision method through Figs. 1 to 4, so further explanation will be omitted.

[0062] The method for providing an interactive generation service according to one embodiment described through FIG. 5 may also be implemented in the form of a recording medium containing computer-executable instructions, such as an application or program module executed by a computer. A computer-readable medium may be any available medium accessible by a computer and includes both volatile and non-volatile media, and both removable and non-removable media. Additionally, a computer-readable medium may include all computer storage media. A computer storage medium includes both volatile and non-volatile, removable and non-removable media implemented by any method or technique for storing information such as computer-readable instructions, data structures, program modules, or other data.

[0063] The method for providing an interactive generation service according to one embodiment of the present invention described above may be executed by an application basically installed on a terminal (which may include a program included in a platform or operating system, etc., basically installed on the terminal), or by an application (i.e., a program) directly installed by a user on a master terminal through an application providing server, such as an application store server, an application, or a web server related to the service. In this sense, the method for providing an interactive generation service according to one embodiment of the present invention described above may be implemented as an application (i.e., a program) that is basically installed on the terminal or directly installed by a user, and may be recorded on a computer-readable recording medium such as a terminal.

[0064] The foregoing description of the present invention is for illustrative purposes only, and those skilled in the art will understand that other specific forms can be easily modified without altering the technical spirit or essential features of the present invention. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive. For example, each component described as a single unit may be implemented in a distributed manner, and components described as distributed may likewise be implemented in a combined form.

[0065] The scope of the present invention is defined by the claims set forth below rather than by the detailed description above, and all modifications or variations derived from the meaning and scope of the claims and equivalent concepts thereof should be interpreted as being included within the scope of the present invention.< / clip>

Claims

1. A user terminal that inputs instruction sentences based on natural language to generate 3D human movements and outputs a 3D human with generated movements based on said instruction sentences; and A generation service providing server comprising: a receiving unit that receives a command sentence from the user terminal; a discrimination unit that infers and defines the motion included in the command sentence by dividing it by body part; a generation unit that generates a motion vector based on generative AI to move the body part based on the motion defined by the body part; and a verification unit that verifies the motion based on the command sentence by applying the motion vector to a pre-constructed 3D human model. An interactive 3D human motion generation system using artificial intelligence including 2. In Paragraph 1, The above-mentioned generation service providing server is, A motion DB construction unit that constructs a motion database by creating a database of a body part where the action of at least one word is performed and motion data for said at least one word; An interactive 3D human motion generation system using artificial intelligence characterized by further including 3. In Paragraph 2, The above-mentioned determination unit is, After dividing the actions included in the above instruction sentence into body parts such as the face, upper body, hands, and lower body, the body part on which the action must be performed is designated, and the action data for each body part of the above instruction sentence is inferred based on a CLIP (Contrastive Language-Image Pre-training) model, Motion data for each body part not specified by the above instruction sentence is selected based on the correlation between the specified motion data for each body part based on the above motion database, and An interactive 3D human motion generation system using artificial intelligence characterized by generating integrated data by integrating motion data for the specified body parts and motion data for unspecified body parts.

4. In Paragraph 1, The above verification unit is, An interactive 3D human motion generation system using artificial intelligence, characterized by applying the motion vector to the 3D human model and determining the motion as abnormal if the motion includes an angle that is not possible with human joint angles or if the physical collision caused by the motion is a collision that cannot be generated by physical simulation.

5. In Paragraph 4, The above-mentioned generation service providing server is, An interpolation unit that performs a procedure to interpolate the above abnormal operation; An interactive 3D human motion generation system using artificial intelligence characterized by further including 6. In Paragraph 1, The above-mentioned generation service providing server is, A visualization unit that transmits the above instruction sentence and the operation vector of the operation verified by the verification unit to the user terminal, and stores the above instruction sentence and the operation vector of the verified operation so as to be mapped to an operation database; An interactive 3D human motion generation system using artificial intelligence characterized by further including 7. In Paragraph 1, The above-mentioned generation service providing server is, A context memory unit that applies a Session-Based Context Management model to remember the context of a command sentence entered at the above-mentioned user terminal; An interactive 3D human motion generation system using artificial intelligence characterized by further including