Method and system for generating 3D objects based on semantic analysis
The method and system analyze text to generate three-dimensional objects that perform actions corresponding to sentence meanings, addressing the limitations of current technology by automatically creating dynamic 3D models with natural motion and diverse elements.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- ネイションエイ インク
- Filing Date
- 2023-09-14
- Publication Date
- 2026-06-04
AI Technical Summary
Current three-dimensional modeling technology is insufficient for generating dynamic objects that perform actions corresponding to the meaning of a sentence, with a lack of tools capable of converting moving objects into three-dimensional space and representing them naturally.
A method and system that analyzes text to identify sentences, defines actions corresponding to the sentence, and generates a three-dimensional object using a trained artificial neural network to perform these actions, allowing for the linking of multiple actions to create a moving 3D object.
Enables the generation of animated 3D objects that perform defined actions based on text input, without requiring users to manually define actions each time, and can acquire various elements like movement, facial expressions, and background information.
Smart Images

Figure 2026518113000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and system for generating a three-dimensional object based on semantic analysis. More specifically, the present invention relates to a method and system for analyzing a text-formatted sentence and automatically generating a three-dimensional object corresponding to the meaning of the sentence.
Background Art
[0002] With the growth of the metaverse market, the demand for three-dimensional (3D, 3-Dimensional) object generation technology, which is a core element, has also increased rapidly. Companies at home and abroad are paying attention to the metaverse that reproduces the real world in the virtual world as the next-generation growth driver, and the interest in three-dimensional modeling technology, which is the core technology of the metaverse and converts the real world into three dimensions and represents it in the virtual world, has also increased significantly.
[0003] Once, three-dimensional modeling technology was mainly utilized in the game field, but recently, it has tended to be extended and applied to the entire industrial area, expanding its application scope to various content industries such as AR / VR, movies, animations, and broadcasts.
[0004] However, despite the high interest in three-dimensional modeling technology, the technical maturity of this technology is not sufficient. For example, in the case of companies that generate AI humans, etc., technology development has been concentrated only on the generation of virtual faces by the CG generation method or the Deep Fake generation method, and the development of technology for recognizing and naturally reproducing the motion of a person moving is not sufficient.
[0005] In addition, as three-dimensional modeling technology is diffusely applied to the entire industrial area, the three-dimensional modeling tool market is also continuing to grow, but the current three-dimensional imaging modeling tools are specialized only in creating static objects (such as architecture, interiors, equipment, etc.), and the research and development of technology for converting moving dynamic objects into three dimensions and representing them in a three-dimensional space is in an insufficient situation.
Summary of the Invention
[0006] The technical problem to be solved through embodiments of the present invention is to provide a method and system for analyzing meaning and generating three-dimensional objects. Specifically, the present invention provides a method and system for analyzing text provided in text format and automatically generating three-dimensional objects corresponding to the meaning of the text.
[0007] To this end, the present invention can provide a method for generating a three-dimensional object that performs actions corresponding to a sentence by applying actions related to the meaning of the identified sentence to a three-dimensional model once the sentence has been identified.
[0008] Another technical problem that the embodiments of the present invention aim to solve is to provide a method and system for generating a three-dimensional object that performs an entire action corresponding to the meaning of a sentence by sequentially combining a plurality of three-dimensional objects that perform partial actions corresponding to the meaning of words in a sentence.
[0009] Furthermore, the present invention relates to a method and system for generating and applying linked actions between different actions to a 3D model so that the 3D model can perform multiple actions naturally in relation to the meaning of a text. [Means for solving the problem]
[0010] To solve the above-mentioned problems, the method for generating a 3D object according to the present invention may include the steps of: identifying a sentence based on user input; analyzing the identified sentence and defining a plurality of actions corresponding to the sentence based on the analysis results; defining a linking action that links the plurality of actions corresponding to the sentence; and generating a moving 3D object that performs the plurality of actions and the linking action using a trained artificial neural network.
[0011] In one example, the linking operation may be a linking operation that connects at least two operations that are adjacent to each other in a time-series flow from among the plurality of operations.
[0012] In one example, the movement of the 3D object may be formed by a plurality of frames corresponding to the first action, which include keypoint information of the 3D model corresponding to the first action; a plurality of frames corresponding to the second action, which include keypoint information of the 3D model corresponding to the second action; and a plurality of frames corresponding to the connecting action, which includes keypoint information of the 3D model corresponding to the connecting action that connects the first action and the second action.
[0013] In one example, each of the multiple frames corresponding to the first operation, the multiple frames corresponding to the second operation, and the multiple frames corresponding to the linked operation may include positional information with respect to key points of the three-dimensional model.
[0014] In one example, if the first and second operations are adjacent operations that follow a time series, the key points of the 3D models of the multiple frames corresponding to the linked operations may be determined or defined based on the key point information of the 3D models of the multiple frames corresponding to the first operation and the multiple frames corresponding to the second operation.
[0015] In one example, if the first operation precedes the second operation based on the time series, the position information of the key points of the 3D model included in the first frame of the multiple frames corresponding to the linked operation may relate to the position information of the key points of the 3D model included in the last frame of the multiple frames corresponding to the first operation, and the position information of the key points of the 3D model included in the last frame of the multiple frames corresponding to the linked operation may relate to the position information of the key points of the 3D model included in the first frame of the multiple frames corresponding to the second operation.
[0016] In one example, the position of the key point of the 3D model included in the first frame of the linking operation may be the same as or adjacent to the position of the key point of the 3D model in the final frame of the first operation, and the position of the key point of the 3D model included in the final frame of the linking operation may be the same as or adjacent to the position of the key point of the 3D model in the first frame of the second operation.
[0017] The 3D object generation system according to the present invention may include a storage unit that includes a pre-provided database, and a control unit that identifies a sentence based on user input, analyzes the identified sentence, defines a plurality of actions corresponding to the sentence based on the analysis results, defines a linking action that links the plurality of actions corresponding to the sentence, and generates a moving 3D object that performs the plurality of actions and the linking action using a trained artificial neural network.
[0018] A program executed by one or more processes in an electronic device and stored on a computer-readable recording medium may include instructions that cause the following steps to be performed: identifying a sentence based on user input; analyzing the identified sentence and defining a plurality of actions corresponding to the sentence based on the analysis results; defining a concatenation action that links the plurality of actions corresponding to the sentence; and using a trained artificial neural network, generating a moving three-dimensional object that performs the plurality of actions and the concatenation action.
[0019] The method for generating a three-dimensional object according to the present invention may include the steps of: identifying text based on user input; analyzing the identified text and defining a storyboard corresponding to the text based on the analysis results; acquiring a three-dimensional model and operation information corresponding to the storyboard using a pre-established database; and generating a three-dimensional object that performs the operation corresponding to the text by applying the operation information to the three-dimensional model.
[0020] Furthermore, the 3D object generation system according to the present invention may also include a storage unit that includes a pre-provided database, and a control unit that identifies a text based on user input, analyzes the identified text, defines a storyboard corresponding to the text based on the analysis results, acquires a 3D model and operation information corresponding to the storyboard using the pre-provided database, and generates a 3D object that performs the operation corresponding to the text by applying the operation information to the 3D model.
[0021] Furthermore, the program stored on a computer-readable recording medium according to the present invention is executed by one or more processes in an electronic device and is a program stored on a computer-readable recording medium, and the program may include instructions that cause the execution of the following steps: identifying a document based on user input; analyzing the identified document and defining a story corresponding to the document based on the analysis results; acquiring a three-dimensional model and operation information corresponding to the story using a pre-established database; and generating a three-dimensional object that performs an operation corresponding to the document by applying the operation information to the three-dimensional model. [Effects of the Invention]
[0022] The 3D object generation method and system according to the present invention define actions corresponding to the meaning of a sentence and generate 3D objects that perform the defined actions. This makes it possible to generate animated 3D objects simply by inputting text, without the user having to define the actions of the 3D objects each time.
[0023] Furthermore, the 3D object generation method and system according to the present invention can acquire information that realizes various elements of a 3D model corresponding to a 3D object, such as the movement, facial expression, external form, age, clothing, gender, physical condition, background, location, and situation, based on keywords identified in correspondence with text.
[0024] Also, according to various embodiments of the present invention, the three-dimensional object generation method and system can naturally connect different operations performed by a three-dimensional object by generating concatenated operation information for a plurality of operation information with consecutive operation sequences.
Brief Description of the Drawings
[0025] [Figure 1] Shows a three-dimensional object generation system according to the present invention. [Figure 2] Shows an embodiment of a database included in the three-dimensional object generation system according to the present invention. [Figure 3] It is a flowchart showing a three-dimensional object generation method according to the present invention. [Figure 4] It is a flowchart showing a three-dimensional object generation method according to the present invention. [Figure 5] Shows an embodiment of identifying a sentence based on user input. [Figure 6] Shows an embodiment of defining an operation corresponding to a sentence. [Figure 7] Shows an embodiment of defining an operation corresponding to a sentence. [Figure 8] Shows an embodiment of defining an operation corresponding to a sentence. [Figure 9] Shows an embodiment of defining an operation corresponding to a sentence. [Figure 10] Shows an embodiment of applying operation information to a three-dimensional model. [Figure 11] Shows an embodiment of applying operation information to a three-dimensional model. [Figure 12] Shows an embodiment of applying operation information to a three-dimensional model. [Figure 13] Shows an embodiment of applying operation information to a three-dimensional model. [Figure 14] Shows an embodiment of applying operation information to a three-dimensional model.
Modes for Carrying Out the Invention
[0026] The embodiments disclosed herein will be described in detail below with reference to the accompanying drawings, but regardless of the reference numerals used in the drawings, identical or similar components will be given the same reference numerals, and redundant descriptions thereof will be omitted. The suffixes “module” and “part” used for components in the following description are added or mixed for the sake of ease of writing the specification and do not have any distinguishing meaning or role in themselves. Furthermore, when describing the embodiments disclosed herein, if it is determined that a detailed description of the relevant prior art would obscure the gist of the embodiments disclosed herein, such detailed description will be omitted. In addition, the accompanying drawings are intended solely to facilitate the understanding of the embodiments disclosed herein, and the technical ideas disclosed herein should not be limited by the accompanying drawings and should be understood to include all modifications, equivalents and substitutions that fall within the concept and technical scope of the present invention.
[0027] Terms including ordinal numbers such as "1st," "2nd," etc., may be used to describe various components, but the components are not limited to those defined by these terms. These terms are used solely to distinguish one component from another.
[0028] When it is stated that one component is “connected” or “linked” to another component, it should be understood that it may be directly connected or linked to the other component, but there may also be another component between them. On the other hand, when it is stated that one component is “directly connected” or “directly linked” to another component, it should be understood that there is no other component between them.
[0029] A singular expression includes plural forms unless otherwise clearly indicated in the context.
[0030] In this application, terms such as “includes” or “having” are intended to specify the presence of features, figures, steps, actions, components, parts, or combinations thereof as described in the specification, and should be understood not to preemptively exclude the possibility of the presence or addition of one or more other features, figures, steps, actions, components, parts, or combinations thereof.
[0031] Figure 1 shows a 3D object generation system according to the present invention. Figure 2 shows one embodiment of the database included in the 3D object generation system according to the present invention.
[0032] Referring to Figure 1, the 3D object generation system 100 according to the present invention can identify a sentence based on user input 10 and, by analyzing the identified sentence, define (or determine, generate, extract) a continuity corresponding to the sentence. Based on the learning described later, the 3D object generation system 100 can define a continuity corresponding to the sentence, or define an action or a 3D model. In the present invention, a continuity is information that defines various things such as the action that the 3D model should perform in order to express the meaning or situation corresponding to the sentence, the shape of the 3D model, size, facial expression, gender, etc., and can also be understood as a "storyboard". In the present invention, "definition" can be understood as determining, generating, or extracting information or data.
[0033] The 3D object generation system 100 can acquire a 3D model 1 and motion information corresponding to the text from a pre-established database 131 based on a defined storyboard, and apply the motion information to the 3D model 1 to generate a 3D object 3 that performs actions according to the text. In the present invention, a 3D model is a subject that performs actions corresponding to the text (for example, a person, an animal, an object, etc.), and a dynamic object that performs actions by moving a 3D model is called a 3D object. The 3D objects described in the present invention can be renamed 3D dynamic objects, 3D graphic objects, 3D graphic dynamic objects, 3D modeling objects, 3D animation objects, or objects, dynamic objects, graphic objects, graphic dynamic objects, modeling objects, animation objects, etc. On the other hand, the implementing body of the present invention in the following description is referred to as the 3D object generation system 100, but this may be replaced with "control unit".
[0034] In the present invention, the storyboard may include information defining at least one of the persona of a 3D model and the behavior of a 3D model. That is, the 3D object generation system according to the present invention can define, generate, or extract information from text that defines at least one of the persona of a 3D model and the behavior of a 3D model, based on a trained artificial intelligence model.
[0035] In this case, the persona of the 3D model may include information that defines the visual appearance, personality, etc. of the 3D model.
[0036] The visual appearance of a 3D model may include at least some of the various pieces of information that can define the visual appearance of the 3D model, such as the model's age, gender, height, weight, hairstyle, face shape, skin color, and clothing. Furthermore, the persona may include at least some of the various pieces of information that influence the behavioral characteristics of the 3D model, such as the model's occupation and personality. The 3D model may be defined in various ways depending on the specified text, such as a person, animal, or object.
[0037] On the other hand, the 3D object generation system 100 can define the storyboard based on the input specified sentence and other sentences (or other information, such as information contained in the paragraph containing the specified sentence, or sentences (or information) input before or after the input of the specified sentence).
[0038] On the other hand, persona information for defining a 3D model may not be extracted from text, but may be defined based on separate user input or settings. In this case, the 3D object generation system according to the present invention may include a separate process for generating or defining a 3D model that performs actions according to text. In this case, the 3D object generation system 100 can identify at least one action that the 3D model should perform from the text, and generate a 3D object in which the 3D model performs the identified action.
[0039] On the other hand, the present invention may be understood as defining at least one "action" that a 3D model performs from text, rather than "defining a storyboard" from text.
[0040] The 3D object generation system 100 can receive user input 10 in various forms, such as audio signals, images, and text, and can identify sentences to apply to the 3D model 1 based on the received user input 10.
[0041] For example, when the 3D object generation system 100 identifies the phrase "a person walks" based on user input 10, it can define the action of "walking" based on the analysis results of the phrase and obtain action information related to "walking" from the database 131. This allows the 3D object generation system 100 to generate and output a 3D object 3 that performs a walking action by applying the obtained action information to the 3D model 1. In this case, if the 3D model 1 is defined as a person, the 3D object generation system 100 can generate and output a 3D object 3 in the shape of a person performing a walking action by applying the obtained action information to the 3D model 1, which is prepared in the shape of a person. For the sake of explanation, a "person" will be used as an example of a 3D model below, but the 3D model in this invention is not necessarily limited to a person.
[0042] On the other hand, in this specification, a text may be identified based on user input 10, and the text may be identified in the form of text that describes at least one of the behavior and appearance of a three-dimensional object 3.
[0043] Furthermore, the 3D object 3 may be generated in a way that it is modeled in 3D space, and may be generated to have various shapes such as people, animals, and objects.
[0044] For example, a 3D object 3 may include at least one of the following: multiple key points, lines connecting two different key points, a mesh representing the shape (or volume) of the 3D object 3, and a texture (or image) representing the texture and outline of the 3D object 3.
[0045] In this case, the 3D model 1 may be provided to constitute the 3D object 3, that is, the 3D model 1 can form the framework (or skeleton) of the 3D object 3. For this purpose, the 3D model 1 may be configured to correspond to at least one of the following in the 3D object 3: a plurality of keypoints, a line connecting two different keypoints, and a mesh representing the shape (or volume) of the 3D object 3. In some embodiments, the 3D object 3 may be generated by overlaying a texture onto the 3D model 1.
[0046] On the other hand, the positions of multiple keypoints may be specified to correspond to the joint parts of 3D model 1 (or 3D object 3).
[0047] Furthermore, the 3D model 1 may include, as label data for each keypoint, information for distinguishing each keypoint, the range of movement available for the position of each keypoint, and the range of speed available for changes in each keypoint. This allows the 3D model 1 to be configured to achieve natural motion.
[0048] On the other hand, the storyboard may include information in various categories (or fields), not only information for defining the actions of the 3D model and information for defining the persona, but also information for defining the circumstances in which the 3D model performs its actions and information for defining the background in which the 3D model performs its actions. For example, the storyboard may include information (or definitions) in various categories (or fields), such as the actions of the 3D model, the shape of the 3D model, age, clothing, gender, physical condition, and the background, location, and situation in which the 3D model performs its actions. In other words, the storyboard may include (or define) information related to the shape of the 3D model obtained in response to the text identified above, the actions applied to the 3D model, and the location, background, and situation in which the 3D model is placed.
[0049] For example, the 3D object generation system 100 may define a storyboard in which, in response to the sentence "A cat jumps over a wall," a 3D model having a shape corresponding to a cat performs the action of jumping over a wall at a location where the wall is placed. Specifically, "cat" may be included in the storyboard as a definition for a category related to the external shape of the 3D model, "wall" as a definition for a category related to the background in which the 3D model is placed, and "jump over" as a definition for a category related to the action applied to the 3D model.
[0050] On the other hand, motion information may include information generated so that the 3D model 1 moves according to the actions based on the text. Specifically, motion information may include information related to the movement of each of the one or more keypoints included in the 3D model 1. That is, motion information may include multiple poses or multiple frames (or time), and each pose or frame may include position information for each of the multiple keypoints included in the 3D model 1. In this case, the motion information may specify the positions of each of the multiple keypoints with respect to a single point in 3D space.
[0051] For example, the motion information may include information related to the positional changes and rate of change for each of the multiple keypoints included in the 3D model 1.
[0052] As another example, motion information may be generated such that the position information for each of the multiple keypoints included in the 3D model 1 is specified at a predetermined time period (e.g., 1 / 60 second) for a series of time intervals corresponding to the motion information.
[0053] Furthermore, the motion information may include facial expression information corresponding to facial expression categories generated so that the facial parts of the 3D model 1 move in accordance with the facial expressions based on the text. In this case, the facial expression information may include information related to the movement of each key point belonging to the facial parts of the 3D model 1, and therefore may be provided to be included in the motion information.
[0054] Referring to Figure 2, the database 131 may store multiple 3D models 1 and multiple motion information 20 that are different from each other. Here, multiple 3D models 1 that are different from each other may mean multiple 3D models 1 that are provided to have different external shapes. For example, multiple 3D models 1 that are different from each other may include objects with various external shapes such as cats, people, boys, girls, and dogs. In this case, the number and position of multiple key points, mesh format and texture, etc., of multiple 3D models 1 that are different from each other may be set differently depending on the external shape of each object.
[0055] The database 131 may further store labeling information 60 (or tags, keywords) matched to each of the multiple 3D models 1 and the multiple motion information 20.
[0056] The labeling information 60 may be set to correspond to the results of the text analysis. That is, the labeling information 60 is a keyword for a story defined according to the text, and can represent elements such as actions corresponding to the story and the outline of a 3D model. In other words, the 3D object generation system 100 can define a story corresponding to a text by obtaining keywords from the text. At this time, the 3D object generation system 100 can also classify the keywords according to categories corresponding to the story.
[0057] For example, labeling information 60 may include words such as "person" and "walk" in relation to the action of a person walking. Another example is that labeling information 60 may further include objects, complements, and adverbs in relation to the action performed by the subject and predicate included in the sentence.
[0058] As a result, the 3D object generation system 100 can obtain a 3D model 1 and operation information corresponding to a document from the database 131 using a context (e.g., keywords or words) defined based on the results of the document analysis. In particular, by using morphological analysis, the 3D object generation system 100 can obtain a 3D model 1 with more accurate and diverse forms, as well as operation information 20 related to the operation, by further considering various document components such as objects, complements, and adverbs in addition to the subject and predicate contained in the document.
[0059] In this regard, the database 131 may store an artificial neural network model that outputs a 3D model 1 and operation information corresponding to a text when a story defined from a text is input. In such a case, the 3D object generation system 100 according to the present invention can input the story defined based on the text analysis results into the artificial neural network model provided in the database 131, and obtain the 3D model 1 and operation information corresponding to the text from the artificial neural network model.
[0060] For this purpose, the artificial neural network model is trained using training scenarios and ground truth 3D models and ground truth operation information labeled as label data in the training scenarios. In this case, the 3D models and operation information may include the various types of information mentioned above. Therefore, the artificial neural network model can be trained to output a 3D model 1 and operation information corresponding to the input scenario when a scenario is input, using training scenarios corresponding to the various categories mentioned above and ground truth operation information.
[0061] In this regard, the overall operation of the 3D object generation system 100 according to the present invention will be described below by describing a method for searching for 3D models and operation information corresponding to a storyboard in database 131. However, it should be considered that in the present invention, the 3D models and operation information can also be generated using an artificial neural network model as described above.
[0062] Furthermore, the database 131 may be classified and stored according to the various categories mentioned above. In such a case, the 3D object generation system 100 can match the storyboards corresponding to the text according to the categories and retrieve the data corresponding to the matched categories from the database 131.
[0063] Depending on the embodiment, the database 131 may store multiple artificial neural network models categorized according to the various categories mentioned above, and data corresponding to each category may also be retrieved.
[0064] For this purpose, the 3D object generation system 100 according to the present invention may include an input unit 110, a storage unit 130, a control unit 150, and an output unit 170. On the other hand, Figure 1 shows only the components related to embodiments of the present invention. Therefore, a person of ordinary skill in the art to which the present invention belongs will see that other general-purpose components may be included in addition to the components shown in Figure 1.
[0065] The input unit 110 can receive user input 10 via a pre-installed input device. For example, the input device may refer to various devices such as a keyboard, mouse, touch pad, microphone, or line-in. This allows the input unit 110 to receive user input 10 of various types, such as voice and text. As a result, the control unit 150 can acquire text corresponding to the user input 10.
[0066] The memory unit 130 may store information and instruction words necessary for the operation of the 3D object generation system 100 according to the present invention. Therefore, the memory unit 130 can load one or more programs to perform methods / operations according to various embodiments of the present invention. An example of the memory unit 130 is, but is not limited to, RAM.
[0067] On the other hand, the memory unit 130 may store a database containing multiple operational information. Furthermore, the memory unit 130 (or the database) may store a 3D model 1 to which the operational information is applied.
[0068] The control unit 150 can control the overall operation of the 3D object generation system 100 according to the present invention. Therefore, the control unit 150 may be configured to include at least one of a CPU (Central Processing Unit), MPU (Micro Processor Unit), MCU (Micro Controller Unit), GPU (Graphic Processing Unit), or any type of processor well known in the art of the present invention. Furthermore, the control unit 150 can perform calculations for at least one application or program to perform methods / operations according to various embodiments of the present invention. The 3D object generation system 100 may comprise one or more control units 150.
[0069] On the other hand, the control unit 150 analyzes the text corresponding to the user input 10, defines a storyboard corresponding to the text from the database 131, obtains a 3D model and motion information based on the storyboard, and generates a 3D object 3 by applying the motion information to the 3D model 1.
[0070] The output unit 170 can output the 3D object 3 generated by the control unit 150. At this time, the output unit 170 can output the 3D object 3 in a way that allows the user to visually confirm it via a display device such as a monitor or television.
[0071] Based on the configuration of the 3D object generation system 100 described above, the method for generating 3D objects will be explained in more detail below.
[0072] Figures 3 and 4 are flowcharts of the 3D object generation method according to the present invention. Figure 5 shows one embodiment of identifying text based on user input. Figures 6 to 9 show one embodiment of defining actions corresponding to text. Figures 10 to 14 show one embodiment of applying action information to a 3D model.
[0073] Referring to Figure 3, the 3D object generation system 100 according to the present invention can identify a text based on user input 10 (S100), analyze the identified text, and define a storyboard corresponding to the text based on the analysis results (S200).
[0074] Specifically, the 3D object generation system 100 can identify a sentence based on the input text if the data entered by the user input 10 is text. Furthermore, if the data entered by the user input 10 is in a form other than text, the 3D object generation system 100 can analyze the input data, convert it to text, and identify a sentence so that an action is defined to apply it to the 3D model 1 based on the converted text.
[0075] Referring to Figure 5 as an example, the 3D object generation system 100 can identify the input text 10a as a sentence when text 10a is input based on user input 10.
[0076] As another example, when the 3D object generation system 100 receives data containing multiple texts based on user input 10, it can perform analysis on the input data to extract multiple texts, divide the extracted multiple texts, and sequentially identify each of the divided texts in the order they appeared in the data as a sentence for generating the 3D object 3.
[0077] To this end, the 3D object generation system 100 can divide multiple texts based on texts corresponding to pre-set sentence codes (e.g., periods) within the text contained in the data. This allows the 3D object generation system 100 to define a storyboard corresponding to each text by sequentially analyzing the divided texts in the order they appeared in the data.
[0078] Another example is that when a 3D object generation system 100 receives an audio signal 10b based on user input 10, it can generate text 11a (or sentences) corresponding to the user input 10 by analyzing the audio signal 10b.
[0079] In this case, the 3D object generation system 100 can use a technique (for example, Speech-to-Text) to convert the audio signal 10b corresponding to the user input 10 into text 11a, and may include an audio model or artificial neural network for this purpose.
[0080] Another example is a 3D object generation system 100, which, when an image is input based on user input 10, can generate text corresponding to the user input 10 by analyzing the image.
[0081] In this case, the 3D object generation system 100 may use techniques for converting an image corresponding to the user input 10 into text (for example, optical character recognition (OCR)), and may include a character detection model, a character recognition model, and an artificial neural network for this purpose.
[0082] Another example is that when data generated according to a pre-set format (e.g., a scenario format) based on user input 10 is input to the 3D object generation system 100, it can extract one or more texts from pre-set areas to represent the outline, background, and behavior of the object in accordance with the format, and sequentially identify the extracted one or more texts as text for generating the 3D object 3.
[0083] In this case, if the pre-set format is a scenario (or script) format, the 3D object generation system 100 can extract one or more texts from the parentheses of areas 11b, 11c, 11d, and 11e that correspond to the stage directions (Actions) in the scenario. This allows the 3D object generation system 100 to sequentially identify each text in the order it was included in the data as text for generating the 3D object 3.
[0084] Furthermore, the 3D object generation system 100 can divide a sentence into multiple words at the morphological level, confirm the meanings corresponding to the divided multiple words, and generate (or extract) at least one keyword corresponding to the meaning of the sentence identified above based on the confirmed meanings. In this case, the generation of at least one keyword corresponding to the meaning of the sentence by the 3D object generation system 100 may also be defined as a storyboard corresponding to the sentence.
[0085] For example, the 3D object generation system 100 can define a storyboard corresponding to a text by using an artificial neural network (e.g., RNN (Recurrent Neural Network) and LSTM (Long Short-Term Memory)) that has been trained to output keywords related to the storyboard corresponding to the multiple words input sequentially from the text.
[0086] In such cases, the 3D object generation system 100 can input the text identified above into a pre-trained artificial neural network to obtain keywords related to the story corresponding to the text, and generate a 3D object that performs the action corresponding to the keywords.
[0087] Another example is that the 3D object generation system 100 can identify the part of speech of each word in a sentence and extract words corresponding to parts of speech related to the object's behavior as keywords.
[0088] In this case, the 3D object generation system 100 can use an artificial neural network that has been pre-trained to identify the part of speech of words, or it can use a pre-defined language model to identify the part of speech of words.
[0089] With the configuration described above, the 3D object generation system 100 according to the present invention can generate keywords related to a storyboard from text identified based on user input 10.
[0090] Referring to Figure 6, for example, if the sentence "He sat on the floor and then stood up" is identified, the 3D object generation system 100 can obtain "he," "sit," "stand up," and "floor" as keywords related to the storyboard through morphological analysis of the identified sentence. At this time, based on the sentence structure of the obtained keywords, the 3D object generation system 100 can define the outline of the 3D model corresponding to "he," define the action of sitting on the floor corresponding to "sit" and "floor," and define the action of standing up from the floor corresponding to "he," "stand up," and "floor."
[0091] As another example, if the 3D object generation system 100 identifies the sentence "sit with legs crossed and then stand up," it can define the sitting and standing actions as storyboards corresponding to the analysis results of the identified sentence, thereby generating keywords related to "sitting" and "standing up."
[0092] Furthermore, the 3D object generation system 100 can determine the correlation between multiple words that are different from each other based on the meaning of each of the multiple words separated from the text, and based on the determined correlation, it can determine at least one of the action levels and action sequences for at least one keyword related to the text.
[0093] Here, the action level can represent the level of multiple different actions performed simultaneously at the same time on a single 3D model (e.g., arms, legs, person).
[0094] Furthermore, the operation sequence can represent the order in which multiple different operations are performed on a single 3D model. In other words, the operation sequence can represent the temporal order in which a single 3D object performs its operations.
[0095] For example, the 3D object generation system 100 can generate the keywords "legs," "cross," and "sit" in response to the sentence "sit with your legs crossed." In this case, the 3D object generation system 100 can confirm that the object of "cross" is "legs" based on the case particle "-o" attached to "legs." Therefore, the 3D object generation system 100 can determine that "legs" and "cross" are at the same level of action.
[0096] Furthermore, the 3D object generation system 100 can confirm that "sitting" and "assembling" occur at the same time based on the conjunction "-de" for "assembling," and can determine that "assembling," which is an action performed on the "legs," has a lower action level than "sitting," based on the fact that the action of "sitting" is performed by an object.
[0097] As another example, the 3D object generation system 100 can generate the keywords "sit" and "stand up" in response to the sentence "sit down and then stand up." In this case, the 3D object generation system 100 can confirm that "stand up" occurs after "sit down" based on the auxiliary particle "-te kara" of "sit down." Therefore, the 3D object generation system 100 can determine that "sit down" has an action sequence that precedes "stand up."
[0098] Furthermore, if the storyboard defined above includes various elements related to a 3D object, the 3D object generation system 100 can match each of the multiple keywords corresponding to the storyboard to one of several categories pre-defined by various elements (for example, action, facial expression, external appearance, age, clothing, gender, physical condition, background, location, situation, etc.), and retrieve information corresponding to each keyword from the database based on the matched category.
[0099] For example, the 3D object generation system 100 can match at least one of several keywords corresponding to a storyboard to a category related to facial expressions, so that facial expressions are applied to the face parts of the 3D model 1. In this case, the 3D object generation system 100 can match keywords related to words corresponding to emotions (or sensibilities) from among several words contained in a text to a category related to facial expressions corresponding to the text.
[0100] In one embodiment, the 3D object generation system 100 can generate the keywords "fun" and "run" in response to the phrase "run happily." In this case, the 3D object generation system 100 can match the keyword "fun," which corresponds to emotion, to a category related to facial expressions, and match the keyword "run," which corresponds to action, to a category related to action.
[0101] Furthermore, if the 3D object generation system 100 contains multiple words in a document that include keywords related to the external shape of the 3D object 3 (or 3D model 1), it can match those keywords to categories related to the external shape. In this case, the 3D object generation system 100 can match keywords related to the external shape of the 3D object 3 with keywords related to its operation. As a result, the 3D object generation system 100 can obtain information corresponding to each category (for example, 3D model and operation information) from the database so that the 3D object 3, which has an external shape corresponding to the document, performs the operation corresponding to the document.
[0102] For example, the 3D object generation system 100 can generate the keyword "person" as a keyword related to the external shape of the 3D model 1 in response to the sentence "A person runs." In this case, the 3D object generation system 100 can match keywords related to the action "running" with the keyword "person," which is a keyword related to the external shape of the 3D model 1.
[0103] As another example, the 3D object generation system 100 can generate the keyword "cat" as a keyword related to the outline of the 3D model 1 in response to the sentence "A cat jumps over the wall." In this case, the 3D object generation system 100 can match keywords related to the actions of "wall" and "jumping over" (or "jumping" and "jumping over") with the keyword "cat," which is a keyword related to the outline of the 3D model 1.
[0104] Another example is that the 3D object generation system 100 can generate the keywords "boy" and "girl" as keywords related to the outline of the 3D model 1, corresponding to the sentence "A boy and a girl are standing under the same umbrella." In this case, the 3D object generation system 100 can match keywords related to the actions "umbrella," "under," and "standing" with the keywords "boy" and "girl," which are keywords related to the outline of the 3D model 1.
[0105] Thus, the 3D object generation system 100 can determine the behavior and external form of a 3D object in response to keywords generated by semantic analysis of text, and can also generate 3D objects in response to various keywords related to the 3D object, such as age, clothing, gender, and physical characteristics (e.g., weight, height, etc.).
[0106] Furthermore, if the 3D object generation system 100 generates keywords related to the space in which the 3D object is generated (for example, location, background, outline of the object placed in the space, situation, etc.) through semantic analysis of the text, it can also generate other 3D objects adjacent to the 3D object to correspond to those keywords.
[0107] For this purpose, the database included in the 3D object generation system 100 may store data corresponding to various categories such as motion, external form, facial expression, age, clothing, gender, and physical characteristics.
[0108] With the configuration described above, the 3D object generation system 100 according to the present invention can define at least one action that is most suitable for a sentence by analyzing the sentence identified based on the user input 10.
[0109] On the other hand, the 3D object generation system 100 can also define multiple situations (or multiple storylines or multiple actions) as a result of analyzing the text identified above. In such cases, each of the multiple situations may be realized in different spaces at the same time, or in the same space at different time points.
[0110] Therefore, the 3D object generation system 100 can use a database to acquire 3D models and motion information corresponding to each situation, and by applying the motion information to the 3D models, it can generate 3D objects that perform actions corresponding to each situation.
[0111] In one embodiment, when multiple different situations are applied to a single 3D model, the 3D object generation system 100 may generate the 3D object in such a way that the 3D object performs an action corresponding to the previous situation, and after a predetermined time interval, performs an action corresponding to the next situation, according to the order of the multiple situations determined by the text analysis results.
[0112] Furthermore, the 3D object generation system 100 may define conditional storyboards for specific situations as a result of analyzing the text identified above. The conditional storyboards may be configured so that the 3D object performs an action corresponding to the conditional storyboard when the conditions set based on the text analysis results (or user input) are met.
[0113] For example, a conditional storyboard may be defined to perform a specific action (a stretching action) at a specific time (e.g., 6:00 AM). In such a case, a 3D object generated based on the conditional storyboard may be embodied to perform the specific action at the specific time, depending on the situation.
[0114] As another example, a conditional storyboard may be defined to perform a specific action (e.g., a greeting action) in response to a pre-set user input (e.g., a click on a 3D object). In such a case, the 3D object generated based on the conditional storyboard may be embodied to perform the specific action in response to the user input.
[0115] To this end, the 3D object generation system 100 obtains a 3D model and motion information corresponding to the conditional storyboard from the database, and generates a 3D object by applying the motion information to the 3D model. However, an event flag corresponding to the conditional storyboard may be set for the motion information applied to the 3D model.
[0116] As a result, the 3D object generation system 100 can generate 3D objects that operate in response to specific conditions, and can also generate a single 3D object to which multiple conditional scenarios, each corresponding to different conditions, are applied.
[0117] Furthermore, the 3D object generation system 100 may directly define at least one action as a result of analyzing the text identified above. In such a case, the 3D object generation system 100 can also generate a 3D object that performs the action corresponding to the text by obtaining action information corresponding to the defined at least one action from a database and applying the action information to a pre-defined 3D model.
[0118] In other words, the 3D object generation system 100 can also generate a 3D object through the process of generating at least one keyword based on a text, obtaining at least one of a 3D model and motion information corresponding to the at least one keyword from a database, and applying the motion information to the 3D model.
[0119] Referring again to Figure 3, the 3D object generation system 100 according to the present invention can acquire a 3D model and operation information corresponding to a storyboard using a pre-established database 131 (S300).
[0120] Specifically, the 3D object generation system 100 can retrieve a 3D model and operation information from the database 131 that corresponds to at least one keyword generated above. This allows the 3D object generation system 100 to obtain the retrieved 3D model and operation information.
[0121] For example, if the keyword "run" is generated, the 3D object generation system 100 can retrieve action information corresponding to the running action by searching for action information corresponding to "run" in the database 131.
[0122] Referring to Figure 7, another example is that for the sentence "He sat on the floor and then stood up," the 3D object generation system 100 defines the outline of a 3D model related to "he" as a storyboard, and can retrieve the 3D model corresponding to "he" from the database 131. In addition, the 3D object generation system 100 defines a first action related to "sitting" and "floor," and a second action related to "he," "standing up," and "floor," and can retrieve action information corresponding to the first and second actions from the database 131.
[0123] As a result, the 3D object generation system 100 can obtain a 3D model related to a person, motion information 21 related to sitting on the floor, and motion information 22 related to standing up from the floor, respectively, from the database 131.
[0124] Referring to Figure 8, another example is that when the keywords "legs," "cross," and "sit" are generated, the 3D object generation system 100 can retrieve action information corresponding to "sit" from the database 131 based on the action level set for each keyword. At this time, the 3D object generation system 100 can check from the database 131 for action information corresponding to "sit," such as sitting with legs crossed 21a, sitting in a squatting position 21b, and sitting with legs together 21c.
[0125] As a result, the 3D object generation system 100 can check the motion information corresponding to "legs" and "crossed legs" based on the motion level from among the multiple motion information retrieved in response to "sitting," and acquire the confirmed motion information. In this way, when keywords divided into multiple motion levels are generated, the 3D object generation system 100 can acquire the motion information most suitable for the multiple keywords by performing a search on the database 131 while considering the motion level.
[0126] Referring to Figure 9, another example is that when the keywords "sit" and "stand up" are generated, the 3D object generation system 100 can search the database 131 for the action information 21 corresponding to "sit" based on the action sequence set for the keywords, and obtain the retrieved action information 21. Subsequently, the 3D object generation system 100 can search the database 131 for the action information 22 corresponding to "stand up" based on the action sequence, and obtain the retrieved action information 22.
[0127] As a result, the 3D object generation system 100 can obtain multiple pieces of action information corresponding to a sentence by sequentially arranging the action information 21 related to "sitting" and the action information 22 related to "standing up".
[0128] As another example, if the 3D object generation system 100 generates the keywords "fun" and "run," it can search and retrieve facial expression information corresponding to "fun" from the database 131 and retrieve motion information corresponding to "run."
[0129] Furthermore, if keywords related to the outline of the 3D model 1 are generated in response to a text, the 3D object generation system 100 can retrieve the 3D model 1 corresponding to the keywords related to the outline from the database 131.
[0130] For example, if the 3D object generation system 100 generates the keywords "person" and "run," it can obtain a 3D model corresponding to the keyword related to "person" and motion information corresponding to the keyword "run" from the database 131.
[0131] As another example, when the 3D object generation system 100 generates the keywords "cat," "fence," and "jump over" (or "fly" and "jump over"), it can obtain from the database 131 a 3D model corresponding to the keyword "cat," and action information corresponding to the keywords "fence" and "jump over" (or "fly" and "jump over").
[0132] Another example is that when the keywords "boy," "girl," "umbrella," "down," and "stand" are generated, the 3D object generation system 100 can retrieve 3D models corresponding to the keywords related to "boy" and "girl" from the database 131, and retrieve action information corresponding to the keywords "umbrella," "down," and "stand." In this case, the 3D object generation system 100 may be implemented to retrieve action information from the database 131 that corresponds to the keywords "umbrella," "down," and "stand" among the action information related to "boy" and "girl."
[0133] Furthermore, when a storyboard related to multiple actions is defined in correspondence with a text, the 3D object generation system 100 can retrieve action information corresponding to each of the multiple actions from the database 131 and sequentially link the retrieved action information according to the action order specified based on the storyboard.
[0134] Referring to Figure 4, the 3D object generation system 100, upon acquiring multiple pieces of operation information from the database 131 (S310), can sequentially concatenate the acquired pieces of operation information according to the operation order specified based on the text analysis results (S320).
[0135] To this end, the 3D object generation system 100 can sequentially link multiple different pieces of motion information by, for two pieces of motion information obtained from the database 131 that are adjacent in motion order, linking the final frame of the preceding motion information with the first (1st) frame of the succeeding motion information.
[0136] At this time, the 3D object generation system 100 can generate linked operation information for two adjacent operation information sets so that operations based on multiple operation information sets obtained from the database 131 are naturally linked (S330).
[0137] Linked operation information is a type of operation information that may include information related to the operations performed between a preceding operation and a succeeding operation, so that two different operations are linked together naturally.
[0138] To this end, the 3D object generation system 100 can generate linked motion information such that the final pose of the first motion information, which is the first motion in the sequence of motions obtained from the database 131, changes to the first pose (or first frame) of the second motion information, which is the second motion in the sequence of motions.
[0139] Here, the final posture of the first motion information may be the posture in the final frame among the multiple frames (or time) included in the first motion information, and the initial posture of the second motion information may be the posture in the first frame among the multiple frames (or time) included in the second motion information.
[0140] Referring to Figure 10 as an example, the 3D object generation system 100 can generate linked motion information 50 such that the positions of multiple keypoints 40a in the final posture 31 of the first motion information 21 are moved to the positions of multiple keypoints 40b in the initial posture 32 of the second motion information 22.
[0141] At this time, the 3D object generation system 100 can generate linked motion information 50 such that each of the multiple key points 40a corresponding to the final pose of the first motion information changes position to the corresponding key point among the multiple key points 40b corresponding to the initial pose of the second motion information.
[0142] As another example, the 3D object generation system 100 can generate multiple first keypoints for the first pose included in the linked motion information using the positions of multiple keypoints corresponding to the final pose of the first motion information and the center points of the positions of multiple keypoints corresponding to the initial pose of the second motion information; generate multiple second keypoints for the second pose included in the linked motion information using the positions of multiple keypoints of the first motion information and the center points of the positions of multiple first keypoints; and generate multiple third keypoints for the third pose included in the linked motion information using the positions of multiple first keypoints and the center points of the positions of multiple keypoints of the second motion information.
[0143] As described above, the 3D object generation system 100 can repeat the process of generating multiple keypoints for multiple poses included in linked motion information using the center points between keypoints included in each other's different motions.
[0144] As a result, when multiple poses included in the linked motion information are generated, the 3D object generation system 100 can determine the order of the multiple poses included in the linked motion information based on the proximity of the positions of the multiple key points corresponding to the final pose of the first motion information and the positions of the multiple key points corresponding to the multiple poses included in the linked motion information, and generate the linked motion information according to the determined order.
[0145] Alternatively, when multiple poses included in the linked motion information are generated, the 3D object generation system 100 can determine the order of the multiple poses included in the linked motion information based on the order in which the positions of the multiple key points corresponding to the first pose of the second motion information and the positions of the multiple key points corresponding to the multiple poses included in the linked motion information are farther apart, and generate the linked motion information according to the determined order.
[0146] Another example is that the 3D object generation system 100 generates a connecting motion line by connecting the positions of multiple key points corresponding to the final pose of the first motion information with the positions of multiple key points corresponding to the initial pose of the second motion information, and generates multiple key points on the generated connecting motion line based on a predetermined interval (or predetermined number) between multiple key points of the first motion information and multiple key points of the second motion information.
[0147] As a result, the 3D object generation system 100 can determine the order of multiple poses included in the linked motion information based on the order in which their positions are closest to multiple key points in the first motion information, and generate the linked motion information according to the determined order.
[0148] Alternatively, the 3D object generation system 100 can determine the order of multiple poses included in the linked motion information based on the order in which their positions are furthest from multiple keypoints in the second motion information, and generate the linked motion information according to the determined order.
[0149] Furthermore, the 3D object generation system 100 can determine the operation speed of linked operation information by considering the first operation speed of the first operation information, which is an operation preceding the linked operation information, and the second operation speed of the second operation information, which is an operation following the linked operation information, during the process of generating linked operation information.
[0150] Referring to Figure 11 as an example, the 3D object generation system 100 can determine a first motion speed based on the distance difference 41 between the positions of multiple key points 40a in the final posture 31 of the first motion information and the positions of multiple key points 40c in the posture 33 immediately preceding the final posture 31, and determine a second motion speed based on the distance difference 42 between the positions of multiple key points 40b in the initial posture 32 of the second motion information and the positions of multiple key points 40d in the posture immediately following the initial posture 32.
[0151] As a result, the 3D object generation system 100 can determine the operating speed of the linked motion information such that the closer the orientation is to the final orientation 31 among the multiple orientations included in the linked motion information, the closer the operating speed is to the first operating speed, and the closer the orientation is to the initial orientation 32, the closer the operating speed is to the second operating speed.
[0152] As another example, the 3D object generation system 100 can calculate the center points of the key points in the final pose 31 and the key points in the initial pose 32 by applying the ratio of the first motion speed to the second motion speed in the process of calculating the center points of the key points in the final pose 31 and the key points in the initial pose 32, respectively.
[0153] As a result, the 3D object generation system 100 can generate multiple poses included in the linked motion information by repeatedly generating at least one pose close to the final pose 31 and at least one pose close to the initial pose 32, with respect to the center point to which the motion speed is applied. As a result, the 3D object generation system 100 can generate one or more poses from the final pose 31 and the initial pose 32 in which the distance between keypoints is farther the closer the pose is to the pose with a faster motion speed, and one or more poses in which the distance between keypoints is closer the closer the pose is to the pose with a slower motion speed. As a result, the 3D object generation system 100 can sequentially arrange the multiple poses generated through the above process and generate linked motion information.
[0154] Another example is that the 3D object generation system 100 generates a first posture included in the linked motion information by calculating the center points of multiple key points corresponding to the final posture 31 and the multiple key points corresponding to the initial posture 32, and can determine the number of postures generated between the final posture 31 and the first posture, and the number of postures generated between the initial posture 32 and the second posture, based on the ratio of the first motion speed to the second motion speed.
[0155] As another example, the 3D object generation system 100 generates a connecting motion line by connecting the positions of multiple keypoints corresponding to the final pose 31 and the positions of multiple keypoints corresponding to the initial pose 32. On the generated connecting motion line, the system can generate poses included in the connecting motion information at intervals where the difference from the distance difference 41 corresponding to the first motion speed is small as the pose approaches the final pose 31, and generate poses included in the connecting motion information at intervals where the difference from the distance difference 42 corresponding to the second motion speed is small as the pose approaches the initial pose 32.
[0156] Furthermore, referring again to Figure 4, the 3D object generation system 100 can sequentially link multiple motion information and insert linking motion information between each motion information to connect them (S340).
[0157] Referring to Figure 12, for example, the 3D object generation system 100 can acquire first action information 21 for sitting down and second action information 22 for standing up in response to the sentence "sit down and then stand up," and generate linked action information 50 for the first action information 21 and the second action information 22. In this way, the 3D object generation system 100 can link the first action information, linked action information, and second action information by linking the final posture of the first action information 21 with the initial posture of the linked action information 50, and linking the final posture of the linked action information 50 with the initial posture of the second action information 22.
[0158] With the configuration described above, the 3D object generation system 100 according to the present invention can obtain at least one piece of operational information from the database 131 based on at least one keyword generated in correspondence with a sentence.
[0159] Furthermore, the 3D object generation system 100 according to the present invention can acquire information that constitutes the external shape and facial expression of the 3D object 3 based on at least one keyword generated in correspondence with text.
[0160] Furthermore, the 3D object generation system 100 according to the present invention can naturally link different operations by generating linked operation information for multiple operation information whose operation sequence is consecutive.
[0161] Referring again to Figure 3, the 3D object generation system 100 according to the present invention can generate a 3D object 3 that performs an action corresponding to a sentence by applying action information to a 3D model 1 (S400).
[0162] Specifically, the 3D object generation system 100 can apply the position of at least one key point specified in the first frame among a plurality of frames (or time) included in the motion information to at least one key point corresponding to the aforementioned at least one key point among a plurality of key points included in the 3D model 1 which is provided in advance (or obtained from the database 131).
[0163] Subsequently, the 3D object generation system 100 applies the position of at least one key point specified in the second frame, which is provided immediately after the first frame, to at least one key point among the multiple key points included in the 3D model 1 that corresponds to the aforementioned at least one key point, and repeats the above process for all frames specified in the motion information, thereby applying the motion information to the 3D model 1.
[0164] Referring to Figures 13 and 14, the 3D object generation system 100 can apply the positions of multiple key points 40 corresponding to each pose 35 to multiple key points included in the 3D model 1. At this time, the 3D object generation system 100 can generate a 3D object 3 that performs an action corresponding to the text by setting the multiple key points included in the 3D model 1 to move sequentially according to the poses of the motion information linked above.
[0165] With the configuration described above, the 3D object generation system 100 according to the present invention can generate 3D objects 3 that perform actions corresponding to text. Furthermore, the 3D object generation system 100 can generate more naturally moving 3D objects 3 by generating and applying linked actions between the different actions performed by the 3D model 1.
[0166] On the other hand, the 3D object generation system 100 can also generate 3D objects from 2D images. For example, when a 2D image is input to the 3D object generation system 100, it can detect objects from the 2D image, then identify the parts corresponding to the objects through object segmentation, and extract data related to the object's pose for each frame of the 2D image. As a result, the 3D object generation system 100 can collect and integrate the pose data extracted from each frame of the 2D image for each angle, and based on the integrated result, generate a 3D object that expresses volume and texture.
[0167] In one embodiment, the generation of a 3D object can be performed based on a 3-Dimensional Volume Metric. The 3-Dimensional Volume Metric can be implemented by setting the initial position of the object in a 2D image as a reference point, then estimating the size of the object by capturing the positional information of the object when it moves to another position in the 2D image, and then generating a volume in the form of a 3D mesh.
[0168] More specifically, the 3D object generation system 100 can receive a 2D video input and output a 3D object in mesh format based on the volume and texture of the object in the 2D video. In this case, the input 2D video may include multiple frames in which the shape of the object is represented at various angles and / or positions.
[0169] The 3D object generation system 100 can analyze key points and lines representing an object from a 2D image and, based on this, generate a 3D point cloud as shape information for the object.
[0170] The 3D object generation system 100 can generate 3D mesh objects by performing 3D mesh modeling and texture mapping based on the generated 3D point cloud.
[0171] The 3D object generation system 100, upon receiving a 2D image, can track key points of objects within the 2D image 41 and reconstruct a 3D model based on these key points. Simultaneously, the 3D object generation system 100 can detect and fit lines of objects within the 2D image and reconstruct a 3D model based on these lines. The 3D object generation system 100 may also use predetermined reconstruction parameters during the reconstruction process.
[0172] Furthermore, the 3D object generation system 100 can generate a 3D point cloud (PC) that estimates the 3D shape of an object in a 2D image using a 3D model reconstructed based on keypoints and / or lines.
[0173] The 3D object generation system 100 can perform 3D mesh modeling based on a 3D point cloud (PC). The result of the 3D mesh modeling may be output as mesh information (MI) representing the volume of the 3D object. The 3D object generation system 100 may also use pre-determined modeling parameters for 3D mesh modeling. In addition, the 3D object generation system 100 performs texture mapping based on a color image of a 2D image. The result of the texture mapping may be output as texture information (TI) representing the texture of the 3D object.
[0174] Furthermore, the 3D object generation system 100 can generate 3D objects based on data related to a 3D mesh after constructing data using mesh information (MI) and texture information (TI).
[0175] As a result, the 3D object generation system 100 can integrate the motion information generated above into the 3D object, enabling the 3D object to perform 3D motion based on the motion information.
[0176] In this case, the 3D object may be a 3D mesh object in 3D mesh format, or it may be a pre-stored 3D character object. Alternatively, if the 3D object generation system 100 uses a 3D character object modeled after a famous person or a virtual character instead of a 3D mesh object, by integrating motion information into the 3D character object, it can generate a 3D object that reproduces the movement of an object in a 2D image.
[0177] In this regard, since motion information is data created based on text or 2D images, and 3D objects are objects generated independently of text or 2D images, inconsistencies may occur between motion information and 3D objects. Therefore, the 3D object generation system 100 may perform retargeting to adjust the motion information to match the 3D objects in order to resolve such inconsistencies.
[0178] For example, the 3D object generation system 100 may align the joint positions of the motion information with the joint positions of the 3D object based on data related to key points of the motion information. Here, aligning the joint positions of the motion information with the joint positions of the 3D object may mean, for example, adjusting the joint positions of the motion information so that they coincide with the joint positions of the 3D object.
[0179] For example, when applying motion information to a 3D object, the physical characteristics of the model on which the motion information is based may differ from those of the 3D object. Therefore, to eliminate inconsistencies caused by differences in physical characteristics between objects, the positions of each joint (i.e., keypoints) in the motion information can be adjusted to match the positions of the joints (i.e., keypoints) of the 3D object.
[0180] In this case, the lines of the operation information can also be adjusted based on the adjusted key points of the operation information. For example, assuming that the operation information extracted from the database includes key points K1(1,1,1) and K2(2,2,2), and a line L1 connecting key points K1 and K2, retargeting can be used to adjust key points K1 and K2 to K1(1,2,2) and K2(2,3,3), respectively, thereby adjusting line L1 to connect (1,2,2) and (2,3,3).
[0181] Furthermore, the 3D object generation system 100 can realize the movement of a 3D object in a manner that mimics the movement of motion information, based on the joint positions of the consistent motion information.
[0182] To this end, the 3D object generation system 100 can adjust the movement of motion information, such as the position, movement distance, and range of motion of each key point and line, based on the adjusted joint positions (i.e., key points) and lines of the motion information. Furthermore, by integrating the adjusted movement of motion information into the 3D object, the movement of the 3D object can be realized.
[0183] The 3D object generation system 100 can output and store the movement of the realized 3D object.
[0184] On the other hand, summarizing the present invention from an operational standpoint, the 3D object generation system 100 according to the present invention can identify a text based on user input and analyze the identified text. Furthermore, the 3D object generation system 100 can define a plurality of actions corresponding to the text based on the analysis results and define a linking action that connects the plurality of actions corresponding to the text. Moreover, the 3D object generation system 100 can generate a moving 3D object that performs the plurality of actions and the linking action using a trained artificial neural network.
[0185] Here, the linking operation may be a linking operation that connects at least two operations that are adjacent to each other in a time-series sequence from among the plurality of operations, as described above.
[0186] Furthermore, the movement of the 3D object may be formed by a plurality of frames corresponding to the first movement, each containing keypoint information of the 3D model corresponding to the first movement among the plurality of movements; a plurality of frames corresponding to the second movement, each containing keypoint information of the 3D model corresponding to the second movement among the plurality of movements; and a plurality of frames corresponding to the connecting movement, each containing keypoint information of the 3D model corresponding to the connecting movement that connects the first movement and the second movement.
[0187] In other words, a three-dimensional object can be understood as a shape in which the three-dimensional model moves as a result of the playback of multiple frames corresponding to the first action, the second action, and the linked action, respectively.
[0188] As described above, each of the multiple frames corresponding to the multiple operations and linked operations, for example, the multiple frames corresponding to the first operation, the multiple frames corresponding to the second operation, and the multiple frames corresponding to the linked operations, may include positional information relative to the key points of the 3D model.
[0189] Furthermore, if, among the multiple frames, the actions are adjacent to each other along the time-series flow, for example, the first action and the second action are adjacent to each other along the time-series flow, the key points of the 3D model of the multiple frames corresponding to the linked action may be defined based on the key point information of the 3D model of the multiple frames corresponding to the first action and the multiple frames corresponding to the second action.
[0190] In this case, if the first operation precedes the second operation based on the time series, the position information of the key points of the 3D model included in the first frame of the multiple frames corresponding to the linked operation may relate to the position information of the key points of the 3D model included in the last frame of the multiple frames corresponding to the first operation, and the position information of the key points of the 3D model included in the last frame of the multiple frames corresponding to the linked operation may relate to the position information of the key points of the 3D model included in the first frame of the multiple frames corresponding to the second operation. Here, the relationship between the position information of key points may mean that the position information of key points corresponds to the same or close positions to each other.
[0191] For example, the position of the key point of the 3D model included in the first frame of the linking operation may be the same as or adjacent to the position of the key point of the 3D model in the final frame of the first operation, and the position of the key point of the 3D model included in the final frame of the linking operation may be the same as or adjacent to the position of the key point of the 3D model in the first frame of the second operation.
[0192] As described above, the 3D object generation method and system according to the present invention defines actions corresponding to the meaning of a sentence and generates a 3D object that performs the defined action. This makes it possible to generate a moving 3D object simply by inputting a sentence, without the user having to define the actions of the 3D object each time.
[0193] Furthermore, the 3D object generation method and system according to the present invention can acquire information that realizes various elements of a 3D model corresponding to a 3D object, such as the movement, facial expression, external form, age, clothing, gender, physical condition, background, location, and situation, based on keywords identified in correspondence with text.
[0194] Furthermore, according to various embodiments of the present invention, the 3D object generation method and system can naturally link different actions performed by a 3D object by generating linked action information for a plurality of action pieces whose action sequences are consecutive.
[0195] Furthermore, the present invention described above can be embodied as computer-readable code or instruction words on a medium on which a program is recorded. That is, the various control methods according to the present invention can be provided in the form of a program, either integrated or individually.
[0196] On the other hand, computer-readable media include all types of recording devices that store data readable by a computer system. Examples of computer-readable media include HDDs (Hard Disk Drives), SSDs (Solid State Disks), SSDs (Silicon Disk Drives), ROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices.
[0197] Furthermore, the computer-readable medium may include storage and may be a server or cloud storage accessible by electronic devices via communication. In this case, the computer can download the program according to the present invention from the server or cloud storage via wired or wireless communication.
[0198] Furthermore, in this invention, the computer described above is an electronic device equipped with a processor, i.e., a CPU (Central Processing Unit), and its type is not particularly limited.
[0199] On the other hand, the above detailed description should not be interpreted restrictively in any way, but should be considered illustrative. The scope of the invention should be determined by a reasonable interpretation of the appended claims, and all modifications within the scope of the equivalents of the invention are included within the scope of the invention.
Claims
1. Steps to identify text based on user input, The steps include analyzing the identified text and defining a number of actions corresponding to the text based on the analysis results, A step of defining a linking operation that links the multiple operations corresponding to the above sentence, A method for generating a three-dimensional object, characterized by comprising the step of generating a moving three-dimensional object that performs the plurality of operations and the concatenation operations using a trained artificial neural network.
2. The method for generating a three-dimensional object according to claim 1, characterized in that the linking operation is a linking operation that links at least two operations that are adjacent to each other in a time series from among the plurality of operations.
3. The movement of the aforementioned three-dimensional object is, Among the aforementioned multiple operations, a plurality of frames corresponding to the first operation, which include key point information of a three-dimensional model corresponding to the first operation, Among the aforementioned multiple operations, a plurality of frames corresponding to the second operation, which include key point information of the three-dimensional model corresponding to the second operation, A method for generating a three-dimensional object according to claim 2, characterized in that it is formed by a plurality of frames corresponding to the linking operation, each containing key point information of the three-dimensional model corresponding to the linking operation that links the first operation and the second operation.
4. The method for generating a three-dimensional object according to claim 3, characterized in that each of the multiple frames corresponding to the first operation, the multiple frames corresponding to the second operation, and the multiple frames corresponding to the linking operation includes positional information with respect to the key points of the three-dimensional model.
5. When the first and second operations are adjacent operations that follow a time-series flow, the key points of the three-dimensional model of the plurality of frames corresponding to the linked operations are: The method for generating a three-dimensional object according to claim 4, characterized in that it is defined based on keypoint information of the three-dimensional model of a plurality of frames corresponding to the first operation and a plurality of frames corresponding to the second operation.
6. Based on the timeline, if the first action precedes the second action, The position information of the key points of the 3D model included in the first frame of the multiple frames corresponding to the aforementioned linking operation is: In relation to the position information of the key points of the three-dimensional model included in the final frame among the multiple frames corresponding to the first operation, The positional information of the key points of the 3D model included in the final frame among the multiple frames corresponding to the aforementioned linking operation is: The method for generating a three-dimensional object according to claim 5, characterized in that it relates to the position information of a key point of the three-dimensional model included in the first frame of a plurality of frames corresponding to the second operation.
7. The position of the key point of the three-dimensional model included in the first frame of the coupling operation is the same as or adjacent to the position of the key point of the three-dimensional model in the final frame of the first operation. The method for generating a three-dimensional object according to claim 6, characterized in that the position of the key point of the three-dimensional model included in the final frame of the linking operation is the same as or adjacent to the position of the key point of the three-dimensional model in the first frame of the second operation.
8. A storage unit including a pre-established database, A three-dimensional object generation system comprising: a control unit that identifies text based on user input, analyzes the identified text, defines a plurality of actions corresponding to the text based on the analysis results, defines a linking action that links the plurality of actions corresponding to the text, and generates a moving three-dimensional object that performs the plurality of actions and the linking action using a trained artificial neural network.
9. A program executed by one or more processes in an electronic device and stored on a computer-readable recording medium, The aforementioned program, Steps to identify text based on user input, The steps include analyzing the identified text and defining a number of actions corresponding to the text based on the analysis results, A step of defining a linking operation that links the multiple operations corresponding to the above sentence, A program stored on a computer-readable recording medium, characterized by including a command word that causes a trained artificial neural network to generate a moving three-dimensional object that performs the plurality of operations and the linked operations.
10. Steps to identify text based on user input, The steps include analyzing the identified text and defining a storyboard corresponding to the text based on the analysis results, The steps include: using a pre-established database to acquire a 3D model and motion information corresponding to the storyboard; A method for generating a three-dimensional object, comprising the step of generating a three-dimensional object that performs an action corresponding to the text by applying the action information to the three-dimensional model.
11. The steps of acquiring the three-dimensional model and motion information are as follows: A method for generating a three-dimensional object according to claim 10, comprising the steps of: defining a story related to a plurality of actions in correspondence with the aforementioned text; obtaining action information corresponding to each of the plurality of actions from the database; and sequentially concatenating the obtained action information according to the action order specified based on the story.
12. The step of sequentially linking the aforementioned operation information is: A method for generating a three-dimensional object according to claim 11, further comprising the step of generating linked operation information such that the final posture of the first operation information, which is the preceding operation among the plurality of operations, changes to the initial posture of the second operation information, which is the subsequent operation.
13. The step of generating the aforementioned coupling operation information is: A method for generating a three-dimensional object according to claim 12, comprising the step of determining the operating speed of the linked operation information, taking into consideration the first operating speed of the first operation information and the second operating speed of the second operation information, among the plurality of operations.
14. The steps of acquiring the three-dimensional model and motion information are as follows: If the storyboard includes multiple elements related to the three-dimensional object, the process involves matching each of the multiple keywords corresponding to the storyboard to one of the multiple categories predetermined by the multiple elements. A method for generating a three-dimensional object according to claim 10, comprising the step of obtaining information corresponding to each keyword from a database based on the matching categories.
15. The step of identifying the aforementioned text is, Steps in which data is entered based on user input, The steps include analyzing the aforementioned data and converting it into text, A method for generating a three-dimensional object according to claim 10, comprising the step of identifying the text such that a storyboard is defined based on the converted text.
16. The step of defining the aforementioned storyboard is: The steps include dividing the identified sentence into multiple words of morpheme unit, The steps include confirming the meanings corresponding to the aforementioned group of words, A method for generating a three-dimensional object according to claim 10, comprising the step of generating at least one keyword corresponding to the meaning of the identified sentence based on the confirmed meaning.
17. The step of defining the aforementioned storyboard is: The steps include: confirming the correlation between multiple words that are different from each other based on the meaning of each of the multiple words separated from the aforementioned text; A method for generating a three-dimensional object according to claim 16, comprising the step of determining at least one of the operation levels and operation sequences for the at least one keyword based on the confirmed correlation.
18. A storage unit including a pre-established database, A three-dimensional object generation system comprising: a control unit that identifies text based on user input, analyzes the identified text, defines a storyboard corresponding to the text based on the analysis results, acquires a three-dimensional model and operation information corresponding to the storyboard using a pre-established database, and generates a three-dimensional object that performs the operation corresponding to the text by applying the operation information to the three-dimensional model.
19. A program executed by one or more processes in an electronic device and stored on a computer-readable recording medium, The aforementioned program, Steps to identify text based on user input, The steps include analyzing the identified text and defining a storyboard corresponding to the text based on the analysis results, The steps include: using a pre-established database to acquire a 3D model and motion information corresponding to the storyboard; A program stored on a computer-readable recording medium, characterized by including a command word that causes the program to perform the steps of: applying the operation information to the three-dimensional model to generate a three-dimensional object that performs an action corresponding to the text.